Dutch neuromorphic computing startup Innatera has launched the production version of its microcontroller, Pulsar, with both analog and digital spiking neural network accelerator fabrics, designed to take on processing of sensor data in endpoint devices.
Versus the T1, a pre-production device announced a year ago, Pulsar has added an FFT accelerator, a power management unit that allows the device to go into different power saving or deep sleep states, and several new interfaces, including a camera parallel interface. The overall pipeline has also been streamlined, Innatera co-founder and CEO Sumeet Kumar told EE Times.
“Sensor applications are notoriously power-constrained,” Kumar said. “Very often what developers need to do is trade off between application complexity, accuracy, and power dissipation.”
Sumeet Kumar (Source: Innatera)
Like the T1, Innatera claims Pulsar can lower latency up to 100× and lower energy consumption up to 500×, compared to deep neural networks implemented on mainstream digital hardware, without degradation of accuracy.
Innatera calls its Pulsar design a microcontroller; there are actually multiple cores and accelerators on chip, including its spiking neural network fabric. It can replace application processors in some cases, especially for always-on sensing as it consumes less than a milliwatt (600 µW for radar-based presence detection and 400 µW for audio scene classification).
Innatera’s main secret sauce is its accelerator for spiking neural networks (SNNs), which has both analog and digital fabrics.
“The idea is that as an application developer, you have the option to pick whether to deploy on the analog or the digital fabric,” he said. “The key considerations are how aggressive your power constraints are…and at what timescale the features occur inside your input data.”
The analog fabric is a massively parallel array of neurosynaptic cores that can be configured for different networks, taking up about a quarter of the silicon area of the Pulsar chip. This fabric allows inference in under a milliwatt and under a millisecond. However, it is best suited to applications with fast-moving signals, as analog states cannot be maintained indefinitely, even with buffering. Analog therefore suits applications like audio streaming.
Slower-moving patterns, that perhaps need to be recognized in the hundreds of milliseconds timeframe, are better suited to the digital fabric—particularly if a bigger SNN is needed and the power budget is a little more relaxed, Kumar said.
Innatera’s Pulsar includes analog and digital spiking neural network accelerators plus accelerators for CNNs, spike encoding and decoding, and FFT. (Source: Innatera)
Both the analog and the digital fabrics accept spikes, run inference and output a precisely timed spike, he said, but the datapaths are different. The analog fabric is a current-mode engine; spikes are converted to current, which is manipulated all the way through the data path. The digital fabric does not manipulate current. Rather, data is converted to digital values and an array of synapses, each with its own weight memory, are activated by incoming spikes. When spikes arrive, packets of charge are produced, which are accumulated until a threshold is met, when an output spike is fired.
While the digital fabric is synchronous (it has a clock), processing happens in an asynchronous way to allow for the required timing properties of spiking networks.
“[Digital] neurons don’t need to be synchronized with one another,” Kumar said. “They’re not waiting for one another to fire, they fire discretely any time they see data which is relevant to them. Firing is independent to all other neurons in the system. So, the underlying fabric is synchronous, but the processing style on top of it is completely asynchronous.”
Hardware encoders
Pulsar includes hardware encoders and decoders—essential for converting data into and out of the spiking domain (custom encoders could also be implemented on the chip’s CPU if required, Kumar said).
Also on-chip are hardware CNN and FFT accelerators, widely used in signal processing stages that might be used as part of a sensor data processing pipeline, and a variety of sensor interfaces.
A RISC-V CPU is used for housekeeping, for implementing custom functions that are not covered by the on-chip accelerators, and any control or logic tasks that might happen before or after inference.
“The main reason we put all of these compute fabrics on the chip is to provide developers with the flexibility that they need to pick and choose the right tool for their job,” Kumar said. “Real-world pipelines are often a combination of different models and they run in a different sequence, so having very efficient heterogeneous acceleration capabilities on chip allows the developer to do whatever they want with it.”
Diagram showcasing inference power and latency performance for Innatera’s Pulsar versus existing solutions.
Inference power and latency performance for Innatera’s Pulsar versus existing solutions (Source: Innatera)
Software stack
Innatera’s software stack, Talamo, has been further built out to interface directly with Pytorch; the company has built a Pytorch extension for SNNs. A Pytorch-based simulator can enable power consumption estimations from the early stages. Talamo is also compatible with TensorFlow.
“Most importantly, we’ve managed to achieve ease of development,” Kumar said. “You no longer need a neuromorphic Ph.D. to be building spiking neural networks and running them on our chips—and this is actual feedback that has come from customers that have been sampling the hardware and the software through 2024.”
On Talamo’s roadmap are features set to make model porting even simpler, Kumar said.
Customer traction
Innatera technology is already in customer deployments, with a “significant” pipeline, Kumar said. Customer applications are across presence detection, people counting, anomaly detection and classification and vital signs monitoring.
Image of Pulsar evaluation kit.
Pulsar evaluation kit. (Source: Innatera)
A reference design built with Socionext performs human presence detection using radar to avoid false positives because it does not use motion as a proxy for presence. This design could be used in home appliances to avoid detecting pets, while making sure to detect humans who are not moving, he said. This is useful in smart doorbells and other appliances that switch off when a person is not present. Motion detection using radar usually requires 14 mW of power for processing; Innatera’s presence detection uses just 0.6 mW on average (power draw is higher during inference).
Another reference design uses a Melexis infrared sensor for people counting. The target application is a smoke detector for hotels that counts how many people are in the room in the event of a fire, without invading privacy, but it can be used in any occupancy application, Kumar said. The solution can detect humans while filtering out other heat sources like laptops or coffee cups.
Other customers are using Innatera chips for audio scene classification and keyword spotting.
“There are some sensor customers that will be integrating these chips into their own modules to go out into the market,” Kumar said. “All of this has been the progress since last year, so it’s a very, very promising pipeline. It’s it’s a lot more tangible than it was a few years ago.”
Spiking networks
Biology-inspired SNNs are still a very active field of research. Are we truly at the stage where SNNs are ready for commercial adoption?
“The neuromorphic space is moving pretty quickly,” Kumar said. “In 2020, training spiking neural networks was still quite hard. In the years after that, a number of new algorithms came out, which greatly simplified training, and made it as simple as it is for conventional, non-spiking models today.”
While the spiking space is still evolving quickly, Kumar feels confident that Innatera’s hardware is sufficiently programmable to cope with even novel types of spiking networks today and in the future. The fundamentals of spiking networks—dynamical systems of discrete compute elements with time-based processing, with complex relationships between neurons (like delays)—have always been present in the company’s architecture, he said.
Image of Pulsar, Innatera’s spiking microcontroller.
Pulsar, Innatera’s spiking microcontroller. (Source: Innatera)
“These are temporal machines, which are inherently asynchronous, you have internal recurrence inside the neurons, you have the ability to create recurrent connections inside the topology, you can do [novel] topologies [outside of feedforward or recurrent networks]…you have the ability to do all that and train them, and do something meaningful,” he said.
The main limitation of the Innatera fabric is that it is not self-learning, Kumar said, noting that the neuron types are fixed, chosen based on their suitability for a wide range of pattern recognition operations across a wide range of sensor types. While functions cannot be changed, parameters can, he said. This programmability is sufficient to do a large set of pattern recognition and signal processing operations very efficiently.
“The key questions are really, can you represent that network well enough in Python, and can you train it with the infrastructure you’ve got?” he said. “As long as you can answer yes to both questions, then yes, you can implement your network on our chip.”
The next generation of Innatera hardware will continue to add programmability, with the ultimate ambition of integrating even more closely with the sensor, Kumar said.
“If you look into the brain of a locust, it has a natural ability to filter the sound of predators by implementing an FFT very efficiently using spiking neurons,” he said. “That’s really cool, because now there’s a role model for doing signal processing that we spend a lot of energy on in the electronics world with spiking neural networks.”
Future Innatera chips will tap into functions like this to move into spaces like control and learning in robotics, Kumar said.
Pulsar will ramp to volume production by the end of the year.
