
On the chip front, AI edge models run on an advanced type of chip called a neural processing unit or NPU. These chips, such as Google’s Tensor Processing Unit (TPU) or Qualcomm’s Snapdragon, are small, efficient, and don’t use as much energy or create as much heat as traditional CPUs or GPUs. Yet they are extremely powerful, able to perform trillions of operations per second (TOPS).
Another breakthrough is the emergence of “neuromorphic” chips, which attempt to mimic how the human brain works. These chips, such as Loihi from Intel and TrueNorth from IBM, spring to action only when a meaningful event occurs, reducing energy consumption. They also deliver advanced data processing capabilities for real-time applications like robotics or autonomous vehicles.
Then there are AI accelerators specifically designed for edge AI deployments from vendors such as Hailo and BrainChip. And a new generation of small language models (SLM), such as Meta’s Llama 3.2, Google’s Gemma 3, and Microsoft’s Phi series, are now available. These SLMs provide strong performance at reduced scale, enabling organizations to train the AI models in the cloud, and then perform inference at the edge.





















