
AI Chips for Edge Computing: A Buyer's Guide to NPUs, TPUs, and Inference Accelerators

AI Chips for Edge Computing: A Buyer's Guide to NPUs, TPUs, and Inference Accelerators
Edge AI is moving compute off the cloud and onto the device itself. For buyers, that means a whole new set of chips to source — NPUs, AI accelerators, inference-optimized SoCs — many of which we rarely touched five years ago. This is the category we get asked about most lately, so here's what we've learned ordering them.
The Edge AI Chip market
We split the edge AI chip market into three tiers when customers ask what to buy:
| Tier | Performance | Example Chips | Application |
| Low-power MCU | <1 TOPS | STM32N6, ESP32-S3, GAP9 | Sensor fusion, keyword spotting |
| Mid-range SoC | 1-10 TOPS | NXP i.MX 93, Rockchip RK3588, MediaTek Genio | Camera analytics, industrial HMI |
| High-end accelerator | 10-100+ TOPS | Hailo-8, Intel Movidius, NVIDIA Jetson Orin | Multi-camera, autonomous machines |
The tiers aren't a ladder you climb up as your budget grows. Each one exists because the application demands it — a doorbell camera does not need a Jetson, and an autonomous machine can't make do with an MCU.
Key Specs to Compare
TOPS (Tera Operations Per Second): The headline number, but not the full story. TOPS/Watt matters more for battery-powered devices. We've had buyers fixate on raw TOPS, then watch their handheld product die in two hours — check the efficiency figure first.
Memory bandwidth: AI inference is memory-bound. A chip with 4 TOPS and 50 GB/s bandwidth often outperforms one with 10 TOPS and 20 GB/s. In our benchmarks, bandwidth is usually where the "slow" chip reveals itself.
Supported frameworks: TensorFlow Lite, ONNX Runtime, and PyTorch Mobile are the standards. Check framework support before committing. This one is easy to skip — until your firmware team hits a wall with the SDK.
Quantization: INT8 inference is 2-4x faster than FP16 with minimal accuracy loss. Most edge AI workloads use INT8. If your model can't quantize cleanly, that's a problem to catch early, not late.
Popular Edge AI Chips
| Chip | TOPS | Power | Interface | Best For |
| Hailo-8L | 13 | 2.5W | PCIe/M.2 | Vision AI, multi-stream |
| Google Coral TPU | 4 | 1.5W | USB/PCIe/M.2 | Prototyping, low volume |
| Intel Movidius Myriad X | 1 | 1W | USB | Computer vision |
| NVIDIA Jetson Orin Nano | 40 | 7-15W | SO-DIMM | Robotics, autonomous |
| STM32N6 (Neural-ART) | 0.6 | mW | MCU on-chip | Always-on sensor AI |
| Kneron KL720 | 1.4 | 1.5W | M.2 | Audio + vision |
We keep this table on the wall. The Coral is still our default suggestion for prototyping, and most robotics builds we've seen end up on the Orin Nano.
Chinese Alternatives
Chinese AI chip makers are growing fast:
- Horizon Robotics Journey series — Automotive ADAS, up to 128 TOPS
- Rockchip RK3588 — Integrated 6 TOPS NPU, widely used in edge boxes
- Sophgo BM1684 — 17.6 TOPS, popular in surveillance
- Axera AX630A — 3.6 TOPS, ultra-low cost for consumer cameras
For cost-sensitive builds with tight schedules, we usually spec these first. A customer recently asked us to find an alternative to an accelerator with a long lead time — the Rockchip route had them in production while the original part was still weeks out.
Sourcing Considerations
Lead times: Edge AI chips from NVIDIA and Hailo can be 12-20 weeks. Chinese alternatives are typically 4-8 weeks. Plan around this: we order long-lead AI parts at the start of a project, never at the end.
Development boards: Always buy the dev kit first. AI chip software stacks (drivers, SDKs, model converters) vary dramatically in quality. The cheapest way to kill a project is picking a chip whose toolchain nobody on the team can actually drive.
Memory pairing: Many AI accelerators need companion LPDDR4/5. Budget for the full BOM, not just the accelerator chip. We've seen plenty of BOMs blow up when the memory and power parts get added at the last minute.
Thermals: Edge AI chips running sustained inference generate significant heat. Budget for heatsinking in enclosure design. An enclosure that looked fine on paper often needs a rework once the model runs continuously.
Related Articles
- Edge Computing Hardware Guide: Gateways, Industrial PCs, Embedded Controllers — Hardware tiers for edge deployment
- FPGA vs ASIC vs GPU for AI Acceleration — Choose the right AI hardware path
- IoT Module Selection Guide — Wireless connectivity for edge devices
References
Written by Daniel Ross
Quality & Counterfeit Detection Lead · Hong Kong
Daniel runs incoming inspection for PartsCube Global, verifying authenticity of semiconductors and passives before they ship. He spent 8 years in electronics testing and failure analysis.
View all articles by Daniel →Need help sourcing these components?
PartsCube Global stocks all alternatives mentioned in this guide. Search our catalog or submit your BOM for a quote.
Chat on WhatsApp