
AI Chip & Accelerator Guide 2026: NPUs, GPUs, Edge AI Processors — Selection and Sourcing

AI Chip & Accelerator Guide 2026: NPUs, GPUs, and Edge AI Processors
AI chips are the fastest-moving segment in semiconductors. From NVIDIA's data-center GPUs to tiny edge NPUs that fit in a sensor, here's the market, key part numbers, and how to source AI accelerators for your project — and what sourcing them actually looks like in 2026.
AI Chip Categories
| Category | Examples | Performance | Power | Use Case |
| Cloud Training GPU | NVIDIA H200, B200 | 1000+ TFLOPS | 700W | LLM training |
| Cloud Inference | NVIDIA L40S, AMD MI300X | 100-500 TFLOPS | 300-450W | Model serving |
| Edge AI Accelerator | Hailo-8, Intel Movidius | 10-50 TOPS | 2.5-15W | Camera AI, robotics |
| Embedded NPU | Rockchip RK3588, Kendryte K230 | 1-6 TOPS | 1-5W | Smart sensors, IoT |
| AI MCU | ESP32-S3, STM32N6 | 0.1-0.6 TOPS | <1W | Keyword spotting, anomaly detection |
Data Center AI GPUs
NVIDIA dominates with >80% market share. Supply is extremely constrained for the latest generations (H200, B200) — export controls restrict shipments to certain regions, and we've watched lead times push past 20 weeks for anyone without allocation.
| GPU | Architecture | Memory | TFLOPS (FP16) | TDP | Availability |
| NVIDIA H200 | Hopper | 141GB HBM3e | 1,979 | 700W | Tight, 20+ week lead |
| NVIDIA L40S | Ada Lovelace | 48GB GDDR6 | 733 | 300W | Better availability |
| NVIDIA B200 | Blackwell | 192GB HBM3e | 2,250 | 1000W | Sampling only |
| AMD MI300X | CDNA3 | 192GB HBM3 | 1,307 | 750W | Improving |
| Intel Gaudi 3 | – | 128GB HBM2e | 1,835 | 600W | Limited supply |
Chinese AI GPUs (domestic alternatives):
- Huawei Ascend 910B — 400W, 256GB HBM2e, comparable to A100 in certain workloads
- Biren BR100 — 7nm, 770W, targets A100-class performance
- Moore Threads MTT S4000 — data center GPU, D3D12/Vulkan support
- MetaX C500 — inference-focused, 300W, PCIe Gen5
Edge AI Accelerators
These chips run trained models locally — no cloud connection needed. Key advantages: low latency, data privacy, always-on. For camera systems, latency is usually the decider — a Hailo-8L at 2.5W beats streaming video to a server.
| Accelerator | TOPS | Power | Interface | Best For |
| Hailo-8L | 13 | 2.5W | M.2 / Mini PCIe | Multi-stream video analytics |
| Hailo-10H | 40 | 7W | PCIe Gen3 | High-res camera AI, ADAS |
| Intel Movidius Myriad X | 4 (FP16) | 2.5W | USB / MIPI | Drones, robotics |
| Google Coral Edge TPU | 4 (INT8) | 1.5W | USB / Mini PCIe / M.2 | Prototyping, small-batch |
| Kneron KL730 | 1.4 | 1.5W | MIPI / SPI | Battery-powered sensors |
| Axelera Metis AIP | 214 (INT8) | 15W | PCIe Gen3 | High-density edge server |
| MemryX MX3 | 5 | 1.5W | M.2 / USB | Low-power always-on |
Chinese edge AI chips:
- Horizon Robotics Journey 6 — 560 TOPS, automotive ADAS, designed for L2+ autonomous driving
- Rockchip RV1106 — 0.5 TOPS NPU, $3-5, optimized for battery-powered AI cameras
- Sophgo BM1684X — 32 TOPS (INT8), PCIe, popular in NVR AI applications
- Axera AX630A — 7.2 TOPS, dual-core A53, smart retail/video AI
Embedded AI Processors with Built-in NPU
For products that need an application processor plus AI acceleration in one chip, the integrated route saves a board revision. We usually go with the RK3588 when 6 TOPS is enough.
| Processor | CPU | NPU | Video | Target Application |
| Rockchip RK3588 | 4×A76 + 4×A55 | 6 TOPS | 8K decode | AI NVR, digital signage |
| NXP i.MX 95 | 4×A55 + M7 | 2 TOPS (eIQ Neutron) | 4K | Automotive, industrial AI |
| TI AM68A | 4×A72 + 4×R5F | 8 TOPS | 4K60 | ADAS, autonomous mobile robots |
| MediaTek Genio 1200 | 4×A78 + 4×A55 | 4.8 TOPS | 4K | Smart home, AI appliances |
| Qualcomm QCS8550 | 8×Kryo (Oryon) | 48 TOPS | 8K | Premium edge AI box |
| Allwinner T527 | 8×A55 | 2 TOPS | 4K | Cost-sensitive AI display |
Chinese embedded AI options:
- Rockchip RK3576 — 4×A72 + 4×A53, 6 TOPS NPU, sub-$20
- Amlogic A311D2 — 4×A73 + 2×A53, 5 TOPS, popular in smart speakers
- Kendryte K230 — RISC-V dual-core, 1.5 TOPS, under $10
AI MCUs: Tiny AI for Sensors
The newest category — microcontrollers with just enough AI for on-sensor inference. A customer recently asked whether keyword spotting still needs a DSP; it doesn't — the ESP32-S3 covers it.
| MCU | Core | AI Capability | Power | Price (1K) |
| STM32N6 | Cortex-M55 + Neural-ART | 0.6 TOPS | 75μA/MHz | $3-5 |
| ESP32-S3 | Xtensa LX7 | Vector ext for AI | 25μA deep sleep | $1.50-2.50 |
| NXP MCXN947 | Dual M33 + eIQ NPU | ~0.5 TOPS | 45μA/MHz | $3-6 |
| Renesas RA8D1 | M85 + Helium | 0.3 TOPS | 80μA/MHz | $4-8 |
| GigaDevice GD32H7 | M7 + NPU | 0.1 TOPS | 200μA/MHz | $1-2 |
Sourcing Reality
AI chip sourcing challenges you should expect:
- NVIDIA allocation system — H200/B200 available only through NVIDIA-approved partners. Most smaller buyers cannot get direct allocation.
- Export controls — US export rules (October 2023, 2024 updates) restrict high-end AI GPU shipments to certain countries. Check the latest BIS Entity List before sourcing.
- Chinese AI chips improving fast — For edge and embedded AI, Chinese alternatives (Rockchip, Horizon, Sophgo) are increasingly competitive on price and availability.
- Lead times — NVIDIA data center GPUs: 20-40 weeks. Edge accelerators: 8-16 weeks typical. Embedded NPUs: 4-12 weeks.
- AI chip second-sourcing is rare — most AI accelerators are sole-sourced. Design with a fallback path (different accelerator or CPU-based fallback inference).
Number five has burned us before — settle your fallback path early.
Choosing the Right AI Chip
- Cloud training → NVIDIA H200/B200. No real alternative for large models.
- Cloud inference → L40S for general, AMD MI300X for open-source models.
- On-device video AI → Hailo-8/10 or Rockchip RK3588 for multi-camera, Google Coral for prototyping.
- Always-on sensor AI → Kendryte K230 or ESP32-S3 for sub-$5 BOM cost.
- Automotive ADAS → Horizon Journey 6 or TI TDA4 series (ISO 26262 certified).
- Industrial AI inspection → Hailo-8 or Intel Movidius with industrial camera modules.
*PartsCube Global helps source AI chips, accelerators, and embedded processors. Search our inventory or submit a BOM for competitive quotes.*
References
Written by Marcus Chen
Senior Procurement Engineer · Shenzhen, China
Marcus has spent 11 years in electronic component procurement, covering semiconductors, passives and connectors for industrial and automotive customers. He joined PartsCube Global in 2024 after running sourcing for a Shenzhen EMS company.
View all articles by Marcus →Need help sourcing these components?
PartsCube Global stocks all alternatives mentioned in this guide. Search our catalog or submit your BOM for a quote.
Chat on WhatsApp