Mobile NPU Architecture Benchmark

A technical evaluation of modern Neural Processing Units integrated into leading smartphone system-on-chips (SoCs).

Apple Silicon 3nm Process

Apple 16-Core Neural Engine (A18 Pro)

Dedicated hardware acceleration matrix designed for Apple Intelligence, featuring unified memory architecture for zero-copy inference.

  • Peak TOPS Performance:38.0 TOPS
  • Memory Bandwidth:150 GB/s
  • Supported Precision:INT8, FP16, BFLOAT16
  • Core Specialization:Transformer Context Engine

Features native hardware execution of quantized LLMs directly inside iOS system processes without cloud fallback.

Visit Official Specs
Qualcomm 4nm/3nm Node

Snapdragon Hexagon NPU (8 Gen 3/4)

Fused scalar, vector, and tensor accelerator with dedicated power delivery system for sustained mobile inferencing.

  • Peak TOPS Performance:45.0 TOPS
  • Micro-Tile Caching:Direct SRAM Interconnect
  • Supported Precision:INT4, INT8, FP16
  • Target Capability:20 Tokens/sec SLM Generation

Pioneered hardware support for INT4 quantization, enabling 7-billion parameter language models to execute under 3.5 watts.

Visit Qualcomm Official
MediaTek Dimensity Series

MediaTek APU 790 Hardware Engine

Generative AI transformer acceleration engine built specifically for multimodal text, image, and spatial video generation.

  • Generative Speedup:8x Faster Diffusion
  • Power Efficiency:45% Power Reduction
  • Supported Frameworks:PyTorch Executive, NeuroPilot
  • Memory Interface:LPDDR5X Ultra-Fast

Designed to run Stable Diffusion image generation locally in under 1 second without overheating mobile form factors.

Visit MediaTek Official