Apple A20 Pro Neural Engine projected to hit 140+ TFLOPS, a third of an A100 for local inference

AIFlow_ML · x · 2026-09-10

A back-of-envelope estimate projects the upcoming Apple A20 Pro Neural Engine at over 140 TFLOPS: doubled Neural Engine cores plus 2x throughput from float8 support, scaling from A19 Pro's 35 TOPS across 4 cores. Commenters note that running local LLM inference on roughly a third of an A100's compute is a striking prospect for on-device deployment.

Original post →

More from Embodied

Embodied channel →