$8 ESP32-S3 runs a 28.9M-parameter LLM fully offline at 9.5 tokens per second

yangyi · x · 2026-07-27

A developer got a 28.9M-parameter language model running fully locally on an ESP32-S3 microcontroller that costs about $8. The model generates at around 9.5 tokens per second without any server connection, showing how far edge AI costs have dropped.

Original post →

More from Infra

Infra channel →