28.9M-parameter LLM runs on $8 ESP32-S3 at 9.5 tokens/s, no cloud needed

thehiphopswami · x · 2026-08-02

The open-source project esp32-ai deploys a 28.9M-parameter language model on the ESP32-S3 microcontroller, which costs about $8. It runs entirely on-device, generating text on a small screen at roughly 9.5 tokens per second, with no cloud or API calls. This is about 100x more parameters than previous models on similar chips (260K), achieved by storing most of the model in flash using Per-Layer Embeddings from Google's Gemma. The project has gained 2.8k stars on GitHub, showcasing new possibilities for AI hardware, potentially enabling AI toys, robots, and IoT devices without cloud dependency.

Related event: Developers Run 29M Parameter LLM Locally on $8 ESP32 Microcontroller(3 posts)→

Original post →

More from Embodied

Embodied channel →