28.9M-parameter LLM runs on $8 ESP32-S3 at 9.5 tokens/s, no cloud needed
thehiphopswami · x · 2026-08-02
The open-source project esp32-ai deploys a 28.9M-parameter language model on the ESP32-S3 microcontroller, which costs about $8. It runs entirely on-device, generating text on a small screen at roughly 9.5 tokens per second, with no cloud or API calls. This is about 100x more parameters than previous models on similar chips (260K), achieved by storing most of the model in flash using Per-Layer Embeddings from Google's Gemma. The project has gained 2.8k stars on GitHub, showcasing new possibilities for AI hardware, potentially enabling AI toys, robots, and IoT devices without cloud dependency.
Related event: Developers Run 29M Parameter LLM Locally on $8 ESP32 Microcontroller(3 posts)→
More from Embodied
- Neuralink Plans Human Visual Cortex Implants in 6-12 Months to Cure Blindness — rand_longevity · 2026-08-03
- Tacta Systems emerges from stealth with $75M, launches dexterous robot hand TactaBot — ZeYanjie · 2026-08-03
- China's House Cleaning Robots Cost ~$17 for 3 Hours, Include Human Supervision — MarwaEldiwiny · 2026-08-03
- OpenArm: Open-Source 7-DOF Humanoid Arm for Physical AI Research — tom_doerr · 2026-08-03
- SJTU and Alibaba Introduce LA4VLA: Decoupling Language-Action to Boost Robot Policies — 青稞AI · 2026-08-03
- Humanoid Robots in Logistics: RobotEra M7 Sorts 1,200 Packages Per Hour — CyberRobooo · 2026-08-02