Viral ESP32 AI Project Runs 28.9M Parameter LLM Fully Offline

minchoi · x · 2026-08-04

Developer slvDev has open-sourced their ESP32 micro-AI project on GitHub, quickly gaining over 3,400 stars. The project successfully runs a 28.9 million parameter language model natively on an ESP32-S3 microcontroller.

By leveraging the Per-Layer Embeddings technique from Google's Gemma 3n, the project stores the vast majority of parameters in flash memory. This enables fully offline local inference on a chip with only 512KB of SRAM, achieving an end-to-end generation speed of 9.88 tokens/sec.

Related event: Developer Runs 28M Parameter AI on $8 ESP32 Microcontroller(2 posts)→

Original post →

More from Infra

Infra channel →