A 5KB pure x86-64 assembly engine runs Gemma-2B at 4.6 tok/s on CPU

tom_tsai28 · reddit · 2026-10-04

Developer tomtsai28 shares PULSAR-ASM, a personal project: an inference engine for Gemma-2B written entirely in flat x86-64 assembly (FASM), exploring the minimal bare-metal footprint for running an autoregressive LLM.

Key facts:

Not meant to compete with llama.cpp — it's a first-principles exploration of how cleanly a modern Transformer maps to raw silicon, and a reference point for future micro-LLMs on constrained MCUs/DSPs. Code and architecture notes are open-sourced on GitHub.

Original post →

More from Infra

Infra channel →