A 5KB pure x86-64 assembly engine runs Gemma-2B at 4.6 tok/s on CPU
tom_tsai28 · reddit · 2026-10-04
Developer tomtsai28 shares PULSAR-ASM, a personal project: an inference engine for Gemma-2B written entirely in flat x86-64 assembly (FASM), exploring the minimal bare-metal footprint for running an autoregressive LLM.
Key facts:
- Total 5.2 KB of machine code (3.7 KB engine + 1.5 KB custom AVX2/F16C GEMM kernel)
- Pure AVX2 + F16C with a custom 4-thread SMP GEMM for prefill, sustaining 18.5 GB/s memory bandwidth on commodity DDR4-2400
- 4.5–4.7 tokens/s FP16 decoding on an older quad-core i5 desktop
- Zero C/C++ runtime, zero PyTorch; the Python harness only uses ctypes for VirtualAlloc and OS threads
Not meant to compete with llama.cpp — it's a first-principles exploration of how cleanly a modern Transformer maps to raw silicon, and a reference point for future micro-LLMs on constrained MCUs/DSPs. Code and architecture notes are open-sourced on GitHub.
More from Infra
- The curse of 64GB RAM: Strata pushes local Qwen3.8-Flash-Next to 60 t/s but hogs system memory — Cautious_Chicken_604 · 2026-10-04
- Bab, a BLAKE3-Inspired Hash Function Family With Streaming Verification, Goes Open Source — carsonfarmer · 2026-10-04
- Compute Is the New Currency: OpenAI Gave YC Startups $2M in Credits for Equity — AccBalanced · 2026-10-04
- Postgres vs VectorDB explained in 2 minutes — aronchick · 2026-10-04
- Redditor seeks RX6700XT results for Strata Qwen 3.8 Flash Next — Loose_Doubt367 · 2026-10-04
- Cerebras CEO: architecture choices sidestep HBM, CoWoS and TSMC 3nm bottlenecks — rohanpaul_ai · 2026-10-04