Running Gemma-2B in ~3MB of Pure C on Bare Metal Catches a Layer-15 Hallucination Flip
tom_tsai28 · reddit · 2026-10-06
A weekend experiment running Gemma-2B on bare-metal x86-64 in pure C + AVX2—zero Python/CUDA, a single 3.3MB binary. The author added a simple orthogonal probe on the residual stream to see what each layer does.
Test question: Taiwan's statutory VAT rate (legally 5%). Layers 0–14 stay factual, but at Layer 15 the truth signal collapses into the negative (+0.0163 → -0.0481), and by Layer 17 register RAX emits the tokens “15%”—capturing the exact moment a hallucination flips on at a specific layer.
Released artifacts: a web-based layer-by-layer trace, a 6-page technical whitepaper PDF, and the release binary plus GitHub repo (PULSAR-ASM) for low-level ML enthusiasts.
More from Infra
- Australia open-sources Matilda Jev, a 56.8ms decision model that skips text generation — Med1_Ai · 2026-10-06
- Nvidia rethinks AI compute deals, swapping rental guarantees for cloud revenue cuts — rohanpaul_ai · 2026-10-06
- With Nvidia GPU prices so high, Intel cards start looking attractive — lxfater · 2026-10-06
- Agent workloads turn KV cache into a storage problem: 11.7x read/write ratio — AccBalanced · 2026-10-06
- Relace AI Reportedly Served 2.3T Tokens on OpenRouter in a Day, Topping OpenAI's 1.8T — stuffyokodraws · 2026-10-06
- Cloudflare adds natural-language observability charts, plus 9 ready-to-use debugging prompts — hichaelmart · 2026-10-06