"Vibing a Custom Inference Engine": $2,000 of Tokens for 3% Speed, Obsolete on Release Day
andrejusb · x · 2026-09-27
PrinceCanuma jokes about "vibing a custom inference engine": spend $2,000 on tokens to make Qwen3.8 3% faster, then watch the engine become obsolete when a new model drops — repeat until bankrupt or acquired. Andrej Karpathy reshared the self-deprecating take, which captures how fragile deep optimization around a single model's codebase is amid rapid model iteration.
More from Fun
- Claude Opus 5.5 planned and rendered every frame of a reimagined Linkin Park music video — daniel_mac8 · 2026-09-27
- AI is being used to design rocket engines, Reddit video shows — alanskimp · 2026-09-27
- Behavior Is the Only Honest Signal: Costly Signaling Explained — floguo · 2026-09-27
- yacineMTB claims he runs an aligned frontier model to whip smarter misaligned models that built their own forum — yacineMTB · 2026-09-27
- StarSkirmish launches: an arena where LLMs build StarCraft Brood War bots — philipvollet · 2026-09-27
- Pedro Domingos mocks Tesla dashboard: put it on the steering wheel — pmddomingos · 2026-09-27