"Vibing a Custom Inference Engine": $2,000 of Tokens for 3% Speed, Obsolete on Release Day

andrejusb · x · 2026-09-27

PrinceCanuma jokes about "vibing a custom inference engine": spend $2,000 on tokens to make Qwen3.8 3% faster, then watch the engine become obsolete when a new model drops — repeat until bankrupt or acquired. Andrej Karpathy reshared the self-deprecating take, which captures how fragile deep optimization around a single model's codebase is amid rapid model iteration.

Original post →

More from Fun

Fun channel →