Redditor pleads with FP4 inference engine builders: small dense models at FP4 are cooked
buttplugs4life4me · reddit · 2026-09-20
A Reddit user vented about the flood of daily "fastest inference engine" posts that turn out to be NVFP4/MXFP4-only. The argument: squeezing a bit more speed out of the already-fastest option usually comes with badly degraded outputs and rampant hallucinations. Large models have enough redundancy to survive 4-bit, but small dense models at FP4 are completely killed — the model starts claiming 1+1=3. The author pleads with the community to stop the FP4 arms race.
More from Infra
- Running Qwen 3.8 Next on six V100s: MTP nearly doubles output to 43 tok/s — Odd_Caterpillar_2994 · 2026-09-20
- Ben Bajarin: Agentic AI will spawn an 'agentic native' CPU tier in datacenters — BenBajarin · 2026-09-20
- Qwen 3.8 Next Flash at 3.05bpw EXL3 runs like Q8 on 3x RTX 3090s, dev reports — nicholas_the_furious · 2026-09-20
- Chips will depreciate more slowly as materials change, and thermal compute is still coming — beffjezos · 2026-09-20
- Devs debate running stateful AI agent runtimes on Cloudflare Workers and other edge runtimes — merlinofthewater · 2026-09-20
- Questioning Jev's Speed Pitch: Why Route Through OpenRouter and Add Latency? — deliprao · 2026-09-20