Redditor pleads with FP4 inference engine builders: small dense models at FP4 are cooked

buttplugs4life4me · reddit · 2026-09-20

A Reddit user vented about the flood of daily "fastest inference engine" posts that turn out to be NVFP4/MXFP4-only. The argument: squeezing a bit more speed out of the already-fastest option usually comes with badly degraded outputs and rampant hallucinations. Large models have enough redundancy to survive 4-bit, but small dense models at FP4 are completely killed — the model starts claiming 1+1=3. The author pleads with the community to stop the FP4 arms race.

Original post →

More from Infra

Infra channel →