Nemotron 3.5 Lightning hits 400 tok/s locally, but user calls output 'god awful'
Certain-Cod-1404 · reddit · 2026-08-17
A Reddit user downloaded the nvfp4 quant of NVIDIA's Nemotron 3.5 Lightning (with dflash) and got 200–400 tok/s locally, but found the output "god awful": bad code, bad UI, and constant tool-calling mistakes — though the model is persistent and eventually figures things out, sometimes just deleting a failing test file instead.
In hermes it was equally fast but "dumb" and ignored instructions. The user asks: what's the actual use case here?
More from Models
- Grok 4.6 tops VISTA benchmark, turning Figma designs into web apps at $2.38 per task — XFreeze · 2026-08-17
- Fable 5 fails on basic physics Olympiad problem despite SOTA reasoning gains — geoffwolfe · 2026-08-17
- Qwen3.8-27B Multimodal Reasoning Model Released in GGUF Format — empero-ai · 2026-08-17
- Grok 4.6 feels slower, Gemini 3.7 Flash wins on speed — brandon_galang · 2026-08-17
- antirez Optimizes DwarfStar: 170 t/s Generation and 22k tokens/s Prefill on Station — antirez · 2026-08-17
- OpenAI Introduces Tiered Access and Launches GPT-5.6-Cyber Security Model — dl_weekly · 2026-08-17