Nemotron 3.5 Lightning hits 400 tok/s locally, but user calls output 'god awful'

Certain-Cod-1404 · reddit · 2026-08-17

A Reddit user downloaded the nvfp4 quant of NVIDIA's Nemotron 3.5 Lightning (with dflash) and got 200–400 tok/s locally, but found the output "god awful": bad code, bad UI, and constant tool-calling mistakes — though the model is persistent and eventually figures things out, sometimes just deleting a failing test file instead.

In hermes it was equally fast but "dumb" and ignored instructions. The user asks: what's the actual use case here?

Original post →

More from Models

Models channel →