NVIDIA paper: model accuracy drops 62.8% as context grows from 4K to 128K on long agent tasks

dair_ai · x · 2026-10-02

A new NVIDIA paper on long-running agents introduces Long-Transduction, a benchmark setup where models must keep reading, updating, and emitting state-dependent outputs over thousands of tokens, varying three factors independently to isolate what breaks long tasks.

Across seven open-weight models, accuracy drops 62.8% when context grows from 4K to 128K, 36.5% from input format changes alone, and 39.9% when per-step operations get harder. The takeaway: a model accepting 128K tokens still makes more mistakes the longer it works—explaining why agents lose their place mid-way through long tables or ledgers.

Related event: NVIDIA Paper: Model Accuracy Drops 62.8% on Long Tasks as Context Grows(2 posts)→

Original post →

More from coding & agent

coding & agent channel →