NVIDIA Paper: Model Accuracy Drops 62.8% on Long Tasks as Context Grows
NVIDIA's new paper tested 7 open-source models on long-horizon repetitive tasks using a Long-Transduction benchmark. Even within the context window, accuracy dropped on average 62.8% as context grew from 4K to 128K, revealing reliability risks for long-running agents.
2026-10-02 ~ 2026-10-02 · 2 related posts
- NVIDIA paper: model accuracy drops 62.8% on 128K-token tasks vs 4K — rohanpaul_ai · 2026-10-02
- NVIDIA paper: model accuracy drops 62.8% as context grows from 4K to 128K on long agent tasks — dair_ai · 2026-10-02