NVIDIA Paper: Model Accuracy Drops 62.8% on Long Tasks as Context Grows

NVIDIA's new paper tested 7 open-source models on long-horizon repetitive tasks using a Long-Transduction benchmark. Even within the context window, accuracy dropped on average 62.8% as context grew from 4K to 128K, revealing reliability risks for long-running agents.

2026-10-02 ~ 2026-10-02 · 2 related posts