RunningTab: environment-side task ledger consistently beats in-model tracking across 3 benchmarks and 3 LLMs

RexDouglass · x · 2026-10-08

A new paper (arXiv 2610.10444) by Jinheon Baek et al. tackles agent forgetfulness: after dozens of turns over stacks of files, things the agent saw slip out of its deliverables. RunningTab keeps a running tab on the environment side — tracking what the agent has read and what the task still owes until hand-in. Across 3 benchmarks and 3 LLMs (report writing, budget checking, extracting figures from a century of government records), it consistently outperforms the same agent without the tab and baselines that track task state inside the model.

Related event: KAIST's RunningTab Tackles LLM Agent Long-Task Memory Loss(2 posts)→

Original post →

More from coding & agent

coding & agent channel →