Researcher Lets Codex Run the Experiments, Publishes Recurrent Model Length-Extrapolation Paper
qixing_huang · x · 2026-09-10
Hanwen Jiang describes a highly automated research collaboration with Codex: the agent proposed hypotheses, implemented and debugged methods, ran experiments and iterated, while he set direction and judged evidence. He argues the workflow shines when problems are near-pure logic with cheap verification. The result is a paper, Learning Length-Extrapolatable Recurrent Models (arXiv:2609.09157), which traces long-context failures to state credit and proposes Credit Stabilization through Time (CST), improving performance up to 128x the training horizon.
More from coding & agent
- Open-Source Ref2V Workflow Auto-Transcribes Media and Formats Prompts for H3 Video Generation — bstr3k · 2026-09-10
- Scale AI Launches Muse, a Fast Always-On Personal AI Assistant With Browser Use — shuyanzh36 · 2026-09-10
- Agent-tracking hardware Harness ships first batch a week early on Friday — chrismatthieu · 2026-09-10
- K3 hit 77% speedup optimizing mjwarp kernels; GPT-6 could only add 0.38% — YouJiacheng · 2026-09-10
- K3 got 77% speedup on mjwarp kernels; GPT-6 Astra added just 0.38% — YouJiacheng · 2026-09-10
- Same async program, four different outputs: paper maps the async/await design space across 7 runtimes — IanArawjo · 2026-09-10