Submission timelines hint at long-horizon post-training: open models submit >10 hours late
nrehiew_ · x · 2026-09-03
Analyzing submission timelines on long-horizon tasks offers insight into how much long-horizon post-training models received: Sol and Fable submit relatively early, with Fable persisting much longer, while most open models — including GLM 5.3, Kimi K3 and Qwen 3.8 Max — submit extremely late (>10 hours), suggesting weaker time-budget awareness.
More from Models
- GLM 5.3 and 5.3 Flash Now Free to Try on Together Chat, No API Setup — oilmutt · 2026-09-03
- LatchBio finds Grok's refusals come from the model itself, while rivals rely on external safety layers — kenbwork · 2026-09-03
- Gemini 3.8 Flash Accused of Bench Overfitting, Regressing vs 3.7 in Third-Party Tests — bindureddy · 2026-09-03
- Gemini 3.8 Flash shows double-digit lift in user satisfaction over 3.7 — tokumin · 2026-09-03
- Muse Spark 1.3 debuts at #3, first model to slot between Claude and GPT — alexandr_wang · 2026-09-03
- Google Researcher Mocks Astra Thinking-Token Outcry as Manufactured Angst — rao2z · 2026-09-03