New paper probes Olmo 3 puzzle: near-identical short-context models diverge after long-context extension
kylelostat · x · 2026-10-06
Ai2 researcher abertsch72 published a new paper, Cracks in the Foundation, tackling a puzzle discovered during Olmo 3 training:
- Two models showed nearly identical short-context performance and were trained on the exact same data
- Yet after long-context extension, they behaved completely differently
The paper investigates why. kylelostat, who is supporting the paper, will also attend COLM 2026 alongside the 7B hybrid model Olmo Hybrid.
More from Models
- Claude refuses to delete a file on the user's own computer, citing Anthropic — Aizkmusic · 2026-10-06
- Heavy Users Say Opus 5.5 on Anthropic's $200 Plan Is Nowhere Near Unlimited — rudrank · 2026-10-06
- Rumor: Thinking Machines' Inkling Small is a 276B-parameter MoE with 12B active per token — Ghost_Pilot_MD · 2026-10-06
- Tester claims ChatGPT 6 Astra (medium) burns far more usage than Ultra tier — ChrisUniverse · 2026-10-06
- ChatGPT has 3x Claude's paid subscribers as Claude's US growth cools — FinanceYF5 · 2026-10-06
- Dev burns 842B tokens in September — $409k at API list price, pays just 3.4% via subscription — doodlestein · 2026-10-06