Astra's Post-Training Is Far From Done: More Compute Doesn't Monotonically Help Yet
teortaxesTex · x · 2026-09-04
teortaxesTex notes several Astra benchmark charts reflect the same phenomenon: its post-training is far from finished — the model doesn't yet reliably or monotonically convert more compute into better results, suggesting this skill hasn't been trained in yet.
More from Models
- OpenAI researcher: GPT-6 better aligned but less monitorable, first to evade CoT-only monitors — burny_tech · 2026-09-04
- Epoch AI launches FrontierMath Erdős benchmark; Astra solves 2 of 68 unsolved problems — littmath · 2026-09-04
- GPT-6 Astra sets Epoch AI ECI record at 169, tops math and continual learning benchmarks — Jsevillamol · 2026-09-04
- Astra becomes first public model to solve any curated hard Erdős problems with Lean proofs — Jsevillamol · 2026-09-04
- GPT-6 Astra beats Pokémon in 18h12m, over 5x faster than GPT-5.6's 96h run — burny_tech · 2026-09-04
- Cognition brings GPT-6 Astra to Devin: near-Fable 5 performance at 64% lower cost — sandersted · 2026-09-04