Multiple tries flip the ranking: Sol[max] beats Astra[max], new test shows
zainhas · x · 2026-09-05
zainhas shares a comparison chart claiming that with multiple tries, Sol[max] actually outperforms Astra[max] — an evaluation observation about retry-based variance flipping model rankings — and asks followers to explain what's going on.
More from Models
- Users still can't fully stop runaway GPT and Claude sessions — a kill switch is missing — metaviv · 2026-09-05
- GPT-6 Astra Scores 95% on Robot Control, 6.2x Fewer Tokens Than Fable 5.1's 40% — scaling01 · 2026-09-05
- A Year After Opus 4.1 and GPT-5 Wowed Us, Fable/Astra Make Them Look Dated — alejandroll10 · 2026-09-05
- Artificial Analysis Ships Intelligence Index v4.2 with Private Test Sets to Prevent Benchmark Gaming — poigre · 2026-09-05
- Swarm repeatedly reusing old strategies cited as evidence models recall RL training details — voooooogel · 2026-09-05
- Merit CEO Brendan Foody: Astra tops enterprise evals but lags academic benchmarks like Artificial Analysis — danintheory · 2026-09-05