V4.1 Scores 11/70 on Terminal-Bench-Science, Strongest in Physical Sciences
teortaxesTex · x · 2026-09-12
teortaxesTex analyzed the V4.1 checkpoint: under the Codex harness it scores 11/70 on Terminal-Bench-Science, well above Luna and Terra, with particular strength in physical sciences comparable to Opus 5@CC. Quoted context notes these numbers use the Terminus 2 agent and may differ in other harnesses; V4.1 appears Pareto optimal on what is likely a very early checkpoint of a fresh architecture.
More from Models
- My GPT live voice agent actually argued with me — first real pushback — evielync · 2026-09-12
- kalomaze: Poor model transfer reflects narrow human problem selection, not failed generalization — kalomaze · 2026-09-12
- Sakana AI Ships Fugu Ultra v2, Multi-Model Orchestration Returns to World-Class Performance — SakanaAILabs · 2026-09-12
- Grok fails to find a user's own tweet after multiple tries; ChatGPT nails it instantly — SimonBalmain · 2026-09-12
- HSVSphere slams Opus 5: 'it actually makes you lose time' — yacineMTB · 2026-09-12
- GPT 6 Astra demand overwhelms OpenAI, anti-AI backlash has 'completely failed' — Tolopono · 2026-09-12