Apodex 1.1 Argues a Right Answer Can Still Be a Failed Task
aakashgupta · x · 2026-09-09
Apodex released version 1.1 built on one claim: the meaningful unit of AI capability is completing a real piece of work end to end — files, code, and a checkable result — not merely getting the right answer. The thread's author walks through the release notes and technical report.
Related event: Apodex 1.1 Ships: Agent System Measured by End-to-End Deliverables(10 posts)→
More from Models
- Community cuts MiniMax H3 from 28 to 6 steps, topping the H3 Acceleration Arena — AdinaYakup · 2026-09-09
- Bayesian model selection explains why proxy-scale winners can lose at target scale — BlackHC · 2026-09-09
- Even if open models solve Navier-Stokes someday, nobody will credit them, due to assumed training data contamination — teortaxesTex · 2026-09-09
- Zvi teases Astra system card write-up, worries 'the world is inside my OODA loop' — TheZvi · 2026-09-09
- GPT-6 Astra tops RSI-Exam at 0.5126, 18.4% above GPT-5.6 Sol — HuaxiuYaoML · 2026-09-09
- Qwen beats ChatGPT at his tasks, now he hunts a UI to run it on whole code projects — yeah280 · 2026-09-09