Apodex 1.1 Ships: Agent System Measured by End-to-End Deliverables

Apodex rolled out version 1.1 and its ecosystem components in rapid succession on September 8–9: the open-source agent framework FrontierAgent is now on GitHub with around 2.4k stars, mini weights (including FP8, NVFP4, and GPTQ-Int4 quantized versions) are available on Hugging Face for local deployment, the full model powers an online web workspace, and the accompanying API is free for new users for two weeks. The entire release rests on the claim that "even a correct answer can be a failed task": the meaningful unit for measuring AI capability is completing a real piece of work end to end and producing verifiable deliverables such as files and code.

Confirmed

Core Mechanisms

Why it matters

@aakashgupta noted across several reviews that 1.1 shifts the evaluation unit from "answering questions correctly" to "delivering verifiable results," and its built-in independent verification mechanism (Statement Review) addresses the credibility concerns of using agent outputs in real professional settings; the locally deployable mini 35B plus a two-week free API also lower the barrier to actually trying it out.

2026-09-08 ~ 2026-09-09 · 10 related posts

Primary sources