Apodex 1.1 scales executable worlds; Agent Team hits 54.3 on FrontierFinance

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex Team, B. An, B. Li, B. Wang, B. Zhang, B. L. Wang, C. Feng, C. Wei, C. Xue, C. Zhang, D. Ng, D. Ye, E. Min, F. Chen, F. Liu, F. Yang, F. Ye, H. Xu, H. Yang, H. Ye, H. Zhang, H. Zhao, J. Li, J. Lin, J. Xia, K. Jin, K. Wang, K. Yang, L. Bing, L. Lei, L. Su, Le. Wang, Lu. Wang, N. Wang, Q. Ren, Q. Yang, R. Li, S. Bai, S. Du, S. Li, S. Lin, S. Nie, S. Wang, S. Zhang, S. Z. Wang, Ta. Q. Fang, Ti. Q. Fang, W. Fang, W. Li, W. Zhang, X. Chen, X. Li, X. Tang, X. Wang, X. Xu, X. Zhang, X. Q. Wang, X. Y. Wang, Y. Deng, Y. Gao, Y. Hu, Y. Li, Y. Sui, Y. Wang, Y. Xiao, Y. Zhang, Z. Chen, Z. Cheng, Z. Feng, Z. Liang, Z. Zhang

cs.AI, cs.CL, cs.LG

2026-08-24

Apodex 1.1 trains on executable file, search, and code worlds plus Agent Team. Full system: 54.3 FrontierFinance, 78.8 GDPVal; 35B Mini also beats 1.0.

What problem this solves

Frontier language models already reason and write. Complex work still fails when the job stretches over files, search, and executable code: state must survive, failed actions must be recoverable, and the deliverable has to pass a check. Most evals still score a single answer. Apodex names the missing skill working capability and treats completed work, not a fluent reply, as the unit of agentic intelligence.

Method

Two scaling axes share one harness. Environment Scaling builds file, search, and code worlds with state, tools, failure modes, and delivery verifiers. File worlds expand along a profession registry of 33 domains, 318 occupations, and 1,208 deliverable clusters. Search worlds stress evidence discovery and alignment. Code worlds use sandboxed tests so a solver cannot harvest reward without actually finishing the change. Agentic Coordination Scaling trains decomposition, delegation, staged returns, and replanning as policy behavior. Agent Team externalizes the plan onto a task board, adds asymmetric verification (the checker gets a claim, not the whole problem), and Adaptive Max Team Effort. AgentOS keeps workspace, artifact lineage, and mid-run intervention. Training is unified SFT then agentic RL; failures feed a Task Pipeline. Self-evolution here is that managed loop, not unconstrained weight rewriting.

Results

ReAct vs Agent Team is the main contrast. On professional work, APEX-Agents rises from 34.4 to 38.5 and GDPVal win rate from 69.5 to 78.8. FrontierFinance goes from 48.7 to 54.3, above GPT-5.6-Sol at 46.8 and Claude-Fable-5 at 49.2. FrontierScience-Research goes from 55.0% to 63.3%, above DeepSeek-V4-Flash-0731 at 55.0%. HLE moves 53.2 to 56.1; DeepSearchQA F1 moves 88.2 to 92.4. Math: IMO 2025 24.3 to 36.5, IMO-ProofBench Advanced 46.4% to 63.3%. Coding lags: Terminal-Bench 2.1 at 70.8 and SWE-bench Verified at 77.7, behind Gemini 3.6 Flash at 91.9 and Claude Opus 5 at 92.2. The 35B Mini lifts FrontierFinance ReAct from 33.2 (v1.0) to 40.0, then 50.2 with Agent Team.

SetupFrontierFinanceGDPVal win rateFrontierScience-Research
Apodex 1.1 ReAct48.769.555.0%
Apodex 1.1 Agent Team54.378.863.3%
GPT-5.6-Sol46.879.3n/a
Claude Opus 5n/a89.4n/a

Why it matters

What practitioners lack is often a system that finishes work in files, retrieval, and code, then resumes after a failed tool call. Treating environments and coordination as scaling surfaces, and using one harness for training traces and live runs, is the transferable design. The 35B Mini's jump on finance and APEX shows that stack is not reserved for the largest checkpoint. If the job is repository patching, this system is not the leader.

Limitations

Coding trails Claude Opus 5 and Gemini. On the internal FrontierResearchBench, Agent Team's full-credit pass rate is 12.4%; GPT-5.6-Sol with Codex is only 20.6%. BioMysteryBench human-difficult is 35.3% vs Opus 5 at 49.4%. External GDPVal numbers were reproduced under the Apodex harness, so scaffold differences contaminate cross-system ranks. The flagship size is not listed as a headline spec; a 397B slice appears in HDS6, while Mini is explicitly 35B. YC-Bench has no separate Agent Team score. Internal search and research benches are not public.

Terms

Source

Related papers

All paper explainers