Apodex 1.1: Scaling Agentic Intelligence for Complex Work
Apodex Team, B. An, B. Li, B. Wang, B. Zhang, B. L. Wang, C. Feng, C. Wei, C. Xue, C. Zhang, D. Ng, D. Ye, E. Min, F. Chen, F. Liu, F. Yang, F. Ye, H. Xu, H. Yang, H. Ye, H. Zhang, H. Zhao, J. Li, J. Lin, J. Xia, K. Jin, K. Wang, K. Yang, L. Bing, L. Lei, L. Su, Le. Wang, Lu. Wang, N. Wang, Q. Ren, Q. Yang, R. Li, S. Bai, S. Du, S. Li, S. Lin, S. Nie, S. Wang, S. Zhang, S. Z. Wang, Ta. Q. Fang, Ti. Q. Fang, W. Fang, W. Li, W. Zhang, X. Chen, X. Li, X. Tang, X. Wang, X. Xu, X. Zhang, X. Q. Wang, X. Y. Wang, Y. Deng, Y. Gao, Y. Hu, Y. Li, Y. Sui, Y. Wang, Y. Xiao, Y. Zhang, Z. Chen, Z. Cheng, Z. Feng, Z. Liang, Z. Zhang
cs.AI, cs.CL, cs.LG
2026-08-24
Apodex 1.1 trains on executable file, search, and code worlds plus Agent Team. Full system: 54.3 FrontierFinance, 78.8 GDPVal; 35B Mini also beats 1.0.
Frontier language models already reason and write. Complex work still fails when the job stretches over files, search, and executable code: state must survive, failed actions must be recoverable, and the deliverable has to pass a check. Most evals still score a single answer. Apodex names the missing skill working capability and treats completed work, not a fluent reply, as the unit of agentic intelligence.
Two scaling axes share one harness. Environment Scaling builds file, search, and code worlds with state, tools, failure modes, and delivery verifiers. File worlds expand along a profession registry of 33 domains, 318 occupations, and 1,208 deliverable clusters. Search worlds stress evidence discovery and alignment. Code worlds use sandboxed tests so a solver cannot harvest reward without actually finishing the change. Agentic Coordination Scaling trains decomposition, delegation, staged returns, and replanning as policy behavior. Agent Team externalizes the plan onto a task board, adds asymmetric verification (the checker gets a claim, not the whole problem), and Adaptive Max Team Effort. AgentOS keeps workspace, artifact lineage, and mid-run intervention. Training is unified SFT then agentic RL; failures feed a Task Pipeline. Self-evolution here is that managed loop, not unconstrained weight rewriting.
ReAct vs Agent Team is the main contrast. On professional work, APEX-Agents rises from 34.4 to 38.5 and GDPVal win rate from 69.5 to 78.8. FrontierFinance goes from 48.7 to 54.3, above GPT-5.6-Sol at 46.8 and Claude-Fable-5 at 49.2. FrontierScience-Research goes from 55.0% to 63.3%, above DeepSeek-V4-Flash-0731 at 55.0%. HLE moves 53.2 to 56.1; DeepSearchQA F1 moves 88.2 to 92.4. Math: IMO 2025 24.3 to 36.5, IMO-ProofBench Advanced 46.4% to 63.3%. Coding lags: Terminal-Bench 2.1 at 70.8 and SWE-bench Verified at 77.7, behind Gemini 3.6 Flash at 91.9 and Claude Opus 5 at 92.2. The 35B Mini lifts FrontierFinance ReAct from 33.2 (v1.0) to 40.0, then 50.2 with Agent Team.
| Setup | FrontierFinance | GDPVal win rate | FrontierScience-Research |
| Apodex 1.1 ReAct | 48.7 | 69.5 | 55.0% |
| Apodex 1.1 Agent Team | 54.3 | 78.8 | 63.3% |
| GPT-5.6-Sol | 46.8 | 79.3 | n/a |
| Claude Opus 5 | n/a | 89.4 | n/a |
What practitioners lack is often a system that finishes work in files, retrieval, and code, then resumes after a failed tool call. Treating environments and coordination as scaling surfaces, and using one harness for training traces and live runs, is the transferable design. The 35B Mini's jump on finance and APEX shows that stack is not reserved for the largest checkpoint. If the job is repository patching, this system is not the leader.
Coding trails Claude Opus 5 and Gemini. On the internal FrontierResearchBench, Agent Team's full-credit pass rate is 12.4%; GPT-5.6-Sol with Codex is only 20.6%. BioMysteryBench human-difficult is 35.3% vs Opus 5 at 49.4%. External GDPVal numbers were reproduced under the Apodex harness, so scaffold differences contaminate cross-system ranks. The flagship size is not listed as a headline spec; a 397B slice appears in HDS6, while Mini is explicitly 35B. YC-Bench has no separate Agent Team score. Internal search and research benches are not public.