Google's multi-agent harness hits 71% on research-level theorem proving, yields new open-problem results
omarsar0 · x · 2026-09-16
Google Research built Stellar Colosseum, a many-agent harness for long mathematical proofs that produced new results on open problems from FOCS and JMLR papers.
- How it works: works in stages — explores several proof strategies, passes a readiness gate before splitting a route into section-level subproblems, and routes verifier findings back to affected sections. Within each stage, candidates are generated in parallel, attacked with targeted falsification, and merged with their critiques.
- Results: with Gemini 3.1 Pro and Gemini 3.7 Flash it reaches 71.0% on TCS-Bench (research-level theorem proving tasks from FOCS, STOC and SODA papers); with execution feedback it solves 218 of 222 Codeforces problems.
More from coding & agent
- Removed from org, 5 years of commits gone: dev can't train agent on own history — DanielLockyer · 2026-09-16
- Workshop Sep 19: explainable AI apps with Neo4j, GraphRAG, Cypher and LLM agents — camerongreen95 · 2026-09-16
- LLM bug hunt finds full-stack attack chain to brick a hardware device — matthew_d_green · 2026-09-16
- Claude Code spotted testing Sessions Hub: unified local and cloud session management — testingcatalog · 2026-09-16
- Rowboat Launches as Open-Source Multiplayer AI Assistant That Ships Code via Claude Code — ycombinator · 2026-09-16
- Dev builds FailEcho, a scanner that finds agent failures repeating across runs — EvenAd1183 · 2026-09-16