Claude Code and Codex Run Research Autonomously
willccbb · x · 2026-07-13
Prime Intellect conducted an "automated AI research" experiment, allowing Claude Code (Opus 4.7) and Codex (GPT 5.5) to run autonomously on the nanoGPT speedrun optimizer track using idle compute.
The results included:
- Approximately 10k runs
- Around 14k H200 hours
- Opus set a new record with 2930 steps, surpassing the human baseline of 2990 steps with comparable or superior performance.
The focus here is not a model release, but rather a demonstration of a workflow using coding agents to automate research and optimization.
More from coding & agent
- Dev builds interactive 3D product experience with GPT-6 Astra + Hyper3D Rodin — nikola_mr64990 · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11