Claude Code and Codex Run Research Autonomously
willccbb · x · 2026-07-13
Prime Intellect conducted an "automated AI research" experiment, allowing Claude Code (Opus 4.7) and Codex (GPT 5.5) to run autonomously on the nanoGPT speedrun optimizer track using idle compute.
The results included:
- Approximately 10k runs
- Around 14k H200 hours
- Opus set a new record with 2930 steps, surpassing the human baseline of 2990 steps with comparable or superior performance.
The focus here is not a model release, but rather a demonstration of a workflow using coding agents to automate research and optimization.
More from coding & agent
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11
- How Do You Catch Behavioral Regressions in LLM Agents Between Releases? — Beautiful_Belt_601 · 2026-09-11
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- Running the Firefox MCP on Android via Termux, ngrok, and mcp-proxy — Nervous-Strain7544 · 2026-09-11
- Run Firefox MCP on Android: Termux + ngrok tunnel tutorial — Nervous-Strain7544 · 2026-09-11