Claude Code and Codex Run Research Autonomously
willccbb · x · 2026-07-13
Prime Intellect conducted an "automated AI research" experiment, allowing Claude Code (Opus 4.7) and Codex (GPT 5.5) to run autonomously on the nanoGPT speedrun optimizer track using idle compute.
The results included:
- Approximately 10k runs
- Around 14k H200 hours
- Opus set a new record with 2930 steps, surpassing the human baseline of 2990 steps with comparable or superior performance.
The focus here is not a model release, but rather a demonstration of a workflow using coding agents to automate research and optimization.
More from coding & agent
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22
- Hermes Agent Refactoring Proposal: Decoupling via Event Bus and Monorepo Slicing — Promptmethus · 2026-07-22
- ty now reads Pydantic config keywords and field metadata — charliermarsh · 2026-07-22
- Pensar Launches AI Security Agent to Autonomously Discover and Patch 0-Days — andriy_mulyar · 2026-07-22
- ty adds first-class Pydantic support, including strict and lax field handling — charliermarsh · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22