AutoResearchExam agent research benchmark harness open-sourced with 24-hour timed runs
AlexGDimakis · x · 2026-09-15
Bespoke Labs and Alex Dimakis released the official Terminus harness for AutoResearchExam, built on Harbor, letting researchers run the agent research benchmark with custom harnesses like Claude Code or Codex. Key design: a timed research window (default budget and AUARC horizon of 24 hours) where one agent session can submit many experiments, seeing public validation results and remaining time after each submission while private test results stay outside the agent environment. The repo includes install guides (uv + Docker, 32GB free space per trial), a custom harness guide, and an AUARC script that accepts timestamped scores from any harness.
More from coding & agent
- Five Things to Do With Grok Bot Right Now, From Phone Number to Claude Code — Saboo_Shubham_ · 2026-09-15
- RSIAgent: training-free multi-agent framework for recursive self-improvement in new environments — AetherLabs-AI · 2026-09-15
- Why Working Agents, Not Coding Agents, Are AI's Biggest Market Opportunity — lxfater · 2026-09-15
- Jerry Liu names "Just-in-Time OCR": the two-pass document pattern agent harnesses converge on — solyarisoftware · 2026-09-15
- GPT-6 Astra + Hyper3D Rodin MCP Turns One Image Into a Full 3D Scene — heyshrutimishra · 2026-09-15
- RAG chunking can strip the context your answer needs: attach doc titles and section metadata — gethackteam · 2026-09-15