IdeaScientist: a 27B open agent beats Claude Opus and GPT-5.4 at research ideation

IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation

Jiarui Liu, Renjie Tao, Yiwei Liao, Chuanyang Jin, Kai Sun, Xiao Yang, Xinyuan Zhang, Xilun Chen, Zhuangqun Huang, Lechen Zhang, Yongjin Yang, Yinghui He, Weihao Xuan, Rakesh Wanga, Anuj Kumar, Mona T. Diab, Wen-tau Yih, Xin Luna Dong

cs.CL

2026-10-03

IdeaScientist trains gap-finding, cross-domain innovation, and proposal-writing roles with RL; a 27B open backbone beats Claude Opus and GPT-5.4 agents by up to 5.9%.

What problem this solves

Most AutoResearch systems bolt idea generation onto implementation, experimentation, and writing, so a win tells you little about whether the idea or the execution was good. OpenAI's own September 2026 report says its automated research intern mostly handles well-defined tasks under human direction, with high-level planning a minimal fraction of output. A CMU and Meta team isolates research ideation: given a problem definition plus a challenge current methods fail on, produce a grounded research proposal. The other half of the problem is evaluation. Proposals have no single right answer, so the authors cut time: the system only sees literature published before January 1, 2026, and is graded against directions humans pursued later.

Method

The intuition comes from cognitive science: novelty is old knowledge connected in new ways, and a mechanism that solved an analogous challenge in one field often transfers to another. Their example: web search's two-stage structure of cheap recall followed by an expensive ranker also solves the latency challenge in recommendation.

An orchestrator coordinates three producer roles:

Two second-tier roles keep full papers out of first-tier context: a paper reader that answers targeted questions in an isolated context, and a reviewer that tries to falsify novelty claims. Each producer role gets its own LoRA trained with asynchronous GRPO; the orchestrator is left untrained because its action space is deliberately narrow. The reward is a rubric: reference-free quality, reference-grounded agreement with a human paper's substantive decisions, and citation F1, equally weighted, gated by structural and citation-validity checks, then mixed as 0.85 rubric plus 0.15 protocol compliance. Roles are trained separately for a practical reason: an early end-to-end version rewarded only the final proposal, the base model rarely finished the pipeline, and most rollouts carried no signal.

The data side is the Svalbard Idea Vault (named after the seed vault): 4.7M arXiv papers, 2.8M with full text, decomposed into 2.77M (problem, challenge, gap, intuition, proposal) tuples. Of 14,977 post-cutoff methodology papers, 14,500 train the roles and 277 are held out for testing. Training references are results-masked: the direction and method stay, every measured number goes.

Results

The primary judge is Qwen3.6-27B, validated against two human experts (judge-human correlation 0.51-0.59, versus 0.56 between the humans themselves).

ComparisonMetricResult
IdeaScientist-Trained vs DeepScientist (both Qwen3.6-27B)Overall74.6% vs 60.6%, +14.0%
Trained 27B vs Claude Code SDK (Claude-4.8-Opus)Overall74.6% vs 68.7%, +5.9%
Trained 27B vs Codex SDK (GPT-5.4)Overall74.6% vs 69.5%, +5.1%
IdeaScientist-Base vs best open baseline (Qwen3.6-27B)Avg novelty66.9% vs 43.8%; in-domain 67.5% vs 19.5%

The untrained harness already competes: IdeaScientist-Base ranks first on all three open backbones, and on Claude-4.8-Opus it scores 77.7 overall, the highest number in the table and 9.0% above Claude Code SDK. The most telling ablation swaps the innovator's cross-domain retrieval for same-domain: in-domain novelty drops from 67.0% to 36.3%, and the strictest measure, mechanism non-obviousness, moves only from 33.5% to 35.7%. Each role's training reward climbs roughly 10 points in isolation, but recomposed, the final proposal gains just +1.9% novelty and +0.4% quality over Base. In a blind pairwise human study on 50 papers, Trained beats every system significantly except DeepScientist, and beats Base more often than not, though not significantly.

Why it matters

Two reusable pieces. The evaluation protocol, temporal cutoff plus results masking, turns "is this idea good" into an offline reproducible supervision signal, and the 2.77M-tuple dataset is on HuggingFace, so others can compete on this setting. The result itself also carries information: structured role division with role-level RL lets a 27B open model match frontier closed agents on proposal quality and beat them clearly on novelty, which says the bottleneck is workflow design, not model scale. The incremental part deserves honesty: RL adds only about two points end to end, and the system stops at proposals. It never runs an experiment.

Limitations

Terms

Source

What people are saying

Related papers

All paper explainers