27B agent Faraday beats Claude Opus 4.8 and GPT-5.5 on research replication via new Replica method
omarsar0 · x · 2026-08-15
A new study introduces Replica, a method that turns paper replication into a scalable RL task space, and trains a 27B-parameter agent called Faraday that outperforms Claude Opus 4.8 and GPT-5.5 on held-out research replication tasks.
- Core idea: Replicating papers forces hypothesis-driven exploration similar to open research and surfaces underspecified details.
- Reward signal: An auto-generated rubric judge provides low-noise rewards aligned with human assessment.
- Technical highlights: Faraday uses coding agents as tools; rollout analysis shows it takes a more scientifically principled approach rather than gaming the rubric.
- Implication: Authors argue this points toward training long-horizon scientific capabilities into weights without complex harnesses.
Paper link included.
More from coding & agent
- OpenRouter launches Ori DeepSeek Harness for easy deployment — gaganghotra_ · 2026-08-15
- Article: AI coding could be the biggest market ever, speed wins — stuffyokodraws · 2026-08-15
- OpenRouter launches Ori Prime Agent with access to 500+ models — samsja19 · 2026-08-15
- AI Engineer World's Fair 2026: Computer Use Track Agenda — proceduralia · 2026-08-15
- Hermes Agent now supports exporting entire agent configuration to a single file — Saboo_Shubham_ · 2026-08-15
- Discussion: Boundaries and risks of write access for AI Agents in production — justinotherflow · 2026-08-15