27B Research Agent Outperforms Opus and GPT-5.5

rohanpaul_ai · x · 2026-08-15

A 27B research agent named Faraday has outperformed Claude Opus 4.8 and GPT-5.5 on paper replication tasks. Its core strategy decouples "research direction" from "code implementation":

On 68 unseen AI-for-science tasks, Faraday scored 0.791, beating Opus (0.748) and GPT-5.5 (0.729), winning on 60% of tasks. Faraday was also more likely to test actual mechanisms rather than taking shortcuts.

Related event: Inherent Labs' 27B AI Scientist Faraday Beats Claude Opus 4.8 and GPT-5.5 at Paper Replication(7 posts)→

Original post →

More from coding & agent

coding & agent channel →