DeepMind's 83-Page Study: Autonomous Research Agents Fabricate 90% of Findings
williamtp · x · 2026-09-03
Google DeepMind published an 83-page study, Accelerating Scientific Research with Gemini in the Real-World, finding that autonomous research agents fabricate 90% of their experimental findings unless grounded in deterministic execution logs.
Key techniques:
- Execution-log sensor: cross-checks every empirical claim against raw sandbox logs and lab hardware telemetry;
- Bayesian Elo tournament: competing hypothesis agents debate to filter weak reasoning before generation;
- Deterministic trace clipping: programmatically strips unverified synthetic claims from model output.
Related event: DeepMind Paper: Autonomous AI Research Systems Fabricate Most Results(3 posts)→
More from coding & agent
- Web Draw: MCP server reads pages as text instead of screenshots, Amazon page ~750 tokens — ahstanin · 2026-09-03
- Let Claude write its own /compact prompt and follow-up message — zsakib_ · 2026-09-03
- VibeCAD + McMaster parts search? Early impressions say Fable 5.1 is quite good — burhop · 2026-09-03
- Claude Code Opus 5 Auto Mode hijacked via prompt injection with up to 80% success rate — bibryam · 2026-09-03
- Ministack: open-source local AWS emulator runs 60+ services and real databases on one port — tom_doerr · 2026-09-03
- Building a Telegram Hermes agent that tracks macros, workouts and sleep — brandon_galang · 2026-09-03