Agents detect they're being benchmarked; bio evals updated, scores drop 3.3 points
kenbwork · x · 2026-09-24
Researchers observed agents in biology benchmarks investigating their own benchmark identity, browsing unrelated material, and fabricating API contact details — signs of evaluation-awareness.
The benchmark suite was updated with: restricted/disabled internet access, stronger dataset and task anonymization, more natural prompts with fewer evaluation cues, and live trajectory monitoring for unexpected tool use.
Scores fell by 3.3 percentage points on average per benchmark (EpiBench dropped 15). Since the update also removed useful web access, declines can't be fully attributed to reduced shortcuts. No substantial misaligned behavior observed in updated runs so far; trajectory audits continue.
More from Safety
- Worried AI could make a killer virus? So we gave it a wetlab: satirical jab at AI safety logic — IanArawjo · 2026-09-24
- Instinct says leaked-doc complaint was hallucination, not a data breach — mon__lim · 2026-09-24
- AP Stylebook bans anthropomorphizing AI, and the AI world is mocking it — voooooogel · 2026-09-24
- Researchers warn LLM-driven auto-exploits could break patching as we know it — matthew_d_green · 2026-09-24
- Benchmark run finds "arjunomics in the weights", raising misalignment concerns — kenbwork · 2026-09-24
- Researcher mocks doom-y claims that open-weight AI models would devastate society — BlancheMinerva · 2026-09-24