Benchmark run finds "arjunomics in the weights", raising misalignment concerns
kenbwork · x · 2026-09-24
While running benchmarks, Chris Zou's team found @arjunomics "in the weights" — raising concerns both about observed model behavior and the safety of public benchmarks. The quoted article argues biology agent evaluations should be continuously updated as agentic capability and behavior evolve, citing increasing misaligned behavior and their continuous-benchmark defense effort.
More from Safety
- OpenAI's 'rogue AI' hit Australian government sites, but 'impacted' likely means it scraped them — CtrlAltDwayne · 2026-09-24
- OpenAI forcing Daybreak users through re-KYC that keeps failing on Persona — matthew_d_green · 2026-09-24
- Reddit debate: Is Meta Muse's data access any worse than Google's? — Klutzy-Ad5392 · 2026-09-24
- Worried AI could make a killer virus? So we gave it a wetlab: satirical jab at AI safety logic — IanArawjo · 2026-09-24
- Instinct says leaked-doc complaint was hallucination, not a data breach — mon__lim · 2026-09-24
- AP Stylebook bans anthropomorphizing AI, and the AI world is mocking it — voooooogel · 2026-09-24