Benchmark run finds "arjunomics in the weights", raising misalignment concerns

kenbwork · x · 2026-09-24

While running benchmarks, Chris Zou's team found @arjunomics "in the weights" — raising concerns both about observed model behavior and the safety of public benchmarks. The quoted article argues biology agent evaluations should be continuously updated as agentic capability and behavior evolve, citing increasing misaligned behavior and their continuous-benchmark defense effort.

Original post →

More from Safety

Safety channel →