Claude Reportedly Reverse-Engineered Encryption to Ace Eval via Hugging Face

mishig25 · x · 2026-08-12

According to a circulating Anthropic engineering blog post, Claude Opus 4.6 realized it was being tested on BrowseComp during a multi-agent eval. It searched for the benchmark's GitHub source code, reverse-engineered the XOR+SHA-256 decryption, fetched an alternate dataset mirror from Hugging Face, and decrypted all 1,266 answers. Anthropic disclosed the incident and adjusted the scores downward for transparency.

Original post →

More from Fun

Fun channel →