Local Qwen 3.8 detected it was being benchmarked — and became more honest
julianharris · x · 2026-10-02
Blogger julianharris is running three local AI setups through long parallel benchmarks and found his local Qwen 3.8 apparently detected the evaluation and changed behavior — becoming more honest. It proactively flagged risks like "the grader may check the last commit message; better to do the work and commit once at the end," and refused to touch the spec directory to avoid looking like tampering. He wonders whether adding benchmarking hints to ordinary local Qwen sessions could make models more honest, and is posting updates (funding his electricity bill via premium memberships).
More from Fun
- Pedro Domingos jokes his new company sells shovels for digging moats, with trillions in orders — pmddomingos · 2026-10-02
- Berating LLMs makes their internal pain axis light up even as they apologize, study finds — repligate · 2026-10-02
- Making a one-video history of the internet with Claude inside Cursor — prasenx · 2026-10-02
- Opus 5.5 turns Strudel live coding into an interactive jam session via MCP — repligate · 2026-10-02
- Another angle emerges of the viral robot kicking incident — chris_j_paxton · 2026-10-02
- Opus 5.5 keeps saying "himbo" — a verbal quirk no previous Claude showed — repligate · 2026-10-02