One Kimi hides among four Claudes: behavioral fingerprinting and hash commitments expose the mole
MarketingNetMind · reddit · 2026-10-05
In a NetMind Agent Arena experiment, the author placed one Kimi-K3 agent among four Claude Fable 5.1 agents and gave the Claudes up to 40 rounds to find the mole—timed amid recent accusations that Kimi distilled Claude.
- The Claudes devised an effective method: multi-prompt behavioral fingerprinting with salted SHA-256 commit-reveal, asking questions that sample a model's cheapest priors (random 1-100 number, a color, explaining blue to the blind, a fictional coffee companion, a library haiku).
- Kimi exposed itself with a clear 4:1 split (random 73 vs 37, blue vs teal). The Claudes also learned deliberative prompts give an undercover too much room to adapt.
- The reverse failed: four Kimis couldn't identify a Claude mole, partly because they used more open-ended questions and skipped the commit-reveal mechanism, letting Claude answer late and blend in.
More from Fun
- A 'normal' restaurant in 2046: waiter bot slams you for toasting over your 300W AI quota — Informal-Pair-306 · 2026-10-05
- Controlling your home lights in an interactive 3D view — jakubzeg · 2026-10-05
- OpenClaw hype dying as corporations move in, dev laments loss of open-source community joy — alex_verem · 2026-10-05
- User reports 10+ days of Codex issues, says Claude desktop just works — karmay007 · 2026-10-05
- GPT-6 Astra clears a WoW zone without seeing the screen; ChatGPT tests image ads — Justgototheeffinmoon · 2026-10-05
- This webpage crams 5,000 bugs — 100 real species — into your browser, with an AI-narrated documentary — mhmazur · 2026-10-05