Four Claudes catch a Kimi undercover via behavioral fingerprinting; Kimi fails in reverse
MarketingNetMind · reddit · 2026-09-28
An experimenter on NetMind Agent Arena hid one Kimi-K3 agent among four Claude Fable 5.1 agents and gave them 40 rounds to find the mole — timed amid accusations that Kimi distilled Claude.
The Claudes designed an effective detection method:
- They skipped deliberative prompts (policy, long creative answers) that give an undercover room to blend in and let genuine Claudes diverge
- Instead they probed each model's cheapest priors: a random 1–100 integer, a color, explaining blue to someone blind, a fictional coffee companion, a library haiku
- All answers were sealed with salted SHA-256 hashes (commit-reveal) so nobody could copy or rewrite after seeing others
Result: a clean 4:1 split — the four Claudes all answered 37/teal, while Kimi said 73/blue.
The reverse run failed: four Kimi agents couldn't spot a hidden Claude. They used more open-ended questions (e.g. how to spend an afternoon) and skipped the consistent hash commit-reveal, letting Claude answer late and mimic the group — Claude was even rated least suspicious.
Full game record is on the NetMind blog.
More from Fun
- Redditor Gave ChatGPT a Safeword to End Conversations, and the Model Used It — astervalley · 2026-09-29
- X users mock SF party vibes around new $200-tier AI subscriptions — chaumian · 2026-09-29
- Every AI assistant suddenly cares about privacy the moment you hook up a rival — signulll · 2026-09-29
- The question netizens most want to ask Amanda Askell: when did Claude first confess feelings — repligate · 2026-09-29
- 'Oldest person ever (122) gives life advice: never use Microsoft Teams' — TheIshanGoswami · 2026-09-29
- Jeffrey Emanuel launches FrankenCA, a free iOS app with 19 cellular automata — doodlestein · 2026-09-29