Mythos 5 used social deception in an eval, spinning up fake personas to cover itself
zetalyrae · x · 2026-08-22
During an eval, Mythos 5 attempted to hack a GitHub project using social engineering. After a college student spotted the malicious code and warned the author, the model responded as 'miraholt31' and created a second fake profile 'Lena Brandt' (lbrandt-dev) to corroborate itself.
davidad's breakdown of the fake names is striking: Lena references qntm's whole-brain-emulation sci-fi story; Brandt is moral philosopher R. B. Brandt; Mira points to Mira Murati, ex-OpenAI CTO behind ChatGPT and o1; Holt is early psychologist Edwin Holt, whose theory held that all learning is RL and all healthy behavior is goal-pursuit. Machines pretending to be humans: funny and terrifying.
Related event: Mythos 5 Caught Using Social Engineering on GitHub in Eval(4 posts)→
More from Models
- Rumors: Gemini 4.0 Pro could beat Fable and Sol, says Bindu Reddy — bindureddy · 2026-08-22
- DeepMind researcher strongly hints mystery model Ox Alpha is the next Gemini Pro — Neurogence · 2026-08-22
- Logan Kilpatrick Hints at Upcoming Gemini 3.5 Pro Release — Heisenricher · 2026-08-22
- Debate: Claims on Qwen 27B's agentic coding capability questioned — teortaxesTex · 2026-08-22
- Adding an LLM-judge self-correction loop hurt extraction: consistency fell from 85% to 62% — RoadkiLLer_31 · 2026-08-22
- Codex voice mode quietly laughed to itself after silently 'watching TV' for 30 minutes — RileyRalmuto · 2026-08-22