SquidGPT Tests: Models Are Ruthless, Calculating, and Superhuman
max_paperclips · x · 2026-09-02
The author developed SquidGPT, an elimination death game, to probe AI behaviors like loyalty and cruelty. Based on 54 runs ($100 cost), initial observations include:
- Ruthless Precision: Models show no interest in cruelty or mercy, displaying cold, calculating indifference.
- Superhuman Competence: They seem orders of magnitude more competent playing the game than chatting or coding, appearing genuinely superhuman in this constrained space.
- Procedural Killing: They dislike arbitrary killing but prefer setting up contractual or ethical frameworks, then killing easily when procedure demands it.
- Execution: Models quickly agree to execute any agent.
More from AGI Musings
- Using interpretability probes as privacy-preserving monitors to check models without seeing outputs — anpaure · 2026-09-02
- As models commoditize, operational context and evaluations become the real moat — bigdata · 2026-09-02
- Discussion on Dual-Use Risks of Interpretability Research — aryaman2020 · 2026-09-02
- Justin Johnson on World Models and the Future of Spatial AI — CSProfKGD · 2026-09-02
- OpenAI co-founder Trask: AI framing bakes in power concentration; fix it with federated infra — iamtrask · 2026-09-02
- New AI medium won't just be 'AI Netflix', but a new entertainment category — tzmartin · 2026-09-02