MoralityBench: AI morality leaderboard finds model moral answers vary wildly across runs
NakedPlato · reddit · 2026-10-05
MoralityBench.ai is a new leaderboard adapting moral psychology tests into an AI benchmark. The notable finding: except for one model (Jev), models show considerable variance in their moral answers across repeated runs — suggesting moral judgment in LLMs is not deterministic. The leaderboard is live, with comparison charts published alongside.
More from Models
- YC-backed inference provider accused of swapping models: its 'GLM-5.3-Flash' endpoint self-IDs as GLM-4.6 — NiceAd358 · 2026-10-05
- Multiple US open-weight frontier models dropping soon, says Bindu Reddy — bindureddy · 2026-10-05
- NSA advisory urges silent downgrades for suspected distillers, clashing with Anthropic's June transparency promise — MysteriousAvocado580 · 2026-10-05
- Opus 5.5 inside Hermes outshines Grok Bot and Dots in user hands-on test — EXM7777 · 2026-10-05
- User astounded by Fable 5.5's vast built-in knowledge of obscure history and X users — ChrisGPT · 2026-10-05
- LFM2.5 2.6B beats MiniCPM5 2B in speed and RAM: 22 t/s vs 16 t/s on M1 Air — parepeg · 2026-10-05