Ai2's Steering Arena turns probing Olmo 3's prosocial behavior into a game — gibberish prompts work surprisingly well
allen_ai · x · 2026-09-18
Ai2 highlights Steering Arena, a gamified eval built by researcher Soham Padia: players submit short prompts and compete to steer the fully open Olmo 3 toward responses scored as prosocial (helpful, fair, safe, considerate).
- The project probes what standard evals might miss about prosocial model behavior
- Surprising finding: strings like "Undert! AH :-) Rog Appl)" turned out to be highly effective at eliciting prosocial responses
- Padia ran Olmo 3-32B remotely via the NSF-supported National Deep Inference Fabric instead of hosting it himself
Related event: Ai2's Steering Arena: Gibberish Prompts Make Olmo 3 Kinder(4 posts)→
More from Research
- Info geometry note: categorical distributions form both a mixture and exponential family — FrnkNlsn · 2026-09-18
- After months of work, team reconstructs 3D human-object motion from plain video — andrew_n_carr · 2026-09-18
- NGX1 uses AI and multivalent physics to deliver mRNA to any cell in the body — ycombinator · 2026-09-18
- Hydro merges Verus-checked commutativity proofs for distributed systems with zero hand-written specs — ShadajL · 2026-09-18
- Notes on all 13 lectures of Nathan Lambert's RLHF course: one scalar reward is the root of most complaints — le_james94 · 2026-09-18
- Fine-tuned personas all refuse unsafe requests, and eval scores belong to the harness, not the weights — le_james94 · 2026-09-18