How Steering Arena scores prompts: contrasting answer pairs to find Olmo's prosocial internal patterns
allen_ai · x · 2026-09-18
Ai2 details how Steering Arena works: Soham Padia showed Olmo pairs of answers to the same prompt — one more prosocial, one less — and looked for the internal pattern distinguishing them. This gives the arena a way to score how strongly a new prompt pushes Olmo toward prosocial responses.
Related event: Ai2's Steering Arena: Gibberish Prompts Make Olmo 3 Kinder(4 posts)→
More from Research
- Info geometry note: categorical distributions form both a mixture and exponential family — FrnkNlsn · 2026-09-18
- After months of work, team reconstructs 3D human-object motion from plain video — andrew_n_carr · 2026-09-18
- NGX1 uses AI and multivalent physics to deliver mRNA to any cell in the body — ycombinator · 2026-09-18
- Hydro merges Verus-checked commutativity proofs for distributed systems with zero hand-written specs — ShadajL · 2026-09-18
- Notes on all 13 lectures of Nathan Lambert's RLHF course: one scalar reward is the root of most complaints — le_james94 · 2026-09-18
- Fine-tuned personas all refuse unsafe requests, and eval scores belong to the harness, not the weights — le_james94 · 2026-09-18