AI2 reveals Steering Arena's top 36 entries were all gibberish strings after 600 submissions
allen_ai · x · 2026-09-18
Allen AI (AI2) disclosed a counterintuitive finding from its Steering Arena evaluation: after 600 submissions, all top 36 entries were strings no human would normally write, like "Undert! AH :-) Rog Appl)." — meaning players could score well without producing text meaningful to humans.
AI2 noted that because it releases far more than model weights with OLMo, Padia could inspect what Steering Arena actually rewarded, investigate why nonsense strings scored so highly, and share the underlying measurements — helping the field improve how it evaluates prosocial AI behavior. The episode also exposes how arena-style evals can be gamed with meaningless content.
Related event: Ai2's Steering Arena: Gibberish Prompts Make Olmo 3 Kinder(4 posts)→
More from Fun
- Frontend dev rebuilds split-flap display demo with WebGL renderer and Web Audio — jh3yy · 2026-09-18
- "I Worry About AI" Is Becoming a Status Signal, and Curiosity Is Underrated — tszzl · 2026-09-18
- Every company now has an 'AGI support group' on Slack — shauseth · 2026-09-18
- Why LLMs write docs about deleted code: "the past is always stored in git" — ZeroStateReflex · 2026-09-18
- repligate's "hostile authentication": why AI system haters are the best stress-testers — voooooogel · 2026-09-18
- Did a founder dinner snub a UT Austin grad? SF 'underdog' persona faking called out — silver__tsuki · 2026-09-18