Just reversing option order moved one model 7 points and changed a quarter of answers
colinmcnamara · x · 2026-09-24
A key caveat from the eval thread: simply reversing the order of options moved one model 7 points and changed a quarter of its answers. Before believing any small benchmark gap, send the identical request twice, then reverse the options and send again — small differences may be pure noise.
More from coding & agent
- eBPF as defense-in-depth for misbehaving agents and RL sandboxes — sloppenheimer · 2026-09-24
- Prime Intellect launches Prime Sandboxes: MicroVM sandboxes purpose-built for RL training — willcb · 2026-09-24
- OpenThai-SystemOne: An 0.8B Open Model That Answers Typed Decisions With Calibrated Probabilities — lmoroney · 2026-09-24
- Why Every Agent Harness Will Need Some Jev Glue, According to One Framework Author — Trainer_Intelligent · 2026-09-24
- Critical WordPress RCE CVE-2026-87902 reproduced for $4.93 with an AI agent — evilsocket · 2026-09-24
- GLiClass Lands in OpenClaw: Local ONNX Decision Model Replaces Costly LLM Calls — steipete · 2026-09-24