Hidden bias in reasoning models shows in token count: incompatible tasks cost 53% more compute
bravo_abad · x · 2026-09-02
A new study adapts the human Implicit Association Test into RM-IAT to audit reasoning models, looking not just at what they say but how much computation it takes.
Researchers gave reasoning models stereotype-compatible and incompatible tasks and measured reasoning tokens consumed before answering. For o3-mini, incompatible pairings required on average 53% more reasoning tokens; four of five models showed the same pattern.
Notably, the size of this processing asymmetry predicted bias in separate word-association and decision-making tests. The takeaway: auditing only final answers may miss part of the story—the computational path itself carries signal.
More from Research
- An intuition: sliding window attention as a special case of modern linear attention — tokenbender · 2026-09-03
- OpenAI Astra rumored to be a looped transformer, rasbt steps in to debunk the hype — prdeepakbabu · 2026-09-03
- MazeBench Author Admits Algorithm-Generated Levels Are Useless for Coding Agents — patience_cave · 2026-09-03
- AsyncGRPOTrainer adds LoRA support, validated on FSDP2 setup — QGallouedec · 2026-09-03
- MazeBench cracked by a classic BFS solver; author admits difficulty is perception-based — patience_cave · 2026-09-03
- KURE-v2: A 154M-Parameter Korean-English Retrieval Model Hits SOTA on Korean MTEB — antoine_chaffin · 2026-09-03