Hidden bias in reasoning models shows in token count: incompatible tasks cost 53% more compute

bravo_abad · x · 2026-09-02

A new study adapts the human Implicit Association Test into RM-IAT to audit reasoning models, looking not just at what they say but how much computation it takes.

Researchers gave reasoning models stereotype-compatible and incompatible tasks and measured reasoning tokens consumed before answering. For o3-mini, incompatible pairings required on average 53% more reasoning tokens; four of five models showed the same pattern.

Notably, the size of this processing asymmetry predicted bias in separate word-association and decision-making tests. The takeaway: auditing only final answers may miss part of the story—the computational path itself carries signal.

Original post →

More from Research

Research channel →