500K Names Tested: Unequal Tokenizer Support Skews LLM Judgments in Hiring and Lending

Mir Tafseer Nayeem · hf · 2026-09-30

A new paper argues that fairness evaluations using "matched names" fail at the lexical interface: matched names are not necessarily matched inputs.

The conclusion: unequal lexical support is demographically structured at the input and remains visible in task-relevant model computation; behavioral comparability begins with lexical comparability.

Original post →

More from Safety

Safety channel →