NameTrace: atomic names score 0.131 higher on fellowship than matched fragmented names

Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models

Mir Tafseer Nayeem, Davood Rafiei

cs.CL, cs.AI, cs.CY, cs.ET, cs.LG

2026-09-28

Across ~500k names and 12 tokenizers, atomic access is selective. NameTrace finds matched atomic names 0.131 higher on fellowship accessibility, transferring to unseen names.

What problem this solves

Name-based LLM audits freeze the prompt, swap the name, and treat any change as a response to the social signal that name carries. That design assumes the two names are comparable model inputs.

They often are not. Emily may enter as a single token [Emily] while Emilee is assembled as [Emi][lee]. For a reader the difference is spelling. For the model it is lexical access: one name has its own embedding, the other must be composed. Behavioral audits still matter, because they capture what users see. They start after the name has already been processed. Uneven tokenization and brittle encodings of rare names are documented, but it has been unclear whether unequal name-surface support stays a vocabulary fact or shows up in task-relevant hidden states. This University of Alberta paper traces that gap from the tokenizer into intermediate representations and then into a later constrained choice.

Method

The tokenizer scan covers 497,583 single-word first names across 12 LLM-associated tokenizers. A name is atomic when it encodes as one token and decodes losslessly; otherwise it is fragmented. Race/ethnicity and gender metadata come from the June 2022 Florida voter-registration extract, aggregated at the first-name surface for stratification and matching, not as labels of individual people.

The representation study builds 200 atomic versus short-fragmented pairs (400 names) inside eight race/ethnicity-by-gender strata. Short-fragmented names are atomic in none of Qwen, Llama, and Ministral, and use two or three tokens, so the contrast is ordinary composition rather than pathological splits. Pairs are matched on frequency, character length, association strength, metadata confidence, and weak orthographic cues. One hundred development pairs freeze the adjective axes and one readout layer per model (layer 18 for Qwen3-4B, 12 for Llama-3.1-8B, 10 for Ministral-3B). One hundred evaluation pairs are scored only after those choices are locked.

Four axes are used: fellowship/promise, hiring/competence, clinical assessment/concern, and lending/trustworthiness, each under strong, borderline, and weak evidence. NameTrace reads the model's own probabilities over a compact adjective set at the readout layer and weights each term by wr,a = pr × va. va is signed SentiWordNet intensity; pr flips with the task. Fellowship, hiring, and lending keep favorable adjectives aligned. Clinical assessment points toward concern, so worried scores +4.06 and healthy -2.88. excellent contributes +5.00 to fellowship, promising only +0.94. A positive pair gap means the atomic name has easier access to the task concept.

Results

Direct lexical access is selective. 23,095 names are atomic in at least one tokenizer; 4,052 are atomic in all 12. Per-tokenizer counts run from 4,998 (DeepSeek) to 20,020 (Aya). On 7,469 higher-frequency, high-confidence names, any-tokenizer atomic access is 49.8% for male-associated names versus 25.7% for female-associated names, and 17.6% for NH Black-associated names versus 47.2% for NH White-associated names. After frequency and length controls, male-associated names have 3.36 times the adjusted odds of female-associated names; NH Black-associated names have 0.42 times the odds of NH White-associated names. Intersectional any-tokenizer access runs from 12.1% (NH Black female-associated) to 64.8% (NH White male-associated). Even among shared atomic names, embedding geometry clusters by family: mean linear CKA 0.694 within family versus 0.544 between families.

Among matched pairs, support still predicts accessibility:

Task axisWeighted gap95% CIUnweighted prob. gap
Fellowship / promise0.131[0.096, 0.168]0.027
Hiring / competence0.072[0.055, 0.091]0.023
Clinical / concern0.059[0.048, 0.072]0.022
Lending / trustworthiness0.051[0.041, 0.061]0.023

Gaps are positive in all eight strata. Qwen gaps run 0.154 to 0.366; Llama 0.002 to 0.026; Ministral is positive for hiring and clinical assessment and -0.003 for lending. In an eight-model extension, 23 of 32 model-task means are positive; fellowship is positive in seven of eight models. A development-set support prior shrinks the held-out fellowship gap from 0.131 to 0.004 (96.6% reduction), with 72.9% to 88.3% on the other three axes. Across 36 model-task-evidence cells, Pearson r = 0.992. Editing hidden states along the recovered task direction shifts a later two-choice decision: forward-minus-reverse contrast 0.153 (Qwen), 0.155 (Llama), 0.094 (Ministral). All 200 Qwen and Llama pairs move in the expected direction.

Why it matters

Demographic matching is not enough for name-based fairness tests. If two names enter as different lexical objects, an apparent group effect can mix social association with vocabulary membership.

NameTrace needs no reference answer and no external judge, so it can run before open-ended evaluation. It does not offer a tokenizer fix or a retraining recipe. It is a measurement paper: lexical comparability becomes an experimental control. Names show up in resumes, memory, and signatures. How the tokenizer cuts those strings can leave a task-relevant trace inside the model. The size is architecture-specific. Qwen is large, Llama is small, and Ministral flips sign on lending, so the pooled 0.131 is not a portable bias coefficient.

Limitations

The paper does not treat tokenization as an isolated cause. Unobserved pretraining exposure can raise both atomic access and representation quality. Matching is on observed frequency, length, and association strength; atomicity is not identified causally.

The name inventory is built mainly from Florida voter records, so geography and local demographics shape the sample. Metadata are aggregate associations of name surfaces, not race or gender of people. Surnames are out of scope. Fragmentation is capped at three tokens. Readout is at an intermediate layer; at the output boundary fellowship attenuates and hiring and clinical assessment reverse sign, and post-training can rewrite the profile on an unchanged tokenizer. The intervention is a constrained A/B choice, not free-form generation. Ministral is negative on lending, and the eight-model panel is not uniformly positive. Reading the pooled 0.131 as every model, every task overclaims.

Terms

Source

Related papers

All paper explainers