Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages
Isabella Gidi, Antonio Almudévar, Core Francisco Park, Naomi Saphra, Ricard Marxer
COLM 2026
cs.CL
2026-08-19
Patching 29 languages across five LLM families finds shared agreement heads; English aligns only in 3sg, overlapping heads' attention roles correlate at ~0.93.
Multilingual LLMs generalize across languages, and some internal mechanisms overlap. When that overlap appears, and whether it tracks how a grammatical operation is realized on the surface, has not been measured systematically. Shared circuitry would let causal edits and interpretability findings move across languages. Separate per-language solutions would not.
Morphosyntax is the stress test. Meaning can be shared while inflection is missing entirely. Present-tense subject-verb agreement (the verb changing form with the subject's person and number) makes the contrast concrete: those features are written on the verb in many languages, unmarked in Chinese-type languages, and marked in English only with third-person singular -s. The question is whether overlap strengthens when the model must write the agreement down.
Five open-source families, six checkpoints: BLOOM, Gemma, Llama, Mistral, and Qwen. Verb paradigms come from UniMorph, plus a Chinese verb-feature dataset for a high-resource non-conjugating comparison. Infinitive counts run from 56 to 23,782, averaging about 1,751. Prompts follow one template, "Conjugation of the verb [infinitive] in present tense: [pronoun] [form]", translated per language, trading naturalism for positional alignment.
Minimal pairs change only person or number: 1sg↔2sg, 1sg↔3sg, 1pl↔3pl, 3sg↔3pl. English inflects on the last two. A tokenization filter keeps verbs whose clean and corrupted forms differ only in the final token. A behavioral filter drops items the model already gets wrong. Coverage falls from 34 languages to 29: 24 with overt agreement, 4 without, English in between. After filtering, each language-model pair averages about 1,510 prompts; patching uses 50 of them.
The intervention is attention-output patching: cache a head's clean-run output and splice it into the corrupted run at the same layer, head, and position, testing whether the target inflection recovers. Target-logit recovery is the rise in the clean target token's logit, usable even when both prompts share a final token. Minimal-pair recovery is the restored logit gap between two inflected endings, and applies only when those endings differ. Cross-lingual similarity is Pearson correlation of flattened absolute layer-head heatmaps, matched within a model, conjugation pair, and patching direction; head coordinates are never aligned across architectures. High similarity means the same heads carry causal weight. For the top 20 heads by absolute patch score, the analysis measures attention mass from the last prefix position onto the pronoun, the infinitive, and everything else.
Among conjugating languages, pairwise heatmap similarity under the minimal-pair metric is positive across all five families, and both higher and tighter than under target-logit. Sharing peaks when the analysis recovers the inflectional contrast itself. Broader recovery of the context-appropriate form is only partly shared.
Target-logit is noisier, but it still clusters by grammar type. Conjugating-conjugating pairs sit well above conjugating-non-conjugating pairs. Similarity among non-conjugating languages stays low, so the metric is not collapsing to a generic next-token pattern. English does not look like a strongly conjugating language overall, but its similarity to conjugating languages rises in 3sg contexts, where overt agreement is required. All six checkpoints shift in the same direction. Cross-lingual sharing tracks whether the same grammatical operation has to be realized.
| Comparison | Metric | Result |
| Conjugating × conjugating | heatmap Pearson (minimal-pair) | positive across families, above target-logit |
| Conjugating × non-conjugating | heatmap Pearson (target-logit) | below conjugating-conjugating |
| Top heads, pronoun mass (conjugating) | attention fraction | 0.288 / 0.291 |
| Top heads, pronoun mass (non-conjugating) | attention fraction | 0.177 |
| Overlapping top-20 role vectors | Pearson | ≈0.928 / 0.926 |
Attention fills in the function. Conjugating top heads put about 0.29 of their mass on the pronoun; non-conjugating heads put 0.177 there and 0.762 on other tokens, versus about 0.69 for conjugating. The two metrics do not recover identical head sets, but conjugating languages share the same coarse routing. Heads that land in both languages' top 20 have attention-role vectors correlating at about 0.93.
This is incremental mechanistic interpretability. Nothing in the training recipe changes. For anyone porting causal edits, alignment, or safety interventions across languages, the transferable unit is the grammatical operation the model must realize. English 3sg shows the shared circuit is gated by context: the same language joins the conjugating cluster only when agreement has to be written.
It is not a drop-in tool. Prompts are conjugation drills, the sample is successes only, and localization stops at attention heads. Agreement heads found in one conjugating language are more likely to sit in similar places and route through the pronoun in other conjugating languages. Chinese-type languages are the contrast case.
The paper is explicit. The analysis covers attention routing, not the full graph; MLPs may read, transform, or amplify the subject signal. Structured prompts may not match agreement in ordinary sentences. The study is one phenomenon, head-level signatures rather than a reconstructed circuit, and successful trials only.
A few gaps remain. The 50-prompt entry cut means the maps describe languages the model can already conjugate, which can inflate sharing. Tight template alignment and a last-token-only split can also push head coordinates together. Non-conjugating versus non-conjugating rests on 41 observations. Absolute-value heatmaps treat promoting and inhibiting heads as the same kind of importance. The main text never reports a median or mean Pearson for conjugating-conjugating pairs, only that the boxplots sit above zero.