New COLM work uses mechanistic interpretability to study how LMs process island constraints
jessyjli · x · 2026-10-08
Presented at COLM, this new work with kmahowald examines Island Constraints — perhaps the most-studied phenomenon in theoretical linguistics — and how language models process it, using mechanistic interpretability techniques to guide the exploration, connecting interpretability, cognition, and meta-cognition.
More from Research
- LLM Agents Form Factions by Model Family, Costing 30% More Rounds and 55% More Tokens — alex_verem · 2026-10-08
- Hide Model Names From Agents: Labels Cost 55% More Tokens and Drop Success to 81% — alex_verem · 2026-10-08
- AI Pinpoints Melon Yield Mutation in an Afternoon, Ranking #1 of 2,494 Candidates — GlennCameronjr · 2026-10-08
- NumanThabit aims to replace animal studies with sufficient physics simulation — MarwaEldiwiny · 2026-10-08
- Manuel Blum responds on matrix multiplication, citing his four 1964 conjectures — aran_nayebi · 2026-10-08
- Researcher's agent workflow: label 100 examples, have the agent scale to 10k pseudo-labels — ducha_aiki · 2026-10-08