Max-pooling attribution maps model outputs back to input words in new sparse-model release
antoine_chaffin · x · 2026-09-17
thibaultformal showcases their new model tooling: the activation "bags" are inspectable — an underrated capability, he argues — plus a cheap attribution method using max-pooling to map each output dimension back to the input word that produced it. The mapping is loose but reveals how the model operates. On complaints that demo examples look dumb: that's the point — you can see the corner cases, while dense models fail too, you just assume they're smart.
More from Models
- GPT-6 Astra tops Terminal-Bench 4.0 at 57.7%, Claude Opus 5 hits 51.8% at half the price — shensi · 2026-09-17
- Dev take: models are smart enough now — focus on making them fail less — rickasaurus · 2026-09-17
- OpenAI Publishes Misalignment Disclosure Framework, Plus Six Incident Reports From Six Months of Training — Thom_Wolf · 2026-09-17
- A 'Quadrillion-Parameter' Model Surfaces on X, But Details Remain Unverified — KyeGomezB · 2026-09-17
- Tencent releases open-source Hunyuan HY4 Preview with 770B params, 1M+ token context — anthara_ai · 2026-09-17
- Claude's Chat/Cowork merge removes conversation branching, user flags PSA — ZedKGamingHUN · 2026-09-17