Max-pooling attribution maps model outputs back to input words in new sparse-model release

antoine_chaffin · x · 2026-09-17

thibaultformal showcases their new model tooling: the activation "bags" are inspectable — an underrated capability, he argues — plus a cheap attribution method using max-pooling to map each output dimension back to the input word that produced it. The mapping is loose but reveals how the model operates. On complaints that demo examples look dumb: that's the point — you can see the corner cases, while dense models fail too, you just assume they're smart.

Original post →

More from Models

Models channel →