Decider model gains from unmasking and more data, authors deny benchmaxxing

antoine_chaffin · x · 2026-10-07

antoinechaffin says the recent gains in their decider models come from sane decisions: lifting the causal mask and training on more data (crediting the tasksource dataset). Pushing back on benchmaxxing accusations, he points to a top-1 spot on a brand-new leaderboard and a tiny gap between public and private splits — "some models collapse, ours does not flinch." Roadmap ahead includes better reranking, PII detection, and stronger multimodal capabilities.

Related event: Decider model author defends gains as private split matches public leaderboard(2 posts)→

Original post →

More from Models

Models channel →