Decider model gains from unmasking and more data, authors deny benchmaxxing
antoine_chaffin · x · 2026-10-07
antoinechaffin says the recent gains in their decider models come from sane decisions: lifting the causal mask and training on more data (crediting the tasksource dataset). Pushing back on benchmaxxing accusations, he points to a top-1 spot on a brand-new leaderboard and a tiny gap between public and private splits — "some models collapse, ours does not flinch." Roadmap ahead includes better reranking, PII detection, and stronger multimodal capabilities.
More from Models
- Paradigm releases tech report for Limite 1B-Violetto, a from-scratch math model — tensorqt · 2026-10-07
- ChatGPT Reportedly Removes Free-Tier Text Chat Limits and Upgrades Default Model — hey_abusiddik · 2026-10-07
- Qwen3.8 Flash Next IQ1_M hits 55 tok/s on a 5060 Ti 16GB and still codes well — bobaburger · 2026-10-07
- EmbeddingGemma 2 runs offline multimodal RAG on phones in ~191MB–567MB of RAM — Saboo_Shubham_ · 2026-10-07
- Paradigm evals its math model across 7 hard benchmarks, releases full eval suite — tensorqt · 2026-10-07
- Limite 1B borrows nanogpt speedrun architecture: NorMuon, MUDD variant and XSA — tensorqt · 2026-10-07