InferBench: MiMo V2.6 Pro Nearly Matches GPT-6 Astra at Inferring User Priorities
Gold-Bat-3225 · reddit · 2026-10-07
InferBench is a new benchmark testing how well LLMs infer a user's priorities: 12 models, 20 scenarios, and 2.8k conversations where a simulated user holds a private profile and the model must pick the best option or ask clarifying questions.
Key findings:
- Models picked the best option 89% of the time when priorities were stated up front, dropping to 60% when some were unstated.
- GPT-6 Astra: 76%, MiMo V2.6 Pro: 75%, Gemini 3.1 Pro: 69%.
- Astra always asked clarifying questions when preferences weren't stated and never did when they were; Grok 4.6 asked in 18/64 conversations.
- Open-weight models outperformed expectations.
- Concern: 105 of 288 wrong decisions were submitted with >90% confidence.
More from Models
- Mathematician clarifies: open problem lists driven by curiosity, not centralized corporate effort — littmath · 2026-10-07
- HuatuoGPT-3: 27B open medical LLM hits 70.1 on HealthBench, beating GPT-6 Astra — CUHKSZ · 2026-10-07
- Dev Complains Overnight Long-Running Tasks Keep Hitting Usage Limits — willdepue · 2026-10-07
- Gary Marcus on whether frontier LLMs can solve open math problems without symbolic harnesses — GaryMarcus · 2026-10-07
- From botching 9.9 vs 9.11 to tackling the hardest math problems in two years — Yuchenj_UW · 2026-10-07
- OpenAI claims 372 unsolved problems cracked, averaging about 3 hours each — i_dg23 · 2026-10-07