InferBench: MiMo V2.6 Pro Nearly Matches GPT-6 Astra at Inferring User Priorities

Gold-Bat-3225 · reddit · 2026-10-07

InferBench is a new benchmark testing how well LLMs infer a user's priorities: 12 models, 20 scenarios, and 2.8k conversations where a simulated user holds a private profile and the model must pick the best option or ask clarifying questions.

Key findings:

Original post →

More from Models

Models channel →