Delip Rao calls out new Qwen3.5-9B-based model for benchmarking latency but not accuracy vs Jev
deliprao · x · 2026-10-10
Delip Rao highlighted a new text-only model built on the Qwen3.5-9B base, trained on "somewhere between a billion and a trillion" tokens, and questioned why its official benchmarking compares latency against Jev but omits accuracy entirely. He directly asked Satya Nadella why—implying selective metric display by the vendor.
More from Models
- Microsoft unveils Microsoft-Decision-1, a fast decision model it claims beats LLMs — MaziyarPanahi · 2026-10-10
- OpenAI ships GPT-6 Sol and Luna with Intelligent UI to all ChatGPT users — aidan_mclau · 2026-10-10
- OpenAI researcher: new Personal AGI models more honest, but eval awareness erodes safety measurement — ericmitchellai · 2026-10-10
- Rumors: Google near recursive self-improvement as all three labs close in on RSI — imjustnewatai · 2026-10-10
- Swap 'math' for 'cancer': AI circle mocks LLM self-comparisons to mathematicians — NachoSoto · 2026-10-10
- Power user says he's nearly hit Grok's weekly limit, urges xAI to double quotas — nima_owji · 2026-10-10