Marco Polo Benchmark Shows Chinese Model Retention Falling, Sparking Questions
teortaxesTex · x · 2026-09-23
- teortaxesTex flags an odd trend: Chinese models' retention on the Marco Polo benchmark appears to be falling.
- The author posts data screenshots pointing out several inconsistencies in the results and asks why — prompting debate over how Chinese models' long-context performance is actually being evaluated.
Related event: Marco Polo Data Shows Chinese AI Research Output Up, US Retention Down(3 posts)→
More from Models
- Grok 4.7 flops in 100 multi-agent coding evals despite insightful solutions — teortaxesTex · 2026-09-23
- Will rumored GPT-6 'Sol' actually ship inside ChatGPT? — flowersslop · 2026-09-23
- Claude Opus 5.5 Said to Fall Back on Frontier LLM Dev Tasks, Drawing Fire — basedjensen · 2026-09-23
- Higgsfield demo: Claude Opus 5.5 crushes GPT-6 Astra at 3D game generation — VraserX · 2026-09-23
- User Gives Opus 5.5 Creative Tools and Asks What It Dreams About — angrypenguinPNG · 2026-09-23
- Yuchen Jin: Opus 5.5 underwhelms, frontier LLM coding has plateaued — Yuchenj_UW · 2026-09-23