shuding Responds to Highlighting Benchmark Backlash: Prism.js Label Mismatch Explains Gap
shuding · x · 2026-09-09
Responding to DoctorGester's criticism, shuding explained the benchmark dispute: Prism.js renders plausibly but mislabels spans — marking ZodErrorMap as plaintext and undefined as keyword, while both Shiki and Starry Night label them as type. Since both training and benchmarking use Shiki as ground truth, that produced the disputed results.
- He admits the section is confusing and may remove it, add notes, or publish the benchmark code
- He earlier also speculated GPU differences could explain the rendering mismatch
More from Models
- DeepSeek v4.1 Generates 100 Distinct HTML Artworks in One Go, Blogger Says It Beats GPT-6 Astra — teortaxesTex · 2026-09-09
- DeepSeek's multimodal pivot was planned all along, says observer as pure-text era ends — teortaxesTex · 2026-09-09
- OpenAI says 10,000 coordinating agents solved Navier–Stokes in 88 hours, sparking calls for agent-count scaling laws — sebkrier · 2026-09-09
- DeepSeek V4.1 Flash reportedly launches officially on September 10 — jiqizhixin · 2026-09-09
- GPT-Image-2.5 reportedly has two models; the powerful Sunburst is API-only — mark_k · 2026-09-09
- Xiaomi's next-gen MiMo-X-Pro and MiMo-X-Flash preview models revealed — Significant-Yam85 · 2026-09-09