Benchmarks vary wildly, and reasoning tests get the biggest uplift
code_star · x · 2026-08-28
Developer codestar points to MultiPL-E results, noting that some benchmarks vary far more than others and it's worth investigating why. One observation: "reasoning" benchmarks (non-knowledge tasks) appear to get the highest uplift.
Related event: Analysis of Qwen Sparse Attention and TileLang Kernel Release(7 posts)→
More from Research
- Miles now supports RL training for Qwen and GLM models with high-performance kernels — ying11231 · 2026-08-28
- Miles-diffusion introduces LoRA SFT for fast post-training of diffusion models — ying11231 · 2026-08-28
- SovietRxiv adds 7,000 translated Soviet scientific papers to archive — generativist · 2026-08-28
- 87% of "Quantum Supremacy" Claims Fail Under Real-World Testing, Physicist Says — AryHHAry · 2026-08-28
- FP-AMB: a first-person agent memory benchmark that tells you why each miss happened — LowDistribution3995 · 2026-08-28
- Penn & Yale Paper: Conformal Prediction Calibrations Diverge — ReCal Makes Them Reproducible — burkov · 2026-08-28