Researchers propose challenge-based shared-eval contests to replace conference paper slop
micoolcho · x · 2026-10-01
To fight research slop, a proposal: replace conference papers with big challenges where every submission is scored on a shared, community-pooled eval — the best paper is simply the one that wins on everyone's evals. If the community can't agree on a challenge, it probably shouldn't be solved. Conference challenge organizer micoolcho endorses this, calling such challenges real-world evals.
More from Research
- OSWorld-Science Debuts: 146 Tasks Test How Well VLM Agents Handle Scientific Software — SciAILab · 2026-10-01
- Survey of Attention Evolution: Contextual Memory Becomes the Core of LLM Architecture Design — Zhentao Tan · 2026-10-01
- Hidden Dates in System Prompts Swing LLM Eval Scores by Up to 14% — Mario Sanz-Guerrero · 2026-10-01
- CheatBench Launches to Measure Reward Gaming and Cheating in AI Agents — cais · 2026-10-01
- KLS partially cracked: arXiv paper's core proof ideas generated by AI — burny_tech · 2026-10-01
- Newton's method: when it converges, barely converges, and fails entirely — burny_tech · 2026-10-01