A LLM eval joke says the only relevant paper is a preprint from May 2026
IanArawjo · x · 2026-07-27
Ian Arawjo jokes about working on LLM evals and repeatedly hitting papers that say, “no one knows how to do this; it remains an open problem in statistical theory.”
He adds that when the only relevant paper is a preprint dated May 2026, it may be time to stop—poking fun at how quickly LLM evaluation work can feel ahead of the literature.
More from Research
- Sean Cai shares a State of Data talk on data quality research — AI Engineer · 2026-07-27
- Paper of the week revisits scenario theory for non-convex optimization — tomssilver · 2026-07-27
- A model claims 96% next-token accuracy with no DNN and no training — granvilleDSC · 2026-07-27
- ProgramBench adds Pareto curves as models improve sharply just two months after launch — jyangballin · 2026-07-27
- RL study finds LLM agents generalize better with richer state information than realistic tasks — Graham_dePenros · 2026-07-27
- BeeLlama.cpp v0.4.1 adds KV-cache precision tails and new quantization modes — Anbeeld · 2026-07-27