Researcher warns LLM evals are broken after reading 'truly horrible' model traces
IanArawjo · x · 2026-09-16
hallerite argues it's time for a serious discussion about the state of LLM evals: reading model traces reveals "truly horrible stuff", suggesting a real gap between benchmark scores and actual model behavior. No specific cases are detailed in the post.
More from Models
- Open-ended RL is 'orbit by freefall': only progress per GPU-time matters — tensorqt · 2026-09-17
- LLMs' subtler hallucination: confidently winging plans for tasks they know nothing about — generativist · 2026-09-17
- Users note Claude's 5th-gen models obsessively spotlight corrections over substance — repligate · 2026-09-17
- Reddit weighs the best methods for uncensoring open-weight LLMs — Hefty_Wolverine_553 · 2026-09-17
- Qwen 3.8 27B beats Muse 30B on benchmarks but flubs multi-question prompts, user finds — octagoncat23 · 2026-09-17
- StepAudio 3 Realtime Technical Report published on Hugging Face — _akhaliq · 2026-09-17