Researcher warns LLM evals are broken after reading 'truly horrible' model traces

IanArawjo · x · 2026-09-16

hallerite argues it's time for a serious discussion about the state of LLM evals: reading model traces reveals "truly horrible stuff", suggesting a real gap between benchmark scores and actual model behavior. No specific cases are detailed in the post.

Original post →

More from Models

Models channel →