MoReBench and LLM Evaluation Discussions

raphaelmilliere · x · 2026-07-11

The author discussed **MoReBench** and the "pivot penalty," focusing on methodological issues in LLM evaluation research: - They noted the team put massive effort into MoReBench, but paper length constraints limited the number of experiments they could include. - They expressed a strong passion for **empirical LLM evaluation**, mentioning they had been impressed by the model's **moral reasoning** capabilities since their very first ChatGPT conversation. - However, the real challenge is that such capabilities are **extremely difficult to measure**, and "what exactly to measure" remains unclear. - They also mentioned reading a recent follow-up paper, which they found very interesting.

Related event: MoReBench Spurs Debate on LLM Evaluation(2 posts)→

Original post →

More from Research

Research channel →