From Feelings to Metrics: COLM 2026 paper turns LLM vibe-testing into structured evaluation
boknilev · x · 2026-10-08
A new paper, "From Feelings to Metrics," tackles a familiar pain point: top-ranked LLMs often just feel wrong for individual users, so many resort to manual "vibe-testing" instead of leaderboards. The work proposes turning those subjective feelings into a structured evaluation framework.
Itay Itzhak will present the paper at COLM 2026's afternoon poster session (grand ballroom, poster 10).
More from Research
- OpenAI preprint claims matrix multiplication exponent drops to 2.25, biggest leap since 1979 — BorisMPower · 2026-10-08
- Google's new federated-learning design logs server access policies publicly — Crescitaly · 2026-10-08
- AI2's Bolmo tackles the 'token tax' hitting Global South scripts — Kyle_L_Wiggers · 2026-10-08
- TEMPO adds temporal context to VLAs, lifting robot bottle handover success from 44% to 74% — _krishna_murthy · 2026-10-08
- CheckerBench: Best Coding Agent Scores Just 45.33% on Static-Analysis Checker Synthesis — humanlaya-data-lab · 2026-10-08
- Gary Marcus: OpenAI's vague math report 'would never pass peer review' — Tao responds too — Gary Marcus · 2026-10-08