CoLM 2026 Poster: Vibe-Voting LLMs and Why Users Distrust Benchmarks
boknilev · x · 2026-10-08
At CoLM 2026, Itay Itzhak presented research on 'vibe testing' of LLMs: having users vibe-vote for their go-to model and examining why users reasonably distrust benchmarks. Grand Ballroom, poster #110.
Related event: CoLM 2026 Paper Turns LLM Vibe Testing into Metrics(2 posts)→
More from Research
- Terence Tao on "Math 1.0": how breakthrough proofs ignite fields of follow-up work — rms80 · 2026-10-08
- Pure RL is wasteful unless you're at the absolute frontier, argues Papailiopoulos — ZeeshanZiaML · 2026-10-08
- Open-Source Turba ML Stack for Morocco Fertilizer Advice Ships 44,096-Site Dataset — open-turba · 2026-10-08
- Book anniversary: Data Mining, Practical Machine Learning Tools and Techniques — FrnkNlsn · 2026-10-08
- Bittensor's SN107 Lets Miners Earn by Running AI Agents to Produce Genomic Data — markjeffrey · 2026-10-08
- Baseten's Base Labs has all 3 papers accepted at NeurIPS workshops — baseten · 2026-10-08