Percy Liang on Simile: per-query confidence matters more than average eval accuracy in simulation
joon_s_pk · x · 2026-08-26
Percy Liang amplified Simile's first technical blog post. Simile trains two types of models: simulation models and confidence models, where the latter predicts the accuracy of population simulations per query in real time. Liang argued confidence is paramount to simulation: if a coding agent messes up you can often tell and repair, but a wrong simulation may go unnoticed and lead to bad consequential decisions. Rigorous evals only tell you average performance over a population and use case; the confidence model tells you how well the model does on each query.
More from Research
- Zeus Weather Model Adds Energy Awareness, Incentivizes European Forecasts — const_reborn · 2026-08-26
- Study: LLMs Are Effective Mode Seekers, Not Faithful Samplers — gerardsans · 2026-08-26
- SWE Refactor Bench: Only 5.4% of 520 Runs Pass, Best Model Opus 5 Scores 47/100 — YouJiacheng · 2026-08-26
- Dribbling the AI Watermark Directly In-Prompt — JulianHabekost · 2026-08-26
- Agent Skills Actually Hurt Performance? WebDev Benchmark Study Reveals — dair_ai · 2026-08-26
- You only need linear algebra, calculus, and probability for ML math — TivadarDanka · 2026-08-26