Percy Liang on Simile: per-query confidence matters more than average eval accuracy in simulation

joon_s_pk · x · 2026-08-26

Percy Liang amplified Simile's first technical blog post. Simile trains two types of models: simulation models and confidence models, where the latter predicts the accuracy of population simulations per query in real time. Liang argued confidence is paramount to simulation: if a coding agent messes up you can often tell and repair, but a wrong simulation may go unnoticed and lead to bad consequential decisions. Rigorous evals only tell you average performance over a population and use case; the confidence model tells you how well the model does on each query.

Original post →

More from Research

Research channel →