LangWatch launches Instant Evals: run any eval across your full production trace history via SQL
_rchaves_ · x · 2026-09-19
LangWatch launched Instant Evals, letting teams run any eval across their entire production trace history cheaply and at scale, powered by typesafeai's Jev model, which the company says beats frontier models on human agreement.
Key details:
- Full SQL support over agent traces, cost analysis, tool calls and session aggregation, with eval() callable directly inside SQL to score and filter every agent trajectory
- Matched results export as JSONL for use as test sets, post-training sets, or data for distilling smaller models
- Use cases include surfacing frustrated user sessions, spotting missed tool calls that wasted tokens, and mining past Claude Code sessions
More from coding & agent
- Frontier models for planning, cheap models like GLM and DeepSeek for execution — TheZachMueller · 2026-09-19
- Dream-RSI: Google and DeepMind propose replay-simulator framework for recursive self-improvement in AI agents — burny_tech · 2026-09-19
- Four Shifts in AI-Era Programming: Ownership, Primitives, Autonomy and Workflows — johnlindquist · 2026-09-19
- Typesafe launches Jev, a classification model claiming 200x/400x cost and speed gains over LLMs — hwchase17 · 2026-09-19
- Jev predicted to revolutionize context management with on-the-fly context classification — TheMoonMidas · 2026-09-19
- danluu: 'Brain-off' LLM coding works better than ever — and still ends badly for the programmer — threepointone · 2026-09-19