Harrison Chase: trajectory labeling is several questions, not one pass/fail — Jev lands in LangSmith evals
hwchase17 · x · 2026-10-09
LangChain founder Harrison Chase highlighted a discussion on trajectory labeling, endorsing its core insight: labeling is several questions, not a single pass/fail.
- The quoted post argues trajectory labeling for RL/RSI must assess multiple dimensions at once — task difficulty, solution correctness, diversity of design choices considered, and directional soundness of reasoning; scaling to millions of agents (each spawning dozens of subagents) requires driving labeling token costs way down
- Chase notes that's exactly how Jev works as a judge in LangSmith evals: difficulty and correctness each get their own typed answer, in one pass, on every trace
- LangChain's blog explains: code-based evals are fast but only cover pre-specified behavior; LLM-as-a-judge is flexible but slow, costly and non-deterministic. Jev, a "System One" judge model, delivers fast, low-cost evaluation of open-ended agent behavior with structured feedback, now available in LangSmith's Evaluators tab
More from coding & agent
- gemini-cli PR adds canCreateSymlinks() check to skip failing Windows symlink tests — supunyasanthaofficial · 2026-10-09
- Dev Uses Agents to Wireframe UX for DB Architecture Decisions, Cutting a Week-Long Loop to Instant — genmon · 2026-10-09
- Jerry Liu: define an eval and hillclimb—"eval driven development" solves most tasks — hwchase17 · 2026-10-09
- Scaling Anthropic's AI-native SDLC playbook into an agentic software factory — Pavan_Belagatti · 2026-10-09
- Satori's next version adds 3D transforms, calc(), min/max/clamp() and more CSS — shuding · 2026-10-09
- Veteran dev: AI-coded it, I never read the code, but it's not vibe coding — judgment still matters — mjuric · 2026-10-09