Structure-only AI audit: Jev predicts outcomes of 2,029 real calls at AUC 0.78 for $3
alexcovo_eth · x · 2026-09-30
A team stripped 2,029 real AI-receptionist calls down to pure structure—turns, tool calls, workflow stages, timing—no transcripts or audio, and fed it to Jev. Zero-shot, it made 38,012 turn-level forecasts at 118 ms median latency and reviewed every call with five typed questions, producing 10,145 answers in 26 seconds with 256 requests in flight. By the halfway point it separated booking vs non-booking calls at AUC 0.78, ranking them correctly 94% of the time near the end—all for $3. Caveat: Jev over-focused on visible errors the agent usually overcomes. The takeaway: structured telemetry alone can cheaply audit and forecast agent call quality.
More from coding & agent
- OpenAI launches new Codex Cloud with configurable environments; Agents API preview adds computer use — charliermarsh · 2026-09-30
- OpenAI upgrades Codex Security Cloud with cyber-capable models via Daybreak Blue by default — OpenAI · 2026-09-30
- OpenAI Codex harness is now open source — BorisMPower · 2026-09-30
- TypeSafe's Workflow Evals: decomposing policies into workflows beats prompts on accuracy, cost and speed — multiply_matrix · 2026-09-30
- xAI launches Grok Bot: AI teammates that log into your tools and finish work — XFreeze · 2026-09-30
- OpenAI launches Decisions API, its answer to Jev — Rasmic · 2026-09-30