First-token logprob gating fixes SKIP recall in a local real-time medical scribe, with zero false skips
r-chop14 · reddit · 2026-10-07
A developer applied the Jev/SemIf-style utterance gating idea to a local real-time clinical scribe and shared the full recipe:
- Pipeline: TEN-VAD segments utterances into a Whisper-compatible backend (a 0.6B medical STT finetune), with CAM++ speaker embeddings for best-effort diarization.
- Gating: each utterance gets a one-shot decode by a small model; first-token top logprobs are summed over NOTE/ACT/SKIP. SKIP utterances are buffered until the main model next wakes up, with a 45s/40-word debounce pass as a safety net; the main model holds tools to edit the running note, and prompt caching keeps later passes fast.
- Key finding: a 4B model almost never predicts SKIP, and probability mass alone matched plain prompting (81% accuracy). Digging into logprobs surfaced a usable heuristic — NOTE selected, P(SKIP) ≥ 0.05, and ≤8 words is almost always a SKIP. With it, false SKIPs dropped to zero across runs and SKIP recall jumped from 0–50% to 75–100% on natural consults.
- Comparison: on 40 hand-labelled utterances from a real consult, Jev and this approach both hit 95% accuracy, but at 73ms vs 514ms (Jev via a remote endpoint); on a command-heavy synthetic script Jev fell behind (80% vs 100%). The heuristic was tuned on the eval set, so a held-out validation is still needed.
Takeaway: first-token logprob classification is old hat, but pairing it with a simple heuristic is more reliable than asking small models to output a label — at no latency cost.
More from coding & agent
- Herald OS goes open source: an agent-native OS where the AI is the interface, not an app — gekobraa · 2026-10-07
- MCP observability tool adds per-tool health grades to catch silent agent failures — Thirumalaiboobathi · 2026-10-07
- Hermes Gadget SDK sparks community builds on watches and old phones three days after release — Teknium · 2026-10-07
- Glasser aggregates 1,558 data APIs from 29 providers for pay-as-you-go AI agents — goyalshaliniuk · 2026-10-07
- MIT paper: a minimalist agent loop that passes history as code variables beats Letta and ACE at half the cost — rohanpaul_ai · 2026-10-07
- A 3Blue1Brown-Style Explainer Rendered End-to-End in Rust by franken_manim — doodlestein · 2026-10-07