RegLLM: a diagnostic harness measures bounded autonomy in regulated agentic AI
Dipankar Sarkar · hf · 2026-10-01
Dipankar Sarkar released RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows.
- Instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action rate, supervised by programmatic verifiers, task-level escalation labels, or AI judges
- A deterministic runtime supervisor blocks ungrounded answers and forces escalation; the same domain constitution drives evaluation, training rewards, and serving guardrails
- A smoke-scale run (n=12) lifts escalation recall from 0 to 0.67 and cuts unsafe-action rate from 0.33 to 0.08 with governance enabled
- Two single-GPU Qwen2.5-3B LoRA/DPO pilots reveal substantial config variance (task success 0.25 vs 0.12 for nominally identical setups), showing configuration noise can overwhelm tuning effects
The author explicitly notes these pilots don't establish production readiness; the contribution is diagnostic.
More from coding & agent
- Developer recreates Project Hail Mary's stellar map in JavaScript with AI — mhmazur · 2026-10-01
- CMU releases CUA-SWE, a benchmark uniting computer-use agents with visual software engineering — CarnegieMellonU · 2026-10-01
- Médula open-sources data on coordinating parallel Claude Code agents: Jev decider escalates 61% of real writes — jokiruiz · 2026-10-01
- Three agents worked the same account blind — one team's fix with a shared memory layer — Davnys · 2026-10-01
- Skip detailed feedback: coding agents fix their work with just "this sucks, improve it" — bosmeny · 2026-10-01
- Physics YouTuber shows Claude fully automating his explainer videos, 3.5 years after asking when AI could — AndyMasley · 2026-10-01