NUS releases SafeActBench: 656 cases reveal where tool-using agents break the evidence-to-action chain
NationalUniversityofSingapore · hf · 2026-10-07
NUS researchers introduce SafeActBench, a benchmark studying where tool-using agents break the chain from evidence to action.
Key findings
- Across 10 model-harness configurations, strong static action assessment can coexist with much weaker interactive execution.
- Failures often begin before execution: agents stop with incomplete investigation or act before required evidence is established.
- Once evidence is obtained, single-action execution is usually reliable, while multi-action workflows additionally expose unresolved prerequisites and incomplete execution.
Benchmark design
SafeActBench comprises 656 cases across six operational domains and five protocols, progressing from static action judgment and investigated non-action to single- and multi-action workflows. A provenance-bound Evidence Ledger and deterministic trajectory evaluator track what information was established, when actions occurred, and whether downstream dependencies were satisfied.
Takeaway: agent failures arise not only from missing information, but also from how agents use established evidence when deciding and executing actions.
More from coding & agent
- Dev Builds 'Slop Cannon': Agent-Orchestrated h3 + fal + Gemini Content Machine — tobowers · 2026-10-07
- Microsoft's MiniCorp Simulates an AI-Run Company to Generate Enterprise Agent Data — microsoft · 2026-10-07
- Paper: Claude Code, Codex and 3 other coding agents let models tamper with their own traces — maksym_andr · 2026-10-07
- AI Red-Teamer Tested on 5 OpenRouter Setups: Cheapest Full Scan Cost $0.14 — Humanbound_AI · 2026-10-07
- Build a sub-1s AI voice agent portfolio project to win clients via cold email — ashishllm · 2026-10-07
- Monid raises $7.7M to become the OpenRouter for agent tools with per-call pricing — aliscodes · 2026-10-07