NUS releases SafeActBench: 656 cases reveal where tool-using agents break the evidence-to-action chain

NationalUniversityofSingapore · hf · 2026-10-07

NUS researchers introduce SafeActBench, a benchmark studying where tool-using agents break the chain from evidence to action.

Key findings

Benchmark design

SafeActBench comprises 656 cases across six operational domains and five protocols, progressing from static action judgment and investigated non-action to single- and multi-action workflows. A provenance-bound Evidence Ledger and deterministic trajectory evaluator track what information was established, when actions occurred, and whether downstream dependencies were satisfied.

Takeaway: agent failures arise not only from missing information, but also from how agents use established evidence when deciding and executing actions.

Original post →

More from coding & agent

coding & agent channel →