Knowledge-gated verifiable tasks isolate whether LLM agents lack knowledge or competence

Hanlin Tian · hf · 2026-09-03

A new protocol separates task instructions from private convention artefacts to explicitly test LLM agents' dependence on hidden knowledge, disentangling "ignorance" from "incompetence".

Calibration tasks validate the design: agent performance drops to near zero without access to the hidden knowledge, showing the protocol cleanly isolates the knowledge factor. Useful for agent evaluation design.

Original post →

More from coding & agent

coding & agent channel →