KNOWS Benchmark Jointly Evaluates Agents on Search, Tools and Visual Understanding

anmarasovic · x · 2026-10-07

researcher anmarasovic unveiled KNOWS, described as the most ambitious benchmark she has co-created, timed with COLM 2026.

KNOWS jointly and reliably evaluates agents across three axes:

Evaluation is programmatic where possible, with LLM judges where needed. The author is also open to discussing AI safety (especially CoT monitoring), interpretability, user simulation, and agents.

Original post →

More from coding & agent

coding & agent channel →