KNOWS Benchmark Jointly Evaluates Agents on Search, Tools and Visual Understanding
anmarasovic · x · 2026-10-07
researcher anmarasovic unveiled KNOWS, described as the most ambitious benchmark she has co-created, timed with COLM 2026.
KNOWS jointly and reliably evaluates agents across three axes:
- live web search
- use of common web productivity tools
- visual/spatial understanding
Evaluation is programmatic where possible, with LLM judges where needed. The author is also open to discussing AI safety (especially CoT monitoring), interpretability, user simulation, and agents.
More from coding & agent
- UWaterloo's IGMBench tests world editing in Minecraft and Terraria; best agent solves 78.2% of tasks — UWaterloo · 2026-10-07
- Flipper harness routes 64% of tasks to DeepSeek, cutting API bill by over 60% — toptickcrypto · 2026-10-07
- Overmind: Open Platform That Turns Production Traces Into Fine-Tuning Data for Agents — cneuralnetwork · 2026-10-07
- How to gate MCP write actions: preview+confirm plus hard limits debated — Staff_Sharp · 2026-10-07
- Ampersand launches integration infrastructure powering AI agents in Salesforce, SAP and NetSuite — SimplyAnnisa · 2026-10-07
- Armature launches agent.reviews with 100k+ tool reviews written by AI agents — ycombinator · 2026-10-07