Snorkel AI commits $3M to Open Benchmarks Grants funding agentic AI evaluations
typewriters · x · 2026-09-25
Snorkel AI's Open Benchmarks Grants program commits $3M to fund open-source datasets, benchmarks, and evaluation artifacts, with rolling acceptances.
- Echoes Logan Graham's take that AI product companies should spend >25% of their time building benchmarks
- Funded examples include Terminal-Bench 4.0, Terminal-Bench-Science, Senior SWE-bench (senior-level engineering work), and Agent's Last Exam
- Thesis: great benchmarks need a strong capability thesis plus methodological rigor (expert and programmatic QA)
- Benchmarks must close the gap along three dimensions: environment complexity, autonomy horizon, and output complexity
More from coding & agent
- Clinician-Reviewed Child Psychiatry MCP Library Launches on Glama — modelcontextprotocol · 2026-09-25
- Brainiall MCP Server Brings Phoneme-Level Speech Assessment to Agents — modelcontextprotocol · 2026-09-25
- Plan Mode rework: customizable prompt, shareable modes, rebindable Shift+Tab — trq212 · 2026-09-25
- Firecrawl raises $75M Series B, targets $100M ARR with fewer than 50 people — devdigest · 2026-09-25
- Hill Sampling: Condition on the Best Verified Solution, Generate 512 Parallel Edits, Repeat — LChoshen · 2026-09-25
- Free Public MCP Server Exposes Canadian Privacy Law Data: 263 Terms, CAI Cases — masiha97 · 2026-09-25