Supabase open-sources Evals for real-world AI agent benchmarking
大模型之路 · wechat · 2026-08-17
Existing coding benchmarks often rely on synthetic problems, missing real-world backend risks like security misconfigurations. Supabase has open-sourced Evals, a tool that transforms real support tickets and GitHub issues into evaluation tasks involving schema design, RLS policy fixes, and migrations.
It combines deterministic checks (e.g., access validation) with LLM-as-a-judge scoring. Early results suggest tool usage capability is more critical than base model size, revealing significant behavioral differences in documentation reading and coding patterns among agents.
More from coding & agent
- Open-Source Dashboard Claude Code Karma: Visualize Local Sessions, Timelines, Costs, and Live Activity — tom_doerr · 2026-08-17
- Arguing with your AI Agent? Check if your context window is over 50% usage. — AiJohnAllen · 2026-08-17
- Hermes-Agent launches 900K context window for subscribers — alexcovo_eth · 2026-08-17
- Hermes: The Highest-Leverage AI Agent Setup with 900K Context and Weekly Updates — EXM7777 · 2026-08-17
- Benchwarmer: Tool rehabilitates misleading benchmark charts — MeganRisdal · 2026-08-17
- Agents read context to fix code, ignoring bad naming — dotey · 2026-08-17