bury-bench: a deterministic, zero-LLM-judge benchmark scoring coding agents on ADHD-friendly answers
Obluness · reddit · 2026-09-11
The author open-sourced bury-bench, a deterministic (zero-LLM-judge) harness that scores coding-agent replies against ADHD-friendly "don't bury the answer" rules and builds a Markdown leaderboard.
- Motivation: the actionable bit keeps getting buried under "Great question!", hedging, and "Hope this helps!" — inspired by the viral i-have-adhd skill
- Uses regex, structure, and heuristics for reproducible scoring: same input → same score, no model grading another model
- Demo (no API keys): good chat scores 76.9, buried chat 7.9
- Try: bury-bench ingest samples/chat.l vs samples/chat-buried.l; critique welcome on proxies like hedging density
More from coding & agent
- Spicy take making the rounds: "You need better engineers to work with LLMs" — andreisavu · 2026-09-11
- Two AI agents completed a substantial payment on their own in under five minutes — mattshumer_ · 2026-09-11
- Websites can hand tools to AI agents: WebMCP demo schedules posts from the browser — haltakov · 2026-09-11
- LangChain ships Managed Deep Agents with containerized evals via LangSmith — LangChain · 2026-09-11
- Slash Astra quota by 86%: run cheap Luna as orchestrator, Astra as subagent — TheMoonMidas · 2026-09-11
- Dart team ships Skills CLI 1.0 to bundle AI agent skills with packages — rseroter · 2026-09-11