bury-bench: a deterministic, zero-LLM-judge benchmark scoring coding agents on ADHD-friendly answers

Obluness · reddit · 2026-09-11

The author open-sourced bury-bench, a deterministic (zero-LLM-judge) harness that scores coding-agent replies against ADHD-friendly "don't bury the answer" rules and builds a Markdown leaderboard.

Original post →

More from coding & agent

coding & agent channel →