Failure Map: 20,168 open Python boundary-case bug repair tasks released
failuremap-f · reddit · 2026-10-02
Failure Map is a new archive of compact Python debugging tasks. The open release packs 20,168 tasks across 254 categories, each with standard-library implementations, explicit contracts, failed repair attempts, and executable boundary checks — designed for evaluating coding models' repair ability.
Sample cases: deduplication dropping legitimate events, cache expiry subtracting a whole tick and rejecting a still-valid entry, and pagination using >= instead of > repeating a cursor record. A unified prompt template is provided: repair the solve function to satisfy the contract, return Python source only, preserve the signature.
The author stresses evaluation hygiene: report exact model and revision, quantization, prompt, sampling settings, seed, attempts per task and pass counts; run candidate code in isolation with grading fixtures outside its control. Program baselines (checks passed out of 3) are published, but no local-model results yet.
- Download: https://failuremap.org/api/exports/tasks.l.gz
- Methodology: https://failuremap.org/methodology
More from coding & agent
- OpenAI admits Codex can't guarantee no subagent use, undermining result repeatability — RealSharpNinja · 2026-10-02
- Decagon launches Personal Agent Gateway and PACT protocol for AI-agent customers — kimberlywtan · 2026-10-02
- AI Coding Agents Need Guardrails: Unit Tests Keep Them Sane, But Watch for Test Bloat — randal_olson · 2026-10-02
- Why TDD works with AI coding agents: cheap guardrails and regression signals — randal_olson · 2026-10-02
- Sentry founder: AI agents need deterministic workflows tailored to how your team operates — zeeg · 2026-10-02
- Economist runs full AI research pipeline in 45 minutes for just $19 with Expected Parrot agent — soumitrashukla9 · 2026-10-02