Developer Uses Codex MCP to Audit MBPP and HumanEval, Finds Deep Issues
code_star · x · 2026-08-22
X user codestar shares his experience using Codex through an MCP server to audit benchmarks. He created 47 concurrent nac sessions and within 2 hours audited every sample of MBPP and HumanEval variants, correcting deep issues in prompts, test cases, formatting, and even upstream MBPP. He reminds: 'Remember, folks, always look at the data!'
Related event: Developer Uses Codex MCP to Audit HumanEval and MBPP Benchmarks(2 posts)→
More from coding & agent
- Arcee open-sources nac agent harness for long-running tasks with intent alignment — latkins · 2026-08-22
- The case for building agent-friendly products — omooretweets · 2026-08-22
- Community fork of DeepSeek Harness ships a signed, notarized desktop app — jasonkneen · 2026-08-22
- Replit now lets you create, inspect and publish projects from ChatGPT, Claude or any MCP client — amasad · 2026-08-22
- Team Ships 30 Apps in 30 Days Using Reusable SDK — PrajwalTomar_ · 2026-08-22
- Opinion: AI Will Smooth Adoption Hurdles, Making Rust a Winner — ethanniser · 2026-08-22