ApprenticeBench tests AI agents on real construction AP work and continual learning
ysu_nlp · x · 2026-09-11
After four years building construction accounts-payable tooling at Gave, the author joined NeoCognition and pushed for ApprenticeBench: a benchmark for computer use, continual learning, and long-horizon work, starting with construction AP intake.
Unlike tidier existing benchmarks, it mirrors messy real work — mismatched units, wrong cost codes, expired insurance — and asks agents to operate an ERP, study past records, learn company conventions, and process months of invoices with progressively less manager feedback, like a new hire, with no job-specific pre-training.
More from coding & agent
- Dev argues MCP is unnecessary: hand agents an API key and they'll get the job done — sull · 2026-09-11
- Dev to add AI-friendly collision checking CLI to OpenSCAD — _Stocko_ · 2026-09-11
- Cadenya: a year-long indie build of an MCP-powered agent runtime — littlebobbyt · 2026-09-11
- xAI team to livestream building a company from scratch with Grok Bot over three days — soleio · 2026-09-11
- ClawOS: an experimental open-source Arch-based OS with OpenClaw as its primary agent interface — heyneighbor · 2026-09-11
- Grok Bot adds Salesforce, HubSpot, Gong connectors plus installable sales-bot templates — XFreeze · 2026-09-11