ApprenticeBench tests AI agents on real construction AP work and continual learning

ysu_nlp · x · 2026-09-11

After four years building construction accounts-payable tooling at Gave, the author joined NeoCognition and pushed for ApprenticeBench: a benchmark for computer use, continual learning, and long-horizon work, starting with construction AP intake.

Unlike tidier existing benchmarks, it mirrors messy real work — mismatched units, wrong cost codes, expired insurance — and asks agents to operate an ERP, study past records, learn company conventions, and process months of invoices with progressively less manager feedback, like a new hire, with no job-specific pre-training.

Related event: NeoCognition Releases ApprenticeBench, First Job-Level Continual Learning Benchmark(8 posts)→

Original post →

More from coding & agent

coding & agent channel →