MiniCPM5-2B Scores 100 vs Spark-X2.5-4B's 84 on a Real Agent Task

Equivalent-Grass-527 · reddit · 2026-09-16

A hands-on comparison tested two small LLMs on a real customer-service agent task with database-backed validation: look up an order, check policy, find inventory, create a replacement, and schedule pickup. The catch: the nearest Shanghai warehouse was out of stock (Suzhou, 95km away, had 12 units), and 'tomorrow afternoon' had to resolve to Sept 10 given the system date.

The author argues the difference lies in the 'last mile of agency' — finishing the workflow vs pushing decisions back to the human — something standard benchmarks miss, and proposes end-to-end task-completion/database-state evals as more meaningful.

Original post →

More from coding & agent

coding & agent channel →