Alibaba’s HSCodeComp benchmark finds top AI agents still far below human tariff experts
jiqizhixin · x · 2026-07-29
Alibaba Group introduces HSCodeComp, a benchmark that asks AI agents to infer precise 10-digit HS codes from product descriptions and tariff rules.
- The task requires layered reasoning from broad product categories down to narrow exceptions, closer to how a lawyer or doctor applies rules.
- Results are harsh: the best AI agent reaches only 46.8% accuracy, while human experts score 95.0%.
- Adding more reasoning steps can actually hurt, with the model drifting off track instead of improving.
- The paper frames HSCodeComp as a realistic, expert-level benchmark for hierarchical rule application.
More from Research
- Small-model orchestration roughly doubled task completion in a 100-task benchmark — _raydeStar · 2026-07-29
- Paper adds a human-only authorship attestation to a quantum matrix result — burny_tech · 2026-07-29
- Biohub is hiring for an AI wet-lab role to build biology models — proteinrosh · 2026-07-29
- An AI digest scans 92 journals every week and turns them into one RSS feed — Afinetheorem · 2026-07-29
- A weekly PDB-synced leaderboard tracks open cofolding models — rishabh16_ · 2026-07-29
- Agentic AI Summit sets robotics and world models session with Sergey Levine and Jim Fan — dawnsongtweets · 2026-07-29