Nemo corpus NER case study: all models misclassify an initiative as an ORG with 90+ confidence
HankYeomans · x · 2026-09-18
- The author ran entity-classification passes on an ORG example from nvidia's Nemo corpus, comparing PII models, Jev, and other models by probability and confidence scores.
- The hard part is separating an organization from a team within it, e.g. "The Marriott Customer Services Team."
- Every model, including Jev, scored "Network Services Initiative" above 90 as an ORG, though it's an initiative of an organization, not one itself.
- The author suggests RL-like approaches are needed to close this gap, noting even human judgment is open to interpretation.
More from Research
- Epoch launches Benchmark Reviews, auditing 15 AI benchmarks: only 4 verified, 9 flawed — stochasticchasm · 2026-09-18
- RL agents invent their own diagnostic renderings to ground code understanding, sparking RL scaling optimism — teortaxesTex · 2026-09-18
- Grounded SI: egocentric data at scale hinges on hand tracking in wild, long-tail scenarios — micoolcho · 2026-09-18
- GenBio AI co-founders publish 'A world model of the virtual cell' in Cell — HongyiWang10 · 2026-09-18
- Weeks after Navier-Stokes, Hodge Conjecture reportedly cracked: Millennium Problems falling fast — haider1 · 2026-09-18
- Meaning Spark Labs experiments with inference-time metacognitive scaffolding for LLMs — PeterBowdenLive · 2026-09-18