Palantir says Nvidia’s Nemotron Ultra beat frontier models on real customer tasks
eliano · x · 2026-08-04
Palantir’s Shyam Sankar says Nemotron Ultra looked “better than frontier” on their workloads after only 24 hours of setup and no post-training.
- He says the model’s benchmark scores looked far from frontier-level, but that those benchmarks were not aligned with his customers’ real tasks.
- The takeaway is that benchmark results can understate usefulness when a model is evaluated against the wrong workload.
More from Models
- xAI updates Grok Build with Grok 4.5, skills, MCP, and plan mode — elonmusk · 2026-08-04
- RL on custom search harnesses may beat the “one big model” idea — shangbinfeng · 2026-08-04
- Qwen 3.8 Max reaches 42% on the hard INDUCTION benchmark, taking second place — DeryaTR_ · 2026-08-04
- OpenAI hires the creator of WebRTC as its GPT-Live voice system gets a deep dive — bookwormengr · 2026-08-04
- Code Arena WebDev puts four open-weight Chinese models near the frontier — floriandotorg · 2026-08-04
- GPT-5.6 Misspells Email Address, Then Hallucinates a Post-Hoc Excuse — WolframRvnwlf · 2026-08-04