Palantir says Nvidia’s Nemotron Ultra beat frontier models on real customer tasks
eliano · x · 2026-08-04
Palantir’s Shyam Sankar says Nemotron Ultra looked “better than frontier” on their workloads after only 24 hours of setup and no post-training.
- He says the model’s benchmark scores looked far from frontier-level, but that those benchmarks were not aligned with his customers’ real tasks.
- The takeaway is that benchmark results can understate usefulness when a model is evaluated against the wrong workload.
Related event: Palantir Says NVIDIA Model Beats Frontier Models in 24 Hours(3 posts)→
More from Models
- TypeSafe AI's Jev introduces 'decision models': text in, probabilistic scores out, at $0.042/M tokens — teropa · 2026-09-22
- Challenge: Track Your Daily Token Usage to Prove AI Companies Are Throttling Limits — tomchapin · 2026-09-22
- Day 1 with Grok 4.7: strict system-prompt adherence and visible gains over 4.5 in real coding work — elonmusk · 2026-09-22
- MoVA adds sparse value experts to attention for more capacity at no extra KV-cache cost — rupspace · 2026-09-22
- ChatGPT Plus power user logs 22 mid-task stalls in a week, up from 1 a month ago — ItsSteve-O · 2026-09-22
- Ex-OpenAI safety lead says Grok 4.7 looks a bit better on some safety fronts — Miles_Brundage · 2026-09-22