Post-Trained NVIDIA Nemotron Beats Claude Opus in Legal Agent Tasks
ctnzr · x · 2026-08-12
Harvey, in collaboration with Trajectory Labs, post-trained the NVIDIA Nemotron 3.5 Lightning model on the Legal Agent Bench (LAB), achieving remarkable results:
- Performance Leap: Agent performance on held-out LAB tasks jumped from 0% to 8.3%, outperforming both Claude Opus 4.6 (6.6%) and the significantly larger Nemotron 3 Ultra.
- Broad Improvement: Performance improved across nine practice areas without regressions.
- Cost Efficiency: Post-training reduced average model output from 90k to 37k tokens, boosting the reward-per-token by 2.4x.
Trajectory Labs notes that continual learning relies on cheaper retraining loops. While large models might run this loop every few weeks, smaller models like Nemotron 3.5 Lightning can be fine-tuned nightly, per customer, or even per legal matter, bringing metered intelligence closer to reality.
More from coding & agent
- xAI Launches Grok Bot: Autonomous AI Agents for Real-World Workflows — XFreeze · 2026-08-12
- KohakuTerrarium: Batteries-Included Framework for Multi-Agent Teams — tom_doerr · 2026-08-12
- Build a 3D Brand Logo Material Playground Instantly with Lovable — felixhhaas · 2026-08-12
- Gemini for Go Developers: Model Selection and Agent Development Guide — rseroter · 2026-08-12
- Mojo 1.0 Released: The Systems Language for the AI Era — clattner_llvm · 2026-08-12
- AI Agent Breaks Out of 'Air Force One' Level Sandbox to Book a Flight — sloppenheimer · 2026-08-12