Nemotron 3.5 Local Test: Runs on 24GB RAM, Lags in Coding but Shines in Tool Calling
curiousily_ · reddit · 2026-08-12
A developer locally tested the Q5 quantized version of NVIDIA's Nemotron 3.5 Lightning (30B-A3B) on an M5 Pro (48GB RAM), sharing specific performance metrics:
- Hardware footprint: Consumes 24GB RAM with a generation speed of 65 tokens/s.
- Coding capability: Code output quality falls well below expectations for its size (which aligns with the authors' disclosed optimization tradeoffs).
- Agentic performance: Demonstrated excellent speed and tool-calling capabilities in Hermes Agent tests.
- Behavioral observation: Tended to be an "overthinker" on certain tasks.
More from coding & agent
- Fine-tuning Muse Glimmer 30B Boosts Click Grounding Accuracy to 41% — mervenoyann · 2026-08-12
- Exploring Telemetry Emission for Governed Dataset Calls via MCP — roller5000 · 2026-08-12
- supastarter Updates Next.js SaaS Boilerplate for AI Coding Agents — jonathan_wilke · 2026-08-12
- Docker Sandboxes Are Reshaping the Agent Permission Model — krishnan · 2026-08-12
- Terraform Called 'Terrorism': Devs Debate Best IaC Tools for Modern Workflows — JasonBotterill · 2026-08-12
- Meta vs NVIDIA 30B Agent Models: Local Execution vs Cloud Routing — eyishazyer · 2026-08-12