2-bit Quantized Nemotron 3.5 Runs Autonomous Tool Calls Continuously on Just 22GB VRAM

danielhanchen · x · 2026-08-13

UnslothAI has applied 2-bit quantization to the NVIDIA Nemotron 3.5 Lightning model, enabling it to run smoothly on devices with only 22GB of VRAM.

In real-world testing, the quantized model continuously executed tool calls for 10 minutes straight, citing over 80 websites, executing code, and searching for 10 real-world locations. This significantly lowers the hardware barrier for deploying advanced agentic workflows at the edge.

Original post →

More from coding & agent

coding & agent channel →