AWS Launches Serverless Fine-Tuning for NVIDIA Nemotron 3

AWS ML Blog · rss · 2026-07-10

Amazon SageMaker AI announced support for serverless customized fine-tuning of the NVIDIA Nemotron 3 model family, initially supporting the Nano (30B) and Super (120B) versions.

Model Features

Nemotron 3 uses a hybrid Mamba-Transformer MoE architecture and supports a context length of up to 1 million tokens. By activating only a small fraction of parameters (e.g., 12B for the Super version), it achieves high throughput and low computing costs, making it perfect for multi-agent workflows.

Fine-Tuning Techniques

Users can adapt the model to specific domains using three techniques without managing underlying GPU infrastructure:

Original post →

More from Models

Models channel →