Data-Efficient Agent Distillation: Small Model Matches 9x Larger Teacher with Only 19 Samples

ctnzr · x · 2026-08-06

Agent distillation can be surprisingly data-efficient. By using trajectory data from only 19 problems, researchers distilled Inkling-Small into Nemotron-3-Nano-30B-A3B via LoRA SFT.

The small model successfully matched the performance of its teacher model—which is nine times its size—on Kubernetes incident-diagnosis tasks. This is highly encouraging for post-training task-specific models in data-poor regimes.

Original post →

More from coding & agent

coding & agent channel →