17M-parameter model beats frontier LLMs 80% of the time after cheap synthetic-data finetuning

max_paperclips · x · 2026-09-27

MaximeRivest shares finetuning experiments: a 17M-parameter Ettin model finetuned on LLM-generated synthetic data beats Jev 80% of the time, and beats Kimi-k3 40% of the time on human labels. Distilling from Kimi-k3 synthetic data gets 70% close to Jev.

Key numbers:

Takeaway: with 500k+ inputs to classify, finetuning a small model on synthetic data is both faster and cheaper than general LLMs.

Original post →

More from Models

Models channel →