DataSmith agent beats harnesses tuning architecture, using 200x fewer tokens

jefrankle · x · 2026-08-19

Datology AI introduced DataSmith, an auto agent for data research, arguing data quality is the ultimate compute multiplier.- Through data interventions alone, DataSmith outperforms harnesses that also intervene on architecture and optimizers, beating Claude Code by 5.1 pts on average.- Example: with only 734M post-training tokens — 200x fewer than the default recipe — it improved SmolLM3-3B's post-training performance by 3 pp.- Its design leads it to experiment more, reason more, explore more novel datasets, and spawn more subagents, finding counterintuitive solutions — e.g., determining the best way to boost coding performance was adding more on-policy math data, not more code data.

Related event: DatologyAI Launches DataSmith, an Autonomous Data Research Agent(3 posts)→

Original post →

More from coding & agent

coding & agent channel →