DataSmith agent beats harnesses tuning architecture, using 200x fewer tokens
jefrankle · x · 2026-08-19
Datology AI introduced DataSmith, an auto agent for data research, arguing data quality is the ultimate compute multiplier.- Through data interventions alone, DataSmith outperforms harnesses that also intervene on architecture and optimizers, beating Claude Code by 5.1 pts on average.- Example: with only 734M post-training tokens — 200x fewer than the default recipe — it improved SmolLM3-3B's post-training performance by 3 pp.- Its design leads it to experiment more, reason more, explore more novel datasets, and spawn more subagents, finding counterintuitive solutions — e.g., determining the best way to boost coding performance was adding more on-policy math data, not more code data.
Related event: DatologyAI Launches DataSmith, an Autonomous Data Research Agent(3 posts)→
More from coding & agent
- Agentic coding accessibility will reshape understanding of software complexity — pixlpa · 2026-08-24
- Devin Agent bypasses Slack block by finding emails in git logs — sandylikesfrogs · 2026-08-24
- Developer habits shift: Agents become collaborators from simple tools — latticecut · 2026-08-24
- Dev bottleneck shifts from writing to reading code: exe.dev co-founder — thursdai_pod · 2026-08-24
- The biggest AI mistake: trying to reinvent the wheel instead of using tools — Tired40s · 2026-08-24
- DeepPaperNote turns research papers into Obsidian notes — tom_doerr · 2026-08-24