Finetuned 1.5B Qwen hits GPT-4o-level bash generation with 400k synthetic examples, fully open-sourced

Comfortable-Rock-498 · reddit · 2026-10-05

A hobby project finetuned a 1.5B Qwen model on 400k synthetic examples to generate bash commands at near GPT-4o level, using a mostly automated training pipeline. The author has open-sourced everything: the full synthetic dataset (dirac-run/ec-training-data on Hugging Face), GGUF models at 1.5B and 0.6B sizes, and a companion CLI tool on GitHub, inviting anyone to train on or reuse the data. It shows how far small models can be pushed on vertical tasks with synthetic data.

Original post →

More from Research

Research channel →