Watch a Solo Dev Post-Train an 80B Model at Home on V100s: 96 Hours of Distillation, 3340 Samples

jjusko20 · reddit · 2026-09-29

A solo developer is live-streaming a full post-training run: turning AliceAI-Foundation-80B-A3B from base to instruct on home V100s.

Details: they built their own distillation engine (OpenAI-compatible endpoint, works with any teacher) and used Qwen 3.8 27B medium-thinking with a Python sandbox to simulate turn-driven agentic workflows. Generating the dataset — 3340 samples split between 1760 general instruct transcripts and 1580 agentic rows on SWE/harnesses/terminals, covering multiple tool-call syntaxes — took 96 hours across 4 instances at 25tps. Training is a rank-16 QLoRA on q/k/v/o proj only, 2 epochs first, RL planned after. Goal is harness reliability rather than leaderboard scores; dataset may land on HF, and the author is job-hunting in NYC.

Original post →

More from Research

Research channel →