Watch a Solo Dev Post-Train an 80B Model at Home on V100s: 96 Hours of Distillation, 3340 Samples
jjusko20 · reddit · 2026-09-29
A solo developer is live-streaming a full post-training run: turning AliceAI-Foundation-80B-A3B from base to instruct on home V100s.
Details: they built their own distillation engine (OpenAI-compatible endpoint, works with any teacher) and used Qwen 3.8 27B medium-thinking with a Python sandbox to simulate turn-driven agentic workflows. Generating the dataset — 3340 samples split between 1760 general instruct transcripts and 1580 agentic rows on SWE/harnesses/terminals, covering multiple tool-call syntaxes — took 96 hours across 4 instances at 25tps. Training is a rank-16 QLoRA on q/k/v/o proj only, 2 epochs first, RL planned after. Goal is harness reliability rather than leaderboard scores; dataset may land on HF, and the author is job-hunting in NYC.
More from Research
- Hexagon, a New Repository for Math and TCS Research, Enters Public Beta — suchenzang · 2026-09-29
- JHU's Harang Ju to present HCOMP talk on the jagged frontier in human-AI collaboration — windx0303 · 2026-09-29
- LayerSkip: Self-Speculative Decoding Speeds Up LLMs Without a Draft Model — burkov · 2026-09-29
- PolyXOR takes 128-bit hash speed crown; new work shows adversarial collisions in fast hashes — thomasahle · 2026-09-29
- Bio-data hackathon with Anthropic and AWS: find new biological insights in 48 hours — adityaag · 2026-09-29
- Hindsight: Letting Your Agent Learn Without Breaking Policy — nishithreddy · 2026-09-29