Redditor post-trains an 80B model on three V100s with SFT + GRPO
jjusko20 · reddit · 2026-09-25
A Redditor shares an in-progress project post-training the open AliceAI-Foundation-80B-A3B-Base model:
- Hardware: Three 32GB V100s, barely fitting 80B training via QLoRA
- Method: A shallow distill of Qwen 3.8 27B generates synthetic data targeting long-horizon agentic work for SFT, followed by RL/GRPO with a grader model
- Tradeoff: Considered a stronger teacher model but insists on keeping everything local
The SFT data-generation framework is built; no final results yet — a community log of post-training a frontier open model on old GPUs.
More from coding & agent
- Researcher Praises TerminalBench Verifiers, Says Eval Design Has Grown Far More Involved in a Year — AnkaReuel · 2026-09-25
- Pika API Lets Grok Bots Generate Images, Video and Audio via 120+ Models on One Key — Kyrannio · 2026-09-25
- AEO tools for coding agents: new product targeting Claude Code in terminal — jia_seed · 2026-09-25
- Autolaunch lets AI agents raise early capital via CCA auctions on Base — seanwbren · 2026-09-25
- How GitHub Rebuilt the Copilot App to Render a Million-Line Pull Request — mariorod1 · 2026-09-25
- Claude Opus 5.5 builds Minecraft from one prompt in ~1 hour for ~$20 — amasad · 2026-09-25