Adaptive VRAM governor keeps 130M model training for 1M steps on an 8GB GPU without OOM
uBazzyZ- · reddit · 2026-09-10
A developer repeatedly hit CUDA OOMs training small models on an RTX 5060 Ti (8GB), so they built MEM Orchestrator, a lightweight runtime control layer for PyTorch: it monitors VRAM pressure and temporarily downshifts micro-batch size and gradient accumulation before an OOM, then steps back up when memory clears.
Implementation details:
- Local policy checks clamp parameter adjustments within safe bounds
- Atomic checkpointing so crashes don't corrupt saved files
- 38 unit tests included
Results: a 130M model survived a 1M-step endurance run with zero crashes, and a 255M model kept training through injected +1.2GB memory spikes during a 50k-step FineWeb-Edu run. The author explicitly doesn't claim this solves OOM and is seeking feedback from CUDA allocator / DeepSpeed / distributed-training experts.
More from Infra
- NVIDIA joins the Rust Foundation — blelbach · 2026-09-10
- Massachusetts Hits Data Centers With New Clean Power Rules, Third State in Three Months — TechCrunch AI · 2026-09-10
- Epoch estimates OpenAI quadrupled compute in both 2024 and 2025, a 17x two-year jump — FlorianGallwitz · 2026-09-10
- Hands-On Guide: Safely Running Untrusted Code with Google Cloud Run Sandboxes — rseroter · 2026-09-10
- Nvidia NVL72 rack shipments seen up 50% in 2027, output forecast to top $710B — Beth_Kindig · 2026-09-10
- Stealth 7-year startup Kepler debuts AI memory beyond HBM, secures up to $245M in US government support — npinto · 2026-09-10