MOLT: local fine-tuning system for consumer GPUs hits ~1005 targets/s on RTX 4060 with Qwen 3B
MKP_Nimilka · reddit · 2026-10-06
A developer released MOLT, a local fine-tuning system for consumer NVIDIA GPUs aiming to make the path from model+dataset to a working local adapter less fragile. It handles dataset detection/validation, pre-run hardware checks (VRAM, RAM, thermals), automatic microbatch fit testing, 4-bit NF4 QLoRA with BF16 adapters, integrity-checked resumable checkpoints, telemetry (VRAM/temperature/energy/throughput), base-vs-adapter evaluation, local adapter chat, and GGUF export workflows.
Under the hood it experiments with custom CUDA/Triton kernels, CUDA graphs, replay modes and fused optimization. On an RTX 4060 Laptop 8GB, a Qwen 3B run hit 1,005 targets/sec, nearly matching his historical Unsloth run (1,007) — though he explicitly notes repeated matched quality/energy/memory tests are still needed. He's recruiting single-GPU fine-tuning testers.
More from Models
- Zuckerberg says Llama 4 failed because it was staffed like Instagram, not a frontier lab — rohanpaul_ai · 2026-10-06
- Controlled Study Finds No Encoding Dominates: Pixels, Bytes and Tokens Each Win on Different Tasks — delliott · 2026-10-06
- Old 'gpt-next' codename resurfaces; researcher speculates an intentional OpenAI easter egg — RileyRalmuto · 2026-10-06
- Claude leaves 50% of plan quota after burning one model's cap; OpenAI goes to zero — xiaohu · 2026-10-06
- Sander Dieleman on why continuous diffusion language models are making a comeback — LucaAmb · 2026-10-06
- Reflection AI launches Beam, a 501B-A23B model scoring between GLM-5.2 and GLM-5.3 — iaziaz · 2026-10-06