MOLT: local fine-tuning system for consumer GPUs hits ~1005 targets/s on RTX 4060 with Qwen 3B

MKP_Nimilka · reddit · 2026-10-06

A developer released MOLT, a local fine-tuning system for consumer NVIDIA GPUs aiming to make the path from model+dataset to a working local adapter less fragile. It handles dataset detection/validation, pre-run hardware checks (VRAM, RAM, thermals), automatic microbatch fit testing, 4-bit NF4 QLoRA with BF16 adapters, integrity-checked resumable checkpoints, telemetry (VRAM/temperature/energy/throughput), base-vs-adapter evaluation, local adapter chat, and GGUF export workflows.

Under the hood it experiments with custom CUDA/Triton kernels, CUDA graphs, replay modes and fused optimization. On an RTX 4060 Laptop 8GB, a Qwen 3B run hit 1,005 targets/sec, nearly matching his historical Unsloth run (1,007) — though he explicitly notes repeated matched quality/energy/memory tests are still needed. He's recruiting single-GPU fine-tuning testers.

Original post →

More from Models

Models channel →