Running MiniMax H3 on 12GB VRAM: quantization, Turbo LoRAs and attention backends compared
Possible_Mood676 · reddit · 2026-09-11
A Reddit user is setting up MiniMax H3 video generation in ComfyUI on an RTX 4070 (12GB VRAM / 32GB RAM), targeting 5-second 768p I2V/First-to-Last clips. They report that unpruned FP8 models cause heavy paging to system RAM, and are polling the community on: quantization choices (pruned INT8 convrot vs GGUF Q4KM/Q3KM, plus text encoder quants) that preserve faces and audio; whether the 4-step LiteX2V v1.1 Turbo LoRA is still the best speed/quality option versus 8-step variants or other LoRA families; and which attention backend (H3 SLA, SageAttention, Comfy Kitchen) performs best on Ada Lovelace cards. The post is a help request without answers yet, but maps the current decision space for low-VRAM local deployment.
More from Infra
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Your p99 latency benchmark may be lying: a deep dive into coordinated omission — Franc0Fernand0 · 2026-09-11
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11