Running MiniMax H3 on 12GB VRAM: quantization, Turbo LoRAs and attention backends compared

Possible_Mood676 · reddit · 2026-09-11

A Reddit user is setting up MiniMax H3 video generation in ComfyUI on an RTX 4070 (12GB VRAM / 32GB RAM), targeting 5-second 768p I2V/First-to-Last clips. They report that unpruned FP8 models cause heavy paging to system RAM, and are polling the community on: quantization choices (pruned INT8 convrot vs GGUF Q4KM/Q3KM, plus text encoder quants) that preserve faces and audio; whether the 4-step LiteX2V v1.1 Turbo LoRA is still the best speed/quality option versus 8-step variants or other LoRA families; and which attention backend (H3 SLA, SageAttention, Comfy Kitchen) performs best on Ada Lovelace cards. The post is a help request without answers yet, but maps the current decision space for low-VRAM local deployment.

Original post →

More from Infra

Infra channel →