kohya Outlines MiniMax H3 Image Training & Inference Roadmap for musubi-tuner
bdsqlsz · x · 2026-08-11
Kohya, the developer behind the popular musubi-tuner framework, posted a detailed roadmap on GitHub for MiniMax H3 image training and single-frame inference support.
Features already completed and merged into the dev branch include:
- Initial Support: Training and inference for T2VA, FL2VA, and Ref2VA.
- Quantization: ConvRot INT8 quantized base weights, featuring on-the-fly quantization of BF16 checkpoints and automatic detection of ComfyUI pre-quantized artifacts.
- Refactoring: Unification of state dict loading and internal trainer cleanup.
More from Multimodal
- Seedance 2.5 Released: Supports 30s Videos and 50 Multimodal References — azed_ai · 2026-08-11
- Grok's Early Video Generation Model Shows Disastrous Results in Demo — ChromeRoseProtocol · 2026-08-11
- Testing Seedance 2.5: Generate Step-by-Step Object Assembly from Any Prompt — aziz4ai · 2026-08-11
- MiniMax H3 Video Generation Benchmark: RTX 5090 Takes Under 5 Minutes — gabxav · 2026-08-11
- Open-Source Minimax-H3 Hybrid Models Balance Video Quality and Referencing — dampflokfreund · 2026-08-11
- lightx2v Minimax H3 8-step Turbo LoRA Released for ComfyUI — rerri · 2026-08-11