One attention node makes local MiniMax-H3 video generation 2.5x faster on AMD Radeon 8060S iGPU

ShamanFlamingoFR · reddit · 2026-09-21

Three weeks after running MiniMax-H3 video generation locally on a Ryzen AI Max+ 395 (Radeon 8060S iGPU, 128GB unified memory), the author reports per-evaluation time dropping from 45.7s to 18.3s and end-to-end cold renders from 521s to 206s — same machine, same graph, same model. The lever: ComfyUI's ModelAttentionBackend("comfy kitchen attention") node finally has a working kernel.

Key details

Revised conclusion: a 2.5x speedup from just swapping attention at unchanged clocks means the earlier "bandwidth-bound" verdict was wrong — the bottleneck was split attention's intermediate VRAM round-trips, which the fused INT8 kernel eliminates.

Original post →

More from Fun

Fun channel →