SOL Attention + SageAttention Speeds Up H3 Video Gen Up to 2.5x on Apple Silicon
TgoAI · reddit · 2026-09-10
A developer added SOL Attention (dynamic block-sparse attention) and a SageAttention-style INT8 QK path to Vpipe on Apple Silicon. On an M5 Pro 24GB, 6-step DiT: at 832×480, vanilla H3 takes 17 min vs SOL's 9.3 min; at 1344×768 both accelerated paths hit 30 min vs 74 min vanilla (2.5x). Unlike VDN's projection+linear-attention replacement, SOL keeps exact attention for selected blocks and produced no artifacts (e.g. reversed water flow) seen with VDN. The author reimplemented SOL kernels natively in Metal rather than reusing MPS, and the shared backend benefits other image/video models. Open source: github.com/tgo-app-dev/vpipe.
Related event: SOL Attention plus SageAttention speeds Apple Silicon inference 2.5x(2 posts)→
More from Infra
- Epoch: Top AI firms' compute grows 4x/year; OpenAI up nearly 20x since 2023 — Jsevillamol · 2026-09-10
- Qwen3.8 Flash on 128GB Strix Halo: memory math and perplexity of Q4+Q8 n-gram hybrid quants — MarkoMarjamaa · 2026-09-10
- Agent auto-builds a custom Dockerfile to optimize new workspace start times — lucasmeijer · 2026-09-10
- MEM Orchestrator: adaptive memory control plane for LLM training on an 8GB GPU — uBazzyZ- · 2026-09-10
- OpenAI compute has grown ~20x since 2023; leading labs scale 4x/year — The Verge AI · 2026-09-10
- Keras ships ZeroModels: 100+ model families in pure Keras 3, runnable on any backend — fchollet · 2026-09-10