Deep Optimization of Automatic1111 for Apple Silicon: 40% Render Time Reduction
Time-Conversation528 · reddit · 2026-08-12
The author shares how they slashed SD1.x render times in Automatic1111 on Apple Silicon (M1 16GB) by about 40% (down to 8.5s for a 512x512 image), without migrating to ComfyUI.
Key Optimizations:
- Metal Flash Attention: Selectively routed for specific SD attention shapes.
- Reduced CPU/GPU Sync: Stopped committing the Metal command buffer after every attention call, letting native kernels work inside PyTorch's MPS stream.
- Unified-memory-aware Attention: Large attention ops dynamically fall back to chunked sub-quadratic attention based on available memory.
- Operator Fusion: Fused GroupNorm + SiLU and GEGLU in Metal.
- FP16 VAE: Enabled FP16 VAE on M1, reducing decode + transfer time from 1.5s to under 1s.
Engineering Takeaways:
The author emphasizes a core principle: "microbenchmarks nominate, full generations elect." Many optimizations that looked promising locally (like packed QKV, K/V caching, or moving a whole ResBlock into MPSGraph) regressed end-to-end performance and were ultimately deleted. Profiling shows 87% of the remaining time is spent on sampling/UNet. The next step is attempting to replay the UNet workload via native Metal/ggml, which won't be integrated unless it delivers a 20%+ speedup.
More from Infra
- AI Compute Boom Drives TL20 Tech Stocks Up 59% Year-to-Date — TiernanRayTech · 2026-08-12
- How Vercel Migrated Its Core Database Handling 6,000 Deployments Per Minute — evilrabbit_ · 2026-08-12
- Analyst: Server CPU Market Bracing for Unprecedented S-Curve Leap — BenBajarin · 2026-08-12
- ComfyUI Workflows Crawl on 128GB DGX Spark vs RTX 4090 — jungseungoh97 · 2026-08-12
- Server CPU Market to Reach $220B by 2030, Potentially Split in Thirds — BenBajarin · 2026-08-12
- HuggingFace Models See ~30% Sustained Drop in Daily Downloads, Likely Due to Filtering Changes — natolambert · 2026-08-12