Deep Optimization of Automatic1111 for Apple Silicon: 40% Render Time Reduction

Time-Conversation528 · reddit · 2026-08-12

The author shares how they slashed SD1.x render times in Automatic1111 on Apple Silicon (M1 16GB) by about 40% (down to 8.5s for a 512x512 image), without migrating to ComfyUI.

Key Optimizations:

Engineering Takeaways:

The author emphasizes a core principle: "microbenchmarks nominate, full generations elect." Many optimizations that looked promising locally (like packed QKV, K/V caching, or moving a whole ResBlock into MPSGraph) regressed end-to-end performance and were ultimately deleted. Profiling shows 87% of the remaining time is spent on sampling/UNet. The next step is attempting to replay the UNet workload via native Metal/ggml, which won't be integrated unless it delivers a 20%+ speedup.

Original post →

More from Infra

Infra channel →