AFFMAE brings 5x faster high-res vision pretraining to desktop hardware, at ECCV 2026
CSProfKGD · x · 2026-09-07
Researchers present AFFMAE at ECCV 2026: a masking-friendly hierarchical pretraining framework built on AutoFocusFormer's adaptive, off-grid token merging, aimed at high-resolution microscopy segmentation.
Key points:
- David Smerkous heavily optimized AutoFocusFormer code, yielding 5x inference speedup at 1024x1024 and 50% lower memory cost;
- The paper introduces numerically stable mixed-precision Triton kernels and a lightweight point-based decoder reusable as a segmentation head;
- At equal parameter counts, AFFMAE matches MAE fine-tuning on foot process width estimation with a ViT backbone, while pretraining 2x faster with halved peak memory, and up to 5x fine-tuning throughput at 1024px — enabling high-res pretraining on desktop hardware;
- Paper and code are open-sourced; presented at ECCV Poster Session 3, #74.
More from Infra
- The Myth of Self-Hosted AI: 'Local' Models Still Route Through the Cloud — nomad-nostalgia · 2026-09-07
- Homelab With 4x RTX 4090 Weighs vLLM+P2P Patch vs llama.cpp for Qwen Models — dowitex · 2026-09-07
- Could AI run entirely on your phone? It could upend OpenAI's pricing — kevinsurace · 2026-09-07
- Inference engineering is the underrated AI skill: KV cache, batching and p99 latency explained — techNmak · 2026-09-07
- Polymarket puts 73% odds a US state enacts data center moratorium by end of 2026 — Polymarket · 2026-09-07
- Louisiana taco shop says 40% of monthly business comes from Meta's AI data center — Polymarket · 2026-09-07