ByteDance Seed Proposes PISA: Block-Sparse Attention With O(Nlog N) Complexity
ByteDance-Seed · hf · 2026-09-28
ByteDance Seed introduces PISA, a block-sparse attention mechanism that reduces long-context attention cost from quadratic to O(N log N).
- Key idea: Conventional block selection scores all query-block pairs and stays quadratic. PISA uses a pyramid Top-K strategy: pooled keys form O(log N) coarse-to-fine levels, and LogSumExp scoring over bounded candidate sets narrows choices level by level.
- Engineering: Hardware-aware Triton kernels for training and inference fuse hierarchical routing and LogSumExp scoring without materializing the query-key score matrix.
- Results: Comparable performance to baselines on commonsense reasoning benchmarks, with better results on retrieval tasks in language modeling evaluations.
More from Infra
- Developer says local AI is shifting from nice-to-have to infrastructure: control beats privacy — Aiden_Tech_Ai · 2026-09-28
- Meta open-sources Component Benchmark, a hierarchical profiler for TB-scale recommender models — _reachsumit · 2026-09-28
- apple-llm: Node/Python wrapper for the free local LLM built into Apple Silicon Macs — light_2earth · 2026-09-28
- Running 8 watercooled GPUs for local AI: one user's case for watercooling over air cooling — HanchungLee · 2026-09-28
- HF: transformers backend now matches native vLLM speed, no porting needed — ariG23498 · 2026-09-28
- Renting their cluster's compute would have cost over $1 billion on a 5-year deal — ericzelikman · 2026-09-28