Open-sourced inference acceleration for structure-based models ships benchmarked and documented
AllThingsApx · x · 2026-09-11
Anthonycosta highlights a major release: while the team has accelerated many models and written numerous advanced kernels, finding a general, flexible way to speed up the wide variety of structure-based inference models has been a huge challenge. The release is now benchmarked, documented, and open-sourced, with the team inviting community feedback — calling it just the beginning.
More from Infra
- Commentary: Anthropic loads shift to Google plus AWS slice, OpenAI doubles down on Azure — ericwdolan · 2026-09-11
- Engram section analysis: prime-sized tables, 4-grams, and fp8 lookup tables — stochasticchasm · 2026-09-11
- Nvidia claims Vera Rubin delivers 50X throughput per MW and 35X lower token cost vs Blackwell Ultra — Beth_Kindig · 2026-09-11
- Epoch AI: GPT long-context latency scales quadratically, matching price jumps — Jsevillamol · 2026-09-11
- RTX 3090 mini-bench: ninfer cuts TTFT from 3.4s to 29ms, prompt processing ~76x faster — milkipedia · 2026-09-11
- SpaceX lands another AI compute hosting deal worth $1.1B per month — NicoVerderosa · 2026-09-11