Dev benchmarks Bend 2 on M4 Max: parallel kernel 6.4x faster than NumPy
arthurcolle · x · 2026-09-20
A developer shared first-hand benchmarks of a small F32 state-update kernel written with Bend 2 on an M4 Max:
- The parallel CPU version beat NumPy by 2.2x at 2M objects and 6.4x at 8M objects
- A first Metal (GPU) attempt was slower, likely due to suboptimal structuring
- The author is asking for guidance on persistent GPU workloads and irregular spatial trees
More from coding & agent
- Agent memory design: promotion gates, expiry and veto beat 'remember everything' — sujingshen · 2026-09-20
- Early user review: Jev is fast but finds no killer use case over manual coding or Astra/Fable — CtrlAltDwayne · 2026-09-20
- Agent workflow: RFB scripts + tart macOS VM for autonomous UI iteration — craigbalding · 2026-09-20
- cairn-memory open-sourced: agent memory with inspectable source receipts to stop silent rewriting — sujingshen · 2026-09-20
- Jev Bets Agents Don't Need Big LLMs: Claims 20-200x Speed, 40-400x Cost Cuts — curiousperhaps · 2026-09-20
- Ex-Parse cofounder picks Opencode+Muse as favorite cheap non-frontier coding setup — alexandr_wang · 2026-09-20