NVIDIA BioNeMo kernels hit 7.1x speedup on triangle attention in Terray's drug discovery benchmarks
AllThingsApx · x · 2026-09-11
Terray Therapeutics benchmarked NVIDIA's newly released BioNeMo Inference Runtime kernels (CuTeDSL) on its production architecture for TerraBind, its binding affinity and structure prediction model.
Key results:
- TerraBind's dominant cost is the O(n³) triangle attention operation central to modern structure-based drug discovery
- On H100 GPUs, the new kernels are up to 7.1x faster than standard PyTorch attention and up to 1.8x faster than Terray's current optimized cuEquivariance kernel stack
- The advantage grows at larger crop sizes, which matter most for protein–ligand complexes
- Integration was straightforward: kernels running and benchmarked within a day of access, module-level swaps only, no pipeline rewrites required
More from Infra
- Dev shares training dashboard: ~$11/b tokens cost with 'insane' MFU — jon_durbin · 2026-09-11
- Positron AI Raises $875M Series C at $5B Valuation, Deploying 50+ Atlas Racks at Oracle Cloud — Scobleizer · 2026-09-11
- Positron AI raises $230M Series B at over $1B valuation with Arm backing — seanmcdonaldxyz · 2026-09-11
- Cerebras Fast Inference Flips Agent Workflows: Fewer Parallel Agents, Same Output — MatthewBerman · 2026-09-11
- Baseten acquires Blaxel to build integrated cloud infrastructure for AI agents — baseten · 2026-09-11
- Skild AI's S1 learns robot tasks from one video, hits $100M revenue run rate — NVIDIA Blog · 2026-09-11