NVIDIA's BioNeMo Inference Runtime hits public beta, boosting Boltz-2 folding throughput 2.9x
AllThingsApx · x · 2026-09-10
NVIDIA released BioNeMo Inference Runtime in public beta: an open, PyTorch-native library that accelerates biomolecular structure-prediction inference on NVIDIA GPUs.
Key points:
- End-to-end pipeline from input parsing, tokenization and feature generation through GPU inference to PDB/mmCIF output, plus direct PyTorch module integration.
- Optimization at three layers: kernel selection, CUDA Graph capture for module optimization, and Ray-powered GPU replicas for pipeline scaling.
- In a matched benchmark of 1,000 human dimer targets on 8xH100 GPUs, accelerated Boltz-2 delivered 58.5K successfully folded residues per GPU-hour vs 20.2K for a torch-compiled open-source implementation — a 2.90x residue-normalized throughput gain.
- Estimated energy to fold one million comparable targets drops from 35 MWh to 11 MWh.
- Supports Boltz-2, OpenFold2, and Protenix v2, with partners including apheris, SandboxAQ, and Xaira.
More from Infra
- DeepSeek Ships V4.1-Flash: 552B MoE With 8B Active, 75% Less KV Cache Memory — mark_k · 2026-09-11
- d-Matrix adopts NVIDIA NVLink Fusion for rack-scale Raptor XPU inference — BenBajarin · 2026-09-11
- AI chip startup dMatrix bets on compute only, partnering with Nvidia for the rest — BenBajarin · 2026-09-11
- DeepSeek v4.1 Flash is twice as big but not twice as smart; local users better off with v4 Flash — QuixiAI · 2026-09-10
- Static Two-GPU Split Fixes Deterministic Crashes in 34GB Video DiT — And It's Faster — Responsible_Art7438 · 2026-09-10
- Podcast: China's agent rules and Baseten's latest — The Cognitive Revolution · 2026-09-10