ICML Coverage: Tackling Tough Biology and Chemistry Problems
VectorInst · x · 2026-07-10
This is the final piece in the series covering ICML 2026 papers, focusing on "applying machine learning to the hardest problems in biology and chemistry."
Genomics
- dnaHNet is a genomics foundation model that learns to dynamically segment DNA rather than using a fixed vocabulary.
- The authors claim it outperforms leading models in efficiency and is more accurate in predicting protein mutation effects and gene essentiality.
- Two other genomics papers discuss evaluation biases and evolutionary pre-training for genome language models.
Quantum Chemistry
Alán Aspuru-Guzik's team reduces quantization costs from three angles:
- MoLe, which learns coupled cluster outputs using cheap inputs;
- Reducing energy errors by 66% through derivative-informed training;
- Achieving up to a 600× speedup when predicting crystal electron density with ELECTRAFI.
Overall, this collection of work emphasizes that in biological and chemical tasks, model design, training methods, and evaluation biases all significantly impact performance and computational costs.
More from Infra
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22
- A Firecracker-based platform says it can host 6,000 AI agents on one 256 GB server — maritime_sh · 2026-07-22