Protein embedding still needs 1TB of memory and 384 CPU nodes for a single run

ebetica · x · 2026-07-22

A complaint about a protein-embedding workflow that still needs enormous resources: 1 TB of memory and 384 CPU nodes just to embed a single protein in about 10 minutes.

The point is less about one specific model than the operational absurdity of the current setup: even a single inference job can require a very large cluster footprint, which is a useful signal about the compute cost of large-scale bio/ML pipelines.

Original post →

More from Infra

Infra channel →