Hydra replaces attention: Cerberus lifts eQTL AUPRC from 0.668 to 0.692

2026-10-10

Calico replaced Borzoi attention with Hydra blocks. Eight-model Cerberus reads 786 kb and lifts GTEx eQTL AUPRC from 0.668 to 0.692, and effect-size Spearman ρ from 0.321 to 0.379.

What problem this solves

Borzoi and AlphaGenome predict expression, splicing, and polyadenylation from DNA, and they mix long-range sequence with self-attention. Attention scales with the square of length, so these models pool onto coarse bins, mix there, and spend a U-net decoder getting fine coverage back. Training is expensive enough that published settings are whatever search fit the budget. Inference is slow enough to cap how many variants get scored.

The reference genome does not grow. A new assay adds a track, not a new sequence. With sequence diversity capped, the useful change is a long-range block that learns more from the same chromosomes and costs less to run.

Method

Cerberus replaces the transformer tower in Borzoi with Hydra, a bidirectional state-space block built on Mamba-2. A state-space model treats the sequence as a recurrence. A hidden state carries what earlier positions wrote, and the cost grows linearly with length. Mamba lets that recurrence depend on the input, so the model writes harder at positions worth keeping. A one-way scan only sees upstream sequence. Regulatory elements sit on either side of a gene.

Adding an independent forward state-space model to an independent backward one ties the diagonal to the vectors that build the off-diagonal terms. Hydra puts the forward scan, the backward scan, and a free diagonal into one quasiseparable mix. The diagonal is its own function of the input. Forward and backward share one input projection, so the parameter count barely moves.

The block comparison holds the rest fixed: 524 kb of input, a convolution tower down to 128 bp, six mixer blocks, a U-net back to 32 bp, 768 channels, 1,870 human tracks and 769 mouse tracks. At 8,192 tokens, Hydra is 5.7 times faster than the distance-window attention Borzoi inherited from Enformer, and 2.4 times faster than the same attention with rotary embeddings. Rotary embeddings lost 10 of 12 coverage, gene, and eQTL metrics, so the rest of the comparison uses distance windows.

Linear cost made the U-net optional. Running the blocks at 32 bp and dropping both U-net stages cut parameters from 36.1 million to 30.0 million, and training and validation metrics rose. Interleaving transformers with Hydra was less stable and less accurate. Raising the state dimension above 1 changed little, so the released models fix it at 1.

Cerberus is the mean of eight replicates. Each trains on 14 of 16 folds for 50 epochs, about 11 days on one A100. One network has 138.3 million parameters, against 185.9 million for Borzoi. Input length is 786,432 bp. A convolution tower reaches 1,536 channels at 32 bp, then eight Hydra blocks with 20 heads each. No attention and no U-net. The heads predict 8,361 human tracks and 3,102 mouse tracks, adding CLIP, pseudobulks from a polyadenylation-regulator screen, and mRNA-decay time courses from 4sU labeling and actinomycin D shutoff. The loss covers the full output. Edge bins are down-weighted with a super-Gaussian window instead of being cropped out of supervision the way Borzoi cropped them.

Expression QTLs are scored as the log fold change of predicted coverage summed over a gene's exons. Splicing and polyadenylation QTLs use maximum profile shift: log-coverage across the gene locus is normalized to a distribution per allele, and the score is the largest absolute per-bin difference. That stops total expression from standing in for a shape change. The benchmark is GTEx v11 SuSiE fine-mapping. Positives have posterior inclusion probability at least 0.9. Negatives are nominally significant variants that stay below 0.01 in every tissue, matched on alleles, frequency, expression, and distance to the transcription start site, a splice site, or a polyadenylation site. Those covariates separate fine-mapped variants from unselected ones with no sequence model at all. None of these models train on genetic variation, so the QTL axis sits outside their data choices. The union holds 23,836 eQTL, 20,732 sQTL, and 4,991 paQTL SNPs. Scoring uses 48 tissues.

Results

In the matched-block comparison, Hydra reaches 0.679 AUPRC on fine-mapped eQTLs against 0.671 for distance-window attention, higher in 46 of 48 tissues. Effect-size Spearman ρ moves from 0.325 to 0.330, and four replicates do not separate that gap.

ComparisonMetricCerberusBorzoi
eQTL, mean of 48 tissuesAUPRC0.6920.668
eQTL effect sizeSpearman ρ0.3790.321
sQTLAUPRC0.7180.689
paQTLAUPRC0.7090.691

Cerberus leads on both eQTL metrics in 47 of 48 tissues. The margin grows with distance to the transcription start site, from +0.022 AUPRC within 3 kb to +0.056 beyond 100 kb. Absolute accuracy moves the other way. Cerberus falls from 0.792 AUPRC and 0.593 ρ within 3 kb to 0.604 and 0.161 beyond 100 kb. It leads on sQTLs in every tissue and on paQTLs in 46 of 48. The splicing and polyadenylation gain sits within 10 kb of the affected site. Beyond 100 kb neither model beats chance, and those positives were left out of the comparison. For a tibial-artery eQTL 41.7 kb from the PROM1 start site, Cerberus's gene-level effect is about 4.5 times Borzoi's. In-silico mutagenesis places the variant in an MEF2C motif that the alternate allele disrupts.

On expression, the tissue-level ranking favors AlphaGenome: AUPRC 0.703 versus 0.692, higher in 40 of 48 tissues, and effect-size Spearman ρ 0.415 versus 0.379. The class breakdown does not keep that order. AlphaGenome's median lead is 0.021 AUPRC in coding sequence and 0.019 within 3 kb of the start site. Cerberus leads by 0.018 at 3 to 20 kb and by 0.049 at 20 to 100 kb. On 3' UTR variants the median favors Cerberus by 0.019, and a cluster bootstrap over 895 variants does not separate that from zero. The 5' UTR and non-exonic class is 88% of positives and the 3' UTR is 5%, so the tissue-level mean is pulled by the large class. Splicing and polyadenylation were not compared. AlphaGenome's API scorers read annotated junctions and cleavage sites, a different statistic from maximum profile shift, and they return no value for 22% and 48% of variants at the designated gene.

On one A100 80 GB, a single Cerberus forward pass at 1 Mb takes 135.7 ms and 4.1 GB, against 681.1 ms and 36.1 GB for AlphaGenome. At the 786 kb training length the figures are 101.5 ms and 3.2 GB. At a matched 524 kb, Cerberus is the slower pass, 68.9 ms against Borzoi's 45.2 ms, because its mixer runs at 32 bp rather than 128 bp and sees four times as many tokens. Latency then scales roughly with length to the power 0.98 for Cerberus and 1.40 for Borzoi, and the curves cross between 1.05 and 1.31 Mb. The released packages differ too. Cerberus is eight networks. AlphaGenome distills 64 teachers into one student. Eight Cerberus passes at 1 Mb take 1.09 s and still peak at 4.1 GB, against 0.68 s and 36.1 GB.

Inside the trained blocks, the median half-life of the decay grows from 0.9 kb in the first block to 79 kb in the seventh. Deep heads spike their step size on promoter-like elements and CTCF peaks, about 27-fold and 16-fold over background by block seven. Proximal enhancer-like elements are 5.4-fold and distal ones 2.5-fold. GC content does not account for the promoter enrichment.

Why it matters

Weights, training code, and the matched QTL benchmark are released under Apache 2.0, and the stack has moved from TensorFlow to PyTorch. One replicate is 11 days on one A100, so fine-tuning on a new track is a realistic experiment.

For variant ranking, the gain over Borzoi is largest on eQTLs far from the gene, where annotation is thin and a score is more likely to change the order. Coding sequence and the first few kilobases still favor AlphaGenome, which also covers Hi-C and splice junctions under a non-commercial license. Cerberus is the open model that fits in roughly an order of magnitude less GPU memory and can be trained in house.

Limitations

Fine-mapping is an inference. In the replication check cited in the discussion, a third of GTEx blood credible sets that resolve to one variant at PIP above 0.95 recover nothing for the same gene in TOPMed, and the estimated non-causal fraction among those high-confidence variants is 21 to 24%. This paper's curves have the same shape. Agreement between model and QTL rises with PIP and with association significance. Some of the disagreement is the label.

Distal regulation is still weak. The best model here reaches about 0.70 AUPRC on matched negatives, against a chance rate of 0.5. Cerberus falls from 0.79 within 3 kb of the start site to 0.60 beyond 100 kb. Past 100 kb, neither model beats chance on splicing or polyadenylation QTLs. Indels were not scored.

The AlphaGenome comparison covers expression only. That model is a distilled student, scored on the forward strand through an API, because the license does not allow running it any other way. Coverage comparisons with Borzoi are also not on the same folds or the same transforms. The state dimension was barely searched. Nearly half the heads in the first block read in only one direction, and that specialization was left untested. Distillation was not applied either, so the eight-network ensemble is still slower at 1 Mb than AlphaGenome's single student.

Terms

Source

What people are saying

All paper explainers