2026-10-10
tangermeme encodes hg38 chr1 one-hot in under 2 seconds, about 3x faster. On a custom profile head, Captum's divergence exceeds 10^-1; tangermeme stays near 10^-7.
Genomic deep learning now predicts transcription-factor binding, histone marks, chromatin accessibility, 3D architecture, transcription, and splicing straight from nucleotide sequence. Most released work stops at raising predictive accuracy. Using a trained model to read cis-regulatory grammar is a different set of operations: insert a motif and see how predictions move, scramble a span the model may be using, substitute one or a few bases to score a variant, then interpret attribution scores. Those steps barely depend on whether the network is a convolution or a Transformer, or on the optimizer.
No shared, model-agnostic implementation of that layer has stuck. Each model ships its own analysis scripts. Captum covers attribution. Selene, kipoi, CREsted, EUGENe, and gReLU focus on training, fine-tuning, and model zoos, with thinner support for what you do after training.
tangermeme is everything but the model. Sequence edits and model operations are separate and can be stacked. Marginalization substitutes a short sequence into many backgrounds and compares predictions. Ablation does the opposite: alter or shuffle a span the model may be reading. Variant-effect scoring changes one or a few bases, not necessarily contiguous. The operation after the edit need not be a forward pass. It can be DeepLIFT/SHAP, in silico saturation mutagenesis, or a function the user supplies.
Sequence design comes in four modes. Screening generates random sequences, scores them with an objective, and keeps the best. Greedy substitution tries every single-base change and keeps the best one at each step. Motif implantation does the same with whole motifs from a database. Construct marginalization designs a short insert whose average effect, across many backgrounds, moves predictions in a specified direction. That is aimed at reporter libraries, without requiring a known structure for the construct.
Ledidi and TF-MoDISco already have maintainers, so tangermeme does not reimplement them. It only stays compatible. Attribution is batched, so reference sequences do not all have to sit in GPU memory at once, and low precision is supported. Convolutions, LSTMs, Transformers, and custom ops all run. The output can be a scalar or a base-resolution profile. Sequence length is limited by whatever memory it takes to run one example.
One-hot encoding treats the byte of a character as the index and writes unsigned 8-bit integers, skipping the usual intermediate integer array. Unexpected characters fail instead of being written wrong. The timing comparison is tangermeme 0.5.1, single-threaded, against the encoding functions in SeqPro 0.6.1, CREsted 1.4.0, gReLU 1.0.7, Selene SDK 0.6.0, and ChromBPNet 1.0.1. Only the encode call is timed, not alphabet setup.
DeepLIFT/SHAP overrides the backward pass of nonlinear layers so a reference sequence enters the calculation. Which layers count as nonlinear is an internal lookup. A custom op missing from that table, or a reused activation, makes common implementations fail silently, often with no warning in the docs. The convergence delta should be zero, and in practice near machine precision. A large delta means the per-base contributions do not add back to the change in model output. tangermeme warns when the delta is too high, can return it for monitoring, and accepts custom nonlinear functions.
A seqlet is a contiguous run of high-attribution bases. The recursive caller finds variable-length seqlets directly: the attribution sum of the span must be significant against a null, and so must every subspan longer than a minimum (4 bp by default). Attributions are binned, a frequency histogram is convolved into a null for each length, and p-values come from the survival function. Called seqlets are matched with Tomtom-lite to human ChIP-seq motifs in JASPAR 2024, then counted for occurrence, pairwise co-occurrence, and spacing.
One-hot encoding hg38 chromosome 1 takes tangermeme under 2 seconds, about three times faster than the next implementation. Figure 1c puts the encoders on one axis. The spread runs from a few seconds to several tens.
Figure 1d shows convergence deltas from 20 DeepLIFT/SHAP runs. On the count head, tangermeme and Captum both sit near 10^-7. On a profile head that needs a custom op, registering that op in tangermeme keeps the delta near 10^-7. Captum does not make the same registration easy, and its delta rises above 10^-1. The matched count-head result says both implementations can be correct. The profile-head split is the unregistered custom layer.
Seqlet calling was checked on 100 random sequences with the MYC motif CACGTG inserted in the center, attributed with DeepLIFT/SHAP from a MYC BPNet model. The recursive caller returns shorter seqlets. TF-MoDISco's default caller uses a long fixed width and covers far more of each sequence. Sweeping the p-value threshold traces a precision-recall curve, and TF-MoDISco's default lands near it. TF-MoDISco is faster on very few sequences. Once the set is modest or large, the recursive caller is faster in examples per second and in seqlets per second. Wide, abundant calls are intentional in TF-MoDISco, because later steps split and filter them. That design should not be scored as if the caller were a standalone tool.
At the PLD6 promoter, BPNet's MYC-binding prediction is driven almost only by a MYC motif. Beluga, nominally predicting the same thing, is driven by many motifs that also show up in accessibility and transcription initiation. Counting annotated seqlets across MYC peaks, and then their pairwise co-occurrence, separates the two grammars: BPNet collapses onto one motif, Beluga learns a wider lexicon and a wider set of combinations. The demo models are BPNet networks trained here for E2F3 and MYC (ENCODE ENCSR036QIR and ENCSR000EGJ, control ENCSR000BLJ), Beluga from Kipoi (K562 MYC only, 0-indexed target 614), and fold 0 of ENCODE ChromBPNet and ProCapNet.
An attribution track can look clean while the convergence delta is already above 10^-1. Silent failure is expensive, because motif calls and variant rankings then follow the wrong contributions. tangermeme makes that delta something you can check, and puts marginalization, ablation, variant effects, and seqlets behind one composable interface. Model repositories keep loading and training. The license is MIT and the package name is tangermeme. Jacob Schreiber is at the Research Institute of Molecular Pathology in the Vienna BioCenter and at UMass Chan Medical School.
The BPNet versus Beluga comparison is a practical warning. Both tasks are labeled MYC binding. Different training data and different output heads can produce grammars that barely overlap. Counting motifs before you switch models tells you more than one prediction track does.
This is a Nature Methods Brief Communication, published 8 October 2026. The demonstrations outweigh a systematic benchmark. The one-hot timing is single-threaded, one chromosome, and the encode function alone. It is not a speed claim for a full attribution pipeline. The Captum comparison is one profile head and 20 repeats. It shows that a custom layer can blow up the convergence delta. It does not show that Captum is orders of magnitude off in every setting. On the count head the two match.
The TF-MoDISco comparison uses a reimplementation inside tangermeme. Precision and recall treat only the central 6 bp as positive. If a chance MYC-like site elsewhere drove a real seqlet, both callers would be penalized. The paper states that bias.
Skipping a model zoo is deliberate. The cost is that users load BPNet or Beluga themselves. The closing discussion of coding agents is a documentation roadmap, not a feature this paper delivers. The package is largely maintained by one author, and the interface will move.