Stanford team finds frontier LLMs like Opus 5 and GPT6 poor at DNA motif annotation out of the box
Stanford professor Anshul Kundaje clarified in a thread on September 24 that his team's focus is not motif discovery, but biological annotation of already-discovered genomic sequence motifs — i.e., what they are, what they do, and which factors they bind. The team tested Claude and GPT for annotating DNA motifs, and both Opus 5 and GPT6 performed poorly out of the box. Kundaje believes this extremely labor-intensive manual task is well suited for AI learning, constituting a multimodal reasoning task that requires dedicated training; once trained properly, it could unlock accurate annotation of large-scale motif catalogs, and the team is accumulating enough labeled data for this. Two related pieces of information also appeared in the thread: jatinn0 discussed with Kundaje using LLMs to break the motif annotation bottleneck; researcher Joseph Kellogg observed features in Mythos model activations that behave similarly to genomic language models (such as ESM-type pLM/gLM), and jatinn0 proposed a hypothesis based on this — LLMs large enough and pretrained with sequence data may have internalized some form of sequence understanding. Kundaje added that the team has never worked with Mythos.
Confirmed
- Kundaje's team tested Opus 5 and GPT6 on biological annotation of DNA motifs; both performed poorly out of the box
- The discussion's focus is biological annotation of motifs, not motif discovery
- The team is accumulating labeled data, aiming to enable models to accurately annotate large-scale motif catalogs
- The team has not worked with Mythos
Not Yet Confirmed
- jatinn0's claim that "sufficiently large LLMs pretrained with sequence data may internalize sequence understanding" remains a hypothesis
- The similarity in activation directions between LLMs and genomic language models observed by Kellogg awaits systematic validation
Why It Matters
- DNA motif annotation is a major bottleneck in genomics; if LLMs can master it through specialized training, gene regulation research could be significantly accelerated
- The poor out-of-the-box performance of frontier models shows that specialized biological tasks still depend on domain data and dedicated training rather than general-purpose models alone
2026-09-24 ~ 2026-09-24 · 6 related posts
Primary sources
- Kundaje tested Opus 5 and GPT6 on DNA motif annotation — both bad out of the box — anshulkundaje ·
- Anshul Kundaje: trained LLMs could break the genomic motif annotation bottleneck — anshulkundaje ·
- Researchers find Mythos activations behave like genomic language models, hinting LLMs learn DNA motifs — jatin_n0 ·
- Stanford's Anshul Kundaje: Claude/GPT are pretty bad at annotating DNA motifs — anshulkundaje · 2026-09-24
- [source] Researchers find Mythos activations behave like genomic language models, hinting LLMs learn DNA motifs — jatin_n0 · 2026-09-24
- [source] Anshul Kundaje: trained LLMs could break the genomic motif annotation bottleneck — anshulkundaje · 2026-09-24
- [source] Kundaje tested Opus 5 and GPT6 on DNA motif annotation — both bad out of the box — anshulkundaje · 2026-09-24
- Researchers Debate Using LLMs to Crack the Genomic Motif Annotation Bottleneck — jatin_n0 · 2026-09-24
1 near-duplicate retellings: anshulkundaje