Stanford team finds frontier LLMs like Opus 5 and GPT6 poor at DNA motif annotation out of the box

Stanford professor Anshul Kundaje clarified in a thread on September 24 that his team's focus is not motif discovery, but biological annotation of already-discovered genomic sequence motifs — i.e., what they are, what they do, and which factors they bind. The team tested Claude and GPT for annotating DNA motifs, and both Opus 5 and GPT6 performed poorly out of the box. Kundaje believes this extremely labor-intensive manual task is well suited for AI learning, constituting a multimodal reasoning task that requires dedicated training; once trained properly, it could unlock accurate annotation of large-scale motif catalogs, and the team is accumulating enough labeled data for this. Two related pieces of information also appeared in the thread: jatinn0 discussed with Kundaje using LLMs to break the motif annotation bottleneck; researcher Joseph Kellogg observed features in Mythos model activations that behave similarly to genomic language models (such as ESM-type pLM/gLM), and jatinn0 proposed a hypothesis based on this — LLMs large enough and pretrained with sequence data may have internalized some form of sequence understanding. Kundaje added that the team has never worked with Mythos.

Confirmed

Not Yet Confirmed

Why It Matters

2026-09-24 ~ 2026-09-24 · 6 related posts

Primary sources

1 near-duplicate retellings: anshulkundaje