Lior Pachter stress-tests OpenAI corpus papers, questions practical value of approximation algorithms
On October 8, computational biologist Lior Pachter posted a series of threads empirically testing and criticizing papers on approximation algorithms in the OpenAI paper corpus, which sparked a public discussion with theoretical computer scientist rrwilliams (Ryan Williams).
Confirmed
- Pachter tested a paper from the OpenAI corpus on a 2-approximation algorithm for the shortest common superstring problem, and found that building it exactly as described becomes infeasible at a scale of just 1,000 sequencing reads; the problem is related to sequencing read compression.
- Another paper on edit distance gives a randomized (1+epsilon) approximation algorithm succeeding with probability at least 2/3 under unit costs, which Pachter argues produces no alignment output and is useless for biological sequence alignment due to its unit-cost assumption.
- He tied his criticism to the (conditional) edit distance lower bound of Backurs-Indyk from 2015, and revisited his own blog post from that year, "In biology n does not tend to infinity," arguing that problem sizes in biology never approach infinity.
- Pachter's overall stance: even if these approximation results are theoretical milestones, their practical significance is limited — invoking Hardy's remark that they possess "only aesthetic value."
- rrwilliams responded that there is a trade-off: trying to solve all possible instances in a general computational model means the results may not apply to many practical uses; he also raised the question of whether there is a better way to count steps.
Why it matters
- The debate highlights the gap between theoretical computer science and real-world bioinformatics applications: complexity bounds and approximation guarantees can break down entirely at concrete engineering and biological scales.
- The OpenAI paper corpus is used as AI research material, and domain experts questioning the practical usability of the included papers is informative for assessing the corpus's quality.
2026-10-08 ~ 2026-10-08 · 5 related posts
Primary sources
- Researcher Tests OpenAI Paper's 2-Approximation: 4.29M Vertices Needed for Just 1,000 Reads — lpachter ·
- Pachter: OpenAI Corpus Edit Distance Paper Is 'Totally Useless' for Biological Alignment — lpachter ·
- Researchers Debate Tradeoff Between Generic Computational Models and Practical Relevance — rrwilliams ·
- [source] Researcher Tests OpenAI Paper's 2-Approximation: 4.29M Vertices Needed for Just 1,000 Reads — lpachter · 2026-10-08
- [source] Pachter: OpenAI Corpus Edit Distance Paper Is 'Totally Useless' for Biological Alignment — lpachter · 2026-10-08
- Theorist unimpressed by new approximation breakthroughs, echoing Hardy's 'aesthetic value' jab — lpachter · 2026-10-08
- Lior Pachter Ties OpenAI Paper Critique to 2015 Edit Distance Lower Bound Debate — lpachter · 2026-10-08
- [source] Researchers Debate Tradeoff Between Generic Computational Models and Practical Relevance — rrwilliams · 2026-10-08