tangermeme gets major speedups via auto-optimization as community debates AI code-rewrite attribution
jmschrei shared a flurry of results on October 11 from using auto-optimize to overhaul the open-source gene-modeling toolchain, and proposed community norms around "using AI to speed up others' scientific code," sparking discussion about attribution and open-source ethics.
Confirmed
- After auto-optimize automatically searched for improvements to tangermeme's onehotencode, geometric mean runtime dropped from 0.1204 ms to 0.02729 ms; encoding all 455 hg38 contigs now takes about 1.5 seconds, and the change has been merged.
- A newly merged PR rewrites extractloci FASTA reading to use .fai-index-based memory mapping, with a numba multi-threaded kernel performing one-hot encoding directly on the read windows, making data loading ten times faster at 0.35 seconds.
- The auto-optimization loop also reworked the bigWig read/write tool bam2bw and tangermeme's extractloci, delivering a 20x speedup in bigWig I/O; the optimized code was released as the open-source tool figwig.
Community Norms Proposal
- jmschreiber91 argued that taking someone else's code, speeding it up, and publishing it under your own name should be resisted by the community, while fast implementations do provide real value.
- He sees this as both an engineering and a community-norms issue: AI-rewritten versions should respect original-author attribution, a topic growing more pressing as agentic programming spreads.
Why It Matters
- These order-of-magnitude speedups directly improve everyday genomics pipelines, and tools like figwig are already available for community use.
- AI-driven auto-optimization of others' code is becoming the norm; how to balance "the value of fast implementations" against "original authors' attribution rights" is a new norm the open-source scientific software community urgently needs to define.
2026-10-11 ~ 2026-10-11 · 5 related posts
Primary sources
- [source] figwig: auto-optimized multithreaded bigWig reader/writer, 20x faster FIMO — jmschreiber91 · 2026-10-11
- tangermeme's extract_loci now 10x faster: ATAC-seq locus loading drops from 3.64s to 0.35s — jmschreiber91 · 2026-10-11
- [source] tangermeme encodes all 455 hg38 contigs in 1.5 seconds after auto-optimize speedups — jmschreiber91 · 2026-10-11
- Optimized one_hot_encode finishes all 455 hg38 contigs in 1.5s — and the ethics of speeding up others' code — jmschreiber91 · 2026-10-11
- [source] Speeding up others' scientific code with AI needs community norms on attribution — jmschreiber91 · 2026-10-11