Kundaje mocks benchmark: fine-tuned scFMs only beat the most trivial baselines
anshulkundaje · x · 2026-10-06
In post 3 of his scFM critique thread, Anshul Kundaje writes that the paper only shows that, with updated metrics, some scFMs explicitly fine-tuned on perturbation data outperform the most trivial of baselines — which he greets sarcastically with 'Bravo!'. This sets up his core claim in post 4 that scFMs trained on observational data cannot reliably learn causal effects.
More from Research
- User Sim Index is broken: trivial bot scores 95% across behavioral dims — ericzelikman · 2026-10-06
- Used OpenAI Dots as a Free Agent Swarm to Break a 47-Year-Old Math Record — jaxchang · 2026-10-06
- Stanford prof says current experiment designs can't train causal virtual cell models — anshulkundaje · 2026-10-06
- TIDES dataset on multi-party and multi-agent collaboration to debut at COLM 2026 — josephseering · 2026-10-06
- LakeQuest QA benchmark, testing RAG on messy enterprise tables and docs, hits COLM — hllo_wrld · 2026-10-06
- Frontier Data Summit lineup: Chollet, Dawn Song headline batch of new agent benchmarks — StanfordAILab · 2026-10-06