On MATH, curating the workflow bank scored 69.34 vs 68.11 for the full pool
furongh · x · 2026-10-05
More workflows don't always help: on MATH, keeping the full pool scored 68.11 while curating it raised the score to 69.34—even though the full pool had greater potential with perfect selection. A larger bank is only useful if you can select from it well.
More from Research
- Judea Pearl: logic and causal discovery are the two pillars of Western science — yudapearl · 2026-10-05
- Tsinghua NLP's LexReward: taxonomy-driven reward modeling for legal LLMs — TsinghuaNLP · 2026-10-05
- Protein folding post-training lifts LLM reasoning: +3.23pp on all 10 benchmarks — SJUT1 · 2026-10-05
- MotorMind: zero-shot robot manipulation with general VLMs, 95% success on real xArm6 — UIUC-CS · 2026-10-05
- HyperBrowseComp: 423 questions in 13 languages stress-test web-browsing agents — MBZUAI · 2026-10-05
- VaSE: training-free stochastic KV cache eviction for reasoning models, shown at COLM 2026 — robinomial · 2026-10-05