On MATH, curating the workflow bank scored 69.34 vs 68.11 for the full pool

furongh · x · 2026-10-05

More workflows don't always help: on MATH, keeping the full pool scored 68.11 while curating it raised the score to 69.34—even though the full pool had greater potential with perfect selection. A larger bank is only useful if you can select from it well.

Related event: FlowBank Hits NeurIPS: Precompute-and-Reuse Workflow Library Averages 73.40 Across Five Benchmarks(9 posts)→

Original post →

More from Research

Research channel →