FlowBank Hits NeurIPS: Precompute-and-Reuse Workflow Library Averages 73.40 Across Five Benchmarks
@furongh posted a thread introducing their team's NeurIPS-accepted paper FlowBank: Query-Adaptive Agentic Workflows Optimization through Precompute-and-Reuse (arXiv:2606.1129), a workflow-layer optimization framework for LLM multi-agent systems, along with results across five benchmarks and future research directions.
Confirmed
- Core insight: many workflow optimization methods keep only a single "winner" among candidates, but workflows with worse average performance often happen to solve the queries the winner misses, so they shouldn't simply be discarded.
- Approach: "precompute and reuse"—discover complementary workflows offline and compress them into a compact library; at inference time, predict which library member to execute for each query, balancing performance and cost; a workflow generator is retained so brand-new workflows can still be designed on the fly when needed.
- Results: an average of 73.40 across five benchmarks versus 70.40 for the strongest automated baseline evaluated, with reported lower average inference cost; all agentic workflows use the same GPT-4o mini as the executor.
- Library scale effects: on MATH, keeping the full workflow pool scores 68.11, while a curated smaller library scores 69.34—even though the full pool has more potential under a "perfect selection" assumption. The takeaway: a larger library only pays off if you're good at selecting from it.
Unconfirmed
- FlowBank currently builds its workflow library offline; "learning to reuse, adapt, or create workflows during deployment" is a next step proposed by the authors with no corresponding results yet.
- The open question the authors pose to agent developers—what signals determine whether to reuse, adapt, or create a workflow—is exploratory with no conclusions.
Why it matters
The work shows workflow-layer optimization needn't be a choice between "single winner" and "build from scratch per query": a small complementary library plus per-query selection can achieve both higher scores and lower inference cost on the same execution model, offering a viable path to low-cost deployment of multi-agent systems.
2026-10-05 ~ 2026-10-05 · 9 related posts
Primary sources
- [source] FlowBank: NeurIPS paper reuses complementary agent workflows for 73.40 avg vs 70.40 baseline — furongh · 2026-10-05
- Why throw away losing agent workflows? They may solve what the winner misses — furongh · 2026-10-05
- FlowBank curates a compact workflow bank and predicts which member to run per query — furongh · 2026-10-05
- On MATH, curating the workflow bank scored 69.34 vs 68.11 for the full pool — furongh · 2026-10-05
- Next step for FlowBank: let the workflow bank evolve through deployment — furongh · 2026-10-05
- [source] FlowBank scores 73.40 avg across 5 benchmarks vs 70.40 baseline at lower cost — furongh · 2026-10-05
- What signals should tell an agent to reuse, adapt, or design a new workflow? — furongh · 2026-10-05