FlowBank scores 73.40 avg across 5 benchmarks vs 70.40 baseline at lower cost
furongh · x · 2026-10-05
Across five benchmarks, FlowBank averages 73.40 vs 70.40 for the strongest automated baseline, at lower reported average inference cost. All agentic workflows use the same GPT-4o mini executor, so gains come from orchestration rather than the model.
More from Research
- Judea Pearl: logic and causal discovery are the two pillars of Western science — yudapearl · 2026-10-05
- Tsinghua NLP's LexReward: taxonomy-driven reward modeling for legal LLMs — TsinghuaNLP · 2026-10-05
- Protein folding post-training lifts LLM reasoning: +3.23pp on all 10 benchmarks — SJUT1 · 2026-10-05
- MotorMind: zero-shot robot manipulation with general VLMs, 95% success on real xArm6 — UIUC-CS · 2026-10-05
- HyperBrowseComp: 423 questions in 13 languages stress-test web-browsing agents — MBZUAI · 2026-10-05
- VaSE: training-free stochastic KV cache eviction for reasoning models, shown at COLM 2026 — robinomial · 2026-10-05