KAIST's FlyBy teaches small models when to query stronger ones, beating Qwen3-14B at 2.7x lower cost

kaist-ai · hf · 2026-09-29

KAIST AI's paper shows self-refinement mostly consolidates probability mass on already-reachable solutions, splitting failures into execution bottlenecks (reflection helps) and knowledge bottlenecks (external info needed). FlyBy trains 4B/8B models to reason, diagnose, and selectively query stronger models, with SFT + cost-aware RL calibrating whether/what/how much to ask. On 1,158 hard problems across six benchmarks, FlyBy-4B hits 45.96% pass@8, beating Qwen3-14B (41.64%) at 2.7x lower serving cost, and FlyBy-8B reaches 51.81%.

Original post →

More from Models

Models channel →