KAIST's FlyBy teaches small models when to query stronger ones, beating Qwen3-14B at 2.7x lower cost
kaist-ai · hf · 2026-09-29
KAIST AI's paper shows self-refinement mostly consolidates probability mass on already-reachable solutions, splitting failures into execution bottlenecks (reflection helps) and knowledge bottlenecks (external info needed). FlyBy trains 4B/8B models to reason, diagnose, and selectively query stronger models, with SFT + cost-aware RL calibrating whether/what/how much to ask. On 1,158 hard problems across six benchmarks, FlyBy-4B hits 45.96% pass@8, beating Qwen3-14B (41.64%) at 2.7x lower serving cost, and FlyBy-8B reaches 51.81%.
More from Models
- Leaked OpenAI 'dot' details show raising phone to ear triggers ChatGPT Voice — koltregaskes · 2026-09-29
- Google to replace Gemini Gems with Skills starting November 17 — mark_k · 2026-09-29
- Carla v0.1.0: a local llama.cpp loom TUI for growing AI characters — max_paperclips · 2026-09-29
- Leaked OpenAI DevDay reveal called 'just a Grok bot / Meta Muse rip-off' — gaganghotra_ · 2026-09-29
- Running Jev at high frame rate with full-state snap inferences makes it a true System 1 — mathemagic1an · 2026-09-29
- OpenAI model naming rumor: Dots, Orbit and 'o' said to be in the mix — mark_k · 2026-09-29