NeoHorse-1: Recursive self-improvement via agentic post-training with routing harness
TokenRhythm · hf · 2026-09-09
NeoHorse-1, released on Hugging Face, pursues recursive self-improvement through agentic post-training: it combines intelligent routing, structured feedback loops, and curriculum-based distillation to improve model capabilities across agent benchmarks.
More from Research
- Proof verifiers trusted for humans may fail against AI-crafted exploit proofs — yoavgo · 2026-09-10
- Columbia DAP Lab's VLDB 2026 keynote: agentic data environments as the next research frontier — adityagp · 2026-09-10
- Nupur Kumari, author of first thesis on customizing generative image models, joins OpenAI — junyanz89 · 2026-09-10
- Frank Nielsen's info geometry textbook offers a foundational ML entry, now on ChapterPal — burkov · 2026-09-10
- Train on Frontier Papers or Build RL Envs? An Insider Debate on Math Model Training — ctjlewis · 2026-09-10
- GPN-Star precomputed variant-effect scores for human genome and 5 model organisms land on Hugging Face — anshulkundaje · 2026-09-10