River API outperforms Tinker in RL, enabling zero train-infer mismatch for MoEs
marcbhargava · x · 2026-08-18
- Benchmark Result: The River API demonstrated superior performance over Tinker in reinforcement learning runs using identical training code.
- Key Innovation: It successfully achieves zero train-infer mismatch for large Mixture of Experts (MoE) models.
- Optimization Focus: Significant effort was spent on details like routing replay to ensure optimal API results.
- Open Source: The method was validated by teaching Qwen3.6-35B-A3B to play Wordle, with full ablations and open-source code available.
Related event: TogetherAI Open-Sources XoRL for Zero Train-Inference Mismatch in MoE RL(4 posts)→
More from Research
- Weaviate Podcast: Why retrieving more and reranking isn't always better — CShorten30 · 2026-08-18
- Developer proposes mixture-of-attention architecture with centroid clustering for token routing — InfamousTrouble7993 · 2026-08-18
- Model time perception may be fuzzy interpolation heuristics — kalomaze · 2026-08-18
- Speculation on RL deployment speed affecting model generation — kalomaze · 2026-08-18
- ByteDance Paper: Are AI Agents Actually Controllable? — rohanpaul_ai · 2026-08-18
- Optimizer Geometry Changes Scaling Laws: Muon Outperforms AdamW in Hard-Rank Growth — YouJiacheng · 2026-08-18