Factory Router cuts inference costs 63% in production; dynamic routing squeezes token prices
matanSF · x · 2026-09-25
Factory AI says Factory Router now cuts inference costs by 63% in production, up from 42% in late June, while routed session volume grew roughly 4x.
Commentator matanSF argues dynamic model routing creates a more efficient free market across providers: for any given task, whoever fails to offer the best/fastest/cheapest tokens loses that traffic to a competitor — putting persistent pricing pressure on every model vendor.
More from Infra
- Perplexity Launches Fast Search API: 95% of Results in Under 230ms on Rust-Based Photon — perplexity_ai · 2026-09-25
- 100B Model Trained Across 5 Data Centers on Plain Internet Links at 30.8% MFU — markjeffrey · 2026-09-25
- AMD gaining 10 points of GPU share would be 'transformational', analyst argues — Beth_Kindig · 2026-09-25
- Google sees orbital AI data centers reaching cost parity with terrestrial ones by mid-2030s — McDonaghMatthew · 2026-09-25
- Goldman Sachs hikes AI power forecasts: 2030 data center capacity raised to 217GW — McDonaghMatthew · 2026-09-25
- Puro-2B: an open recipe trains a Qwen2-1.5B-beating LLM on RTX 5090s for just $4.4K — IgorCarron · 2026-09-25