Developer slams third-party inference providers: Gemini up 10x, Luna 15s latency
julianharris · x · 2026-09-29
Developer Julian Harris shares his frustrating experience with per-token third-party inference providers, listing failure modes:
- Forced China passthrough: most (or all) Chinese models on Together route data through China
- Highly variable throughput
- Extreme slowness: OpenAI Luna low takes 15s vs 1-2s elsewhere
- Crazy price hikes: Gemini up 10x over 12 months, doubling again in January
- Strict limits: Groq throttles hard because it can't service new customers
He concludes OpenRouter's 10% premium may well be justified.
More from Infra
- Celesto: open-source persistent microVM computers for AI agents, boots in 500ms — aniketmaurya · 2026-09-29
- Redditor runs gpt-oss-120b across a phone, three Macs and two Windows PCs — ANR2ME · 2026-09-29
- BioNeMo team boosts Mixtral-8x7B training throughput 2.21x vs HF BF16 baseline — AllThingsApx · 2026-09-29
- DeepSeek's elastic compute team is hiring heavily, shares sandbox infra for large-scale agent training — teortaxesTex · 2026-09-29
- Nereus: adaptive parallelism boosts 8B PPO throughput up to 7.27x over OpenRLHF — Songlin Jiang · 2026-09-29
- Data center water use isn't about total volume, it's who runs out of local freshwater first — AryHHAry · 2026-09-29