OpenRelay Launches Unified Inference Endpoint: 8 Accelerators, Up to 20% Cheaper
ycombinator · x · 2026-08-08
OpenRelay announced the launch of its unified inference endpoint, aiming to aggregate global compute resources for AI inference. The service highlights several core features:
- Multi-Hardware Support: Supports 8 accelerator SKUs, including NVIDIA, TPU, Trainium, and AMD.
- Global Scale: Processes over 100B tokens/week across 22 locations on 4 continents.
- Cost Efficiency: Claims to be up to 20% cheaper than renting the hardware yourself.
More from Infra
- Agentic Cyberattacks Spark Arms Race, Compute Demand to Outstrip Supply — Justin_Halford_ · 2026-08-08
- Musk Plans Solar-Powered AI Satellites to Solve AI's Energy Bottleneck — r0ck3t23 · 2026-08-08
- Zero-Change Foundry-Compatible Silicon Photonics MEMS Optical Switch — jwt0625 · 2026-08-08
- Does more SMs improve GPU training performance? Stas Bekman explains with numbers — StasBekman · 2026-08-08
- Musk's SpaceX to Build 10GW Nvidia GPU Cluster by 2027, Consuming 30% of Rubin Output — zephyr_z9 · 2026-08-08
- Deploying 304B Model on Dual DGX Sparks: Extreme Memory Optimization — StartupTim · 2026-08-08