Aster Launches High-Speed Inference API
ycombinator · x · 2026-07-16
Y Combinator reshared the launch of Aster Labs' new inference API, "Aster Inference."
Key points include:
- They claim it is the world's fastest inference API, "created by AI research agents."
- Regarding GPU inference performance, Aster provided two metrics: OpenAI's gpt-oss-120b at 644 tps, and GLM 5.2 at 281 tps.
- The team stated they will directly productize the inference optimization discoveries made through automated open-ended research, and will continuously iterate on both their inference products and new AI products.
More from Infra
- Nebius says SlimSpec speeds speculative decoding 8–9% without shrinking the vocabulary — Arindam_1729 · 2026-07-21
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- Chamath says open-sourcing Grok would push AI margins from models to infra and apps — Dan_Jeffries1 · 2026-07-21
- EU AI competitiveness is under pressure as firms double down on chips, ethics, and talent — nordicinst · 2026-07-21
- AI bottlenecks are shifting to memory, optics, yield control and power — thedealdirector · 2026-07-21