Baseten launches server-side web search for open models, claiming 15% lower latency
AccBalanced · x · 2026-09-18
- Baseten introduced Hosted Tools and Grounded Inference, bringing server-side real-time web search to open-weight models via a single configuration.
- Until now, adding web search to open models meant hand-wiring orchestration, managing separate vendor keys, and paying latency taxes on every round trip.
- Baseten claims 15% lower latency vs. client-side execution, no extra vendor key, zero orchestration; launch partners include Exa, Keeable AI, P0, and You.com.
- A former Groq executive noted GroqCloud shipped a similar feature last year, calling the move of search tool calls into the inference server a clear performance win.
Related event: Baseten Launches Hosted Tools with Server-Side Web Search for Open Models(6 posts)→
More from coding & agent
- Vibe-coded realistic shooting system ships as a free game mod on Spawn — TAbrodi · 2026-09-18
- Grok Build ships v1.0.35/36: live background task output, org hooks policy, MCP fixes — XFreeze · 2026-09-18
- LangChain publishes guide on building an agent harness with Jev — LangChain · 2026-09-18
- Dev built an action roguelike in a day with Astra agent, iterating live while playing — majidmanzarpour · 2026-09-18
- Codex finally ships message timestamps, a long-requested feature — GabGarrett · 2026-09-18
- Task boards existed 9 months; multi-agent RL finally taught models to use them — herbiebradley · 2026-09-18