Baseten adds server-side web search for open models, 15% lower latency
baseten · x · 2026-09-17
Baseten launched Hosted Tools and Grounded Inference, bringing real-time server-side web search to open-weight models via a single configuration with partners Exa, Keenable, p0, and You.com — eliminating hand-wired orchestration, extra vendor keys, and per-round-trip latency, with claimed 15% lower latency than client-side execution. You.com also touts its Highlights mode scoring 95.17% on SimpleQA, 35% cheaper than Claude's built-in web search.
Related event: Baseten Launches Hosted Tools with Real-Time Web Search for Open Models(2 posts)→
More from Infra
- OpenAI's Jalapeño chip wasn't AI-designed: 100+ ex-Google TPU engineers and Broadcom were — JFPuget · 2026-09-17
- S3 is supposed to span AZs — devs question cheapest-tier AWS data loss claim — zetalyrae · 2026-09-17
- Build the agent setup first, pick the model second: a Linux + Tailscale + llama.cpp stack guide — max_paperclips · 2026-09-17
- Huawei unveils Ascend 960 chips for 2027 and UnifiedBus linking one million processors — mark_k · 2026-09-17
- Leak: CXMT Supplies TSV Embedded DRAM Dies for Chinese cHBM, Stacking Done In-House or via Local OSAT — zephyr_z9 · 2026-09-17
- OpenAI's Astra gets 3x throughput on NVIDIA Vera Rubin, plus 2x more in 72 hours — MickeySteamboat · 2026-09-17