The Washington Post runs Ask The Post AI on independent inference with 1.79B tokens a month
togethercompute · x · 2026-07-25
The Washington Post’s AI search product runs on independent model inference
The Washington Post says its Ask The Post AI product, launched in November 2024, uses Together AI’s inference platform to answer readers’ questions from the paper’s deeply sourced journalism.
Why it switched
- The Post wanted independent model inference instead of being locked into proprietary APIs.
- It needed enterprise-grade serving for thousands of daily queries with sub-10-second latency.
- Self-hosting was seen as too operationally heavy, while hyperscalers and proprietary providers left too little control over cost and model choice.
What the deployment looks like
- Dedicated endpoints for custom GPU infrastructure
- Open-model hosting for Llama and Mistral
- Hybrid deployment with both dedicated and serverless inference
- Direct engineering support from Together AI on configuration and optimization
Reported results
- 1.79 billion input tokens per month
- 2-second consistent response times
- Fixed monthly pricing instead of variable per-token costs
- Better user engagement through answer-expansion and source-click behavior
More from Companies & People
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- X drama: Anthropic researchers accused of spying on academic customers and racing them to results — basedjensen · 2026-09-11
- Investor argues Palantir-Nvidia partnership should slash Anthropic's IPO valuation — pdamodaran · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11