The Washington Post runs Ask The Post AI on independent inference with 1.79B tokens a month
togethercompute · x · 2026-07-25
The Washington Post’s AI search product runs on independent model inference
The Washington Post says its Ask The Post AI product, launched in November 2024, uses Together AI’s inference platform to answer readers’ questions from the paper’s deeply sourced journalism.
Why it switched
- The Post wanted independent model inference instead of being locked into proprietary APIs.
- It needed enterprise-grade serving for thousands of daily queries with sub-10-second latency.
- Self-hosting was seen as too operationally heavy, while hyperscalers and proprietary providers left too little control over cost and model choice.
What the deployment looks like
- Dedicated endpoints for custom GPU infrastructure
- Open-model hosting for Llama and Mistral
- Hybrid deployment with both dedicated and serverless inference
- Direct engineering support from Together AI on configuration and optimization
Reported results
- 1.79 billion input tokens per month
- 2-second consistent response times
- Fixed monthly pricing instead of variable per-token costs
- Better user engagement through answer-expansion and source-click behavior
More from Companies & People
- Prentis, New AI Lab by Reid Hoffman and Marc Pincus, Seeks $100M — TechCrunch AI · 2026-07-25
- A16z survey names Elad Gil, Marc Andreessen and Elon Musk as top angels — jfischoff · 2026-07-25
- Engineer says a scratch-built data platform now serves 400M users internally — teodorio · 2026-07-25
- Opus 5 powers a live AI tutor prototype for AlphaSchool — mattshumer_ · 2026-07-25
- VaynerX says a $20,000 internal app was prototyped in hours with Claude — tomcrawshaw01 · 2026-07-25
- ECCV Faces Backlash Over Forced Full Registrations for Visa-Blocked Students — gan_chuang · 2026-07-25