AI Conference Day 2: Agent Inference Costs and Observability Steal the Show
sanjaykalra · x · 2026-10-03
Field notes from The AI Conference Day 2, where booth conversations kept circling two questions: what it costs to run agents, and who watches them once live.
- Tensormesh delivered the sharpest tokenomics take: most of an agent's inference bill goes to recomputing context the model already processed; they cache that work and reuse it across GPU fleets, bill cached input tokens at $0, and claim up to 10x GPU spend reductions.
- Chalk targets the other agent cost — stale data — with a Context Engine that assembles live context at inference time, plus Chalk Compute for re-running agents.
The meta-observation: the expo floor's real theme was agent economics and observability, not models.
More from coding & agent
- NVIDIA hands over first Vera CPU to test AI agent code environments — denisyarats · 2026-10-03
- A weekend, a few hundred lines: building a personal AI voice agent with Telnyx, gpt-live and Cloudflare — itsOmSarraf_ · 2026-10-03
- Cloudflare integrates Pi Durable into its agents SDK, and it just works — irvinebroque · 2026-10-03
- Mnemos plugin to visualize agent memory in your platform's canvas — RileyRalmuto · 2026-10-03
- GitHub Copilot CLI v1.0.92-2 fixes Windows sandboxed temp files and duplicate sessionEnd hooks — copilot-cli-release-app[bot] · 2026-10-03
- Running a universal AI memory with Notion across Grok, Claude Code and more — nbaschez · 2026-10-03