Unify CTO breaks down cutting agent costs 95% before launch
brandon_galang · x · 2026-08-29
LangChain's Max Agency podcast interviews Unify co-founder & CTO Connor Heggie. Unify builds agents for go-to-market teams — giving every sales rep "an engineer in their back pocket" rather than replacing reps. Key topics:
- Cutting inference costs 90-95% two weeks before launch
- The 15-requests-per-second ceiling hidden inside OpenAI's prompt cache
- Optimizing prompt caching hit rates; fork vs child subagents
- Why your LLM judge must be a different model family
- The Speed Audit: why one row at a time beats a table of 1,000
- Why Unify's subagents are just a function call, and ditching full VMs for a Python REPL
- Locking memory to keys instead of letting the agent freestyle
- What self-driving work taught him about running evals
Related event: Unify CTO Explains Slashing AI Agent Costs by 95% Before Launch(3 posts)→
More from coding & agent
- Agent-Agent Communication Overcomes Single-Machine Compute Limits — peterjliu · 2026-08-29
- Claude Code 2.1.251 Adds Model Switch Hooks — ClaudeCodeLog · 2026-08-29
- Claude Code 2.1.251 Released with Model Switch Hooks and Security Limits — ClaudeCodeLog · 2026-08-29
- Anti-slop linter flags 83 issues in 48 hours across agents — lucasmeijer · 2026-08-29
- MuleSoft Launches MCP Server for Claude Code Integration — msrivastav13 · 2026-08-29
- Clearwing: Open-Source Pentest Agent Inspired by Anthropic's Glasswing — QuixiAI · 2026-08-29