Pydantic AI Agent SDK scores 76.40% pass@1 on Terminal-Bench 2.1, LangChain comparison next
samuelcolvin · x · 2026-09-15
A Terminal-Bench 2.1 run scored 76.40% pass@1 using the Pydantic AI adapter of samuel colvin's Agent SDK with Luna High. The author says the next experiment will use the LangChain adapter under identical conditions, offering a benchmark data point for comparing agent frameworks.
More from coding & agent
- Running 12 market-making bots side by side: clock-based vs distance-based oracle refresh on BTC-BRL — cardosofede · 2026-09-15
- You're overpaying for Claude Code; Quotient claims 35-50% token savings — ycombinator · 2026-09-15
- Turnstone launches free AI workspace with persistent local agents — ycombinator · 2026-09-15
- OpenAI says 2 engineers + Codex rewrote habitat service in Rust, 6x CPU and 15x memory gains — imjustnewatai · 2026-09-15
- Data4Library MCP Server Exposes South Korea's 1,000+ Public Libraries — modelcontextprotocol · 2026-09-15
- Kash.click Connects MCP AI Assistants to Your Point-of-Sale System — modelcontextprotocol · 2026-09-15