Pydantic AI Agent SDK scores 76.40% pass@1 on Terminal-Bench 2.1, LangChain comparison next

samuelcolvin · x · 2026-09-15

A Terminal-Bench 2.1 run scored 76.40% pass@1 using the Pydantic AI adapter of samuel colvin's Agent SDK with Luna High. The author says the next experiment will use the LangChain adapter under identical conditions, offering a benchmark data point for comparing agent frameworks.

Original post →

More from coding & agent

coding & agent channel →