Healthcare Agent Test: Most Expensive Models Make the Most Mistakes

Ubunta · x · 2026-08-05

The author spent two weeks building a healthcare AI agent for real-world evidence (RWE) and tested 5 frontier models on identical cohort generation tasks. The most expensive model (Opus 5) consumed the most input tokens and threw the most planning and tool-use errors, while Sonnet 4.6 was the most consistent and token-efficient across multi-step workflows.

The author emphasizes that for governed healthcare agents, the model is only half the problem. Deterministic tools, validation contracts, and observability form the other half—and unlike the model, that half is entirely under the developer's control.

Original post →

More from coding & agent

coding & agent channel →