Team open-sources agent interface eval project, shares early findings on testing MCP across clients

nguyenfamjj · reddit · 2026-09-10

An open-source project aimed at product teams is building a systematic way to evaluate agent interfaces like MCP, CLI, and skills: teams define real user tasks, then run them across multiple conditions (different models/harnesses, with or without skills) and score the results.

Two early findings:

The team hasn't cracked the right evaluation model yet and is asking the community how they test MCP/CLI/skill workflows today, and which criteria matter most: task success, permissions, reliability, client compatibility, or cost. Progress updates will follow.

Original post →

More from coding & agent

coding & agent channel →