MCP Evals Must Name the Client Config, Not Just the Server

Sussybaka800869 · reddit · 2026-10-02

A methodology discussion argues that the same MCP tool can work in one client and barely be discovered in another, and adding a skill that explains when to use it changes results again — so "the MCP server passed" omits much of the experiment.

Reef Infra's harness adapters provide pieces for testing these differences: an adapter renders the client's configuration, rules and skills, launches a headless task, and reads that client's session format. Pi, OpenCode, Claude Code, Codex, DeepSeek Harness and Hermes each have separate adapters.

Within one adapter, Reef Infra can compare an existing harness with a proposed change on the same tasks, e.g. testing whether a skill improves completion of a tool-backed task. Cross-client comparisons require separate experiments with matched models, tasks and tool access; the adapter does not make environments equivalent automatically.

The differences are consequential: Codex's Reef Infra setup rejects code extensions and disables hooks, plugins and web search during evaluation; Pi and OpenCode have extension/plugin surfaces; Claude Code can render a plugin entry but activation still needs the relevant marketplace or settings reference. MCP fixtures and success checks remain yours to provide, and the client configuration must be kept with the result because it determines what the measurement means.

Original post →

More from coding & agent

coding & agent channel →