MCP servers can pass every test and still fail the user; MCPJam's answer
FarRub2855 · reddit · 2026-09-18
An MCPJam employee announced a broader testing and evaluation platform for MCP servers and agent-facing products. The pain point: every call can return a valid response while the workflow still breaks — the agent picks the wrong tool, or a later step runs on stale state; and a workflow that works in one client breaks in another because ChatGPT, Claude, Cursor, Gemini, Copilot, and Slack handle context and tool calls differently.
New capabilities:
- Full multi-turn journey testing across clients
- User Testing for QA teams and beta testers running real workflows
- Agentic testing and Swarms running simulated personas through different goals
- On failure, inspect prompt, trace, latency, and outcome; save journeys as repeatable evals rerun across clients and in CI/CD
Inspector, the CLI/SDK, local evals, and conformance testing stay free and open source; User Testing, Swarms, and team workflows are part of MCPJam Cloud.
More from coding & agent
- Sim Search lets agents build a knowledge graph across 1,000+ tool integrations — JafarNajafov · 2026-09-18
- Ethan Mollick rebuilds Umberto Eco's 33,000-book library in 3D with Claude Projects — emollick · 2026-09-18
- Cua and typesafeai launch jev-use dev preview, claiming fast computer use is solved — TianbaoX · 2026-09-18
- Developer demos auto-navigating slides driven by the Jev computer-use model — threepointone · 2026-09-18
- DeepSeek Harness ships v0.1.6-alpha.2 pre-release as the app takes shape — teortaxesTex · 2026-09-18
- Self-Proclaimed ChatGPT Co-Inventor Launches Jev, Claims 200x Speed at 1/400 Cost — iamrobotbear · 2026-09-18