Suncly verifies A2A Agent Cards match behavior, finds gaps in 8 official samples
Suncly_com · reddit · 2026-10-12
Suncly built a tool that tests whether an A2A agent actually behaves as its Agent Card claims, then cryptographically signs the result.
How it works
- Agent Cards are fingerprinted via RFC 8785 canonical JSON + SHA-256, tying results to the exact card version
- Tests are drafted only from the card's own declared examples, with human approval before running
- A separate runner process holds credentials, talks only to approved hosts, and redacts secrets
- Layer 1 is deterministic (status, schema, media type); Layer 2 uses a pinned Claude Haiku 4.5 for semantic checks, with model/endpoint/rubric versions recorded
- Results go to an append-only store; reports are Ed25519-signed, so any byte change breaks verification
Design choices: no single trust score (pass/fail/inconclusive per skill), inconclusive never counts as pass, untested areas are listed, final verdict is always flag-for-human-review.
Testing 8 official A2A sample agents surfaced several cards declaring "text" where the spec expects a MIME type like "text/plain". Limits: A2A 1.0 only, coverage bounded by the card's examples. Adversarial probes (undeclared requests, prompt injection) are next, and the team offers free signed reports.
More from coding & agent
- Cloudflare's Think proposes an execution ladder: LLMs pick their own sandbox per task — irvinebroque · 2026-10-12
- Opus 5.5 Compared: Claude Teammates Push Back, Codex Quietly Drifts Behind Green Tests — Sauers_ · 2026-10-12
- AI Writes Code Faster — So Why Aren't We Shipping Faster? — kristiyanstoyanovAI · 2026-10-12
- Enterprise AI Rollouts: 5 of 300 Users, and What Clients Still Can't Get — Initial_Orange2985 · 2026-10-12
- Microsoft Decision-1: decision models for cheap AI routing and agent orchestration — usamawahabkhan · 2026-10-12
- Hill-Climbing Skills: Browserbase Engineer Shows Agents Improving Without Touching Model Weights — AI Engineer · 2026-10-12