xceval 0.4.0 Released: Full Evaluation Loop Control for Coding Agents
rudrank · x · 2026-07-31
The unofficial CLI tool xceval has released version 0.4.0, built for the Apple Evaluations framework.
The tool is designed to provide deterministic tools for developers or external coding agents, owning the entire evaluation loop:
- Core Features: Discovers runnable targets, executes code, inspects evidence, isolates failures, reruns selected samples, compares against baselines, and enforces quality gates.
- Architecture: The agent stays outside the CLI while xceval provides deterministic interfaces for timeouts, logs, dataset revisions, and machine-readable errors. The final decision on "what is a good answer" remains in the developer's Swift code.
Related event: XCEval Open-Sourced for Xcode 27 AI Agent Evaluation(2 posts)→
More from coding & agent
- Cloning Mystery Games for $7: Kimi K3 + MCP Demo — mervenoyann · 2026-07-31
- Testing Kimi K3 Agent: Automates Game Dev with MCP and Flux — mervenoyann · 2026-07-31
- DeepSeek demonstrates autonomous subagent orchestration without prompts — teortaxesTex · 2026-07-31
- Struggling to find MCP tools? Developer seeks best practices for AI coding agents — kampak212 · 2026-07-31
- Read/write is not enough: A four-dimensional risk model for AI agent permissions — Harshit-24 · 2026-07-31
- Generating synced SFX for videos while preserving voiceover using SFX MCP — Soggy-Weather169 · 2026-07-31