Testing Qwen3.8-27B inside a coding agent: evaluation plan
Binary_orchid · reddit · 2026-08-25
The author shares an evaluation plan for Qwen3.8-27B running locally within the EvoX coding agent harness. The focus is on the model's behavior in complex tasks: repo reading, tool calling, error recovery, and multi-turn consistency, rather than ranking desktop apps.
Planned tasks include:
- Screenshot-to-page: Comparing Low/Medium/XHigh reasoning levels in a clean frontend repo, measuring first render time, correction turns, and total time.
- USGS Earthquake Dashboard (Advanced): Testing rendering, filtering, and interactions with static GeoJSON and live feeds.
The author controls variables (e.g., disabling experience reuse) and measures reasoning token counts and time-to-first-tool-call to balance quality and latency. The community is invited to critique the plan.
More from coding & agent
- Expert Validation Becomes Bottleneck in AI Projects — anmarasovic · 2026-08-25
- RocketRide: open-source visual AI pipeline engine with C++ core hits 7k GitHub stars — abhishek__AI · 2026-08-25
- RocketRide: Visually build AI pipelines inside your IDE — abhishek__AI · 2026-08-25
- Jared Palmer: Stylex Superior to Tailwind for the AI Agent Era — Vjeux · 2026-08-25
- Migrating from OpenClaw to Grok Bot: A real-world agent workflow experiment — heyneighbor · 2026-08-25
- Claude autonomously builds, fixes, and deploys cancellation flow — mhmazur · 2026-08-25