Testing Qwen3.8-27B inside a coding agent: evaluation plan

Binary_orchid · reddit · 2026-08-25

The author shares an evaluation plan for Qwen3.8-27B running locally within the EvoX coding agent harness. The focus is on the model's behavior in complex tasks: repo reading, tool calling, error recovery, and multi-turn consistency, rather than ranking desktop apps.

Planned tasks include:

The author controls variables (e.g., disabling experience reuse) and measures reasoning token counts and time-to-first-tool-call to balance quality and latency. The community is invited to critique the plan.

Original post →

More from coding & agent

coding & agent channel →