Building Evals for a Locally Continued-Pretrained Qwen 3.5 4B Model

funJS · reddit · 2026-09-24

The author shares how to design evals for measuring knowledge internalization in a small locally continued-pretrained (CPT) Qwen 3.5 4B model, using n-transfer subway journey outputs as the test case. Key points: constrain outputs with a strict JSON schema, run a post-CPT SFT pass via Unsloth (alpaca format) to improve schema reliability, and measure exact, partial, and semantic matches. More details on the author's blog.

Related event: Practical Guide: Designing Evals for Continually Pretrained Small Models(4 posts)→

Original post →

More from Models

Models channel →