Building Evals for a Locally Continued-Pretrained Qwen 3.5 4B Model
funJS · reddit · 2026-09-24
The author shares how to design evals for measuring knowledge internalization in a small locally continued-pretrained (CPT) Qwen 3.5 4B model, using n-transfer subway journey outputs as the test case. Key points: constrain outputs with a strict JSON schema, run a post-CPT SFT pass via Unsloth (alpaca format) to improve schema reliability, and measure exact, partial, and semantic matches. More details on the author's blog.
Related event: Practical Guide: Designing Evals for Continually Pretrained Small Models(4 posts)→
More from Models
- Frontier labs' health week: Opus 5.5 cut 40%, 950 agents find new enzyme — HealthcareAIGuy · 2026-09-24
- Monologue launches in-house dictation model mono-1: 55% fewer edits, 3x faster than API pipeline — every · 2026-09-24
- Anthropic cut Opus 5.5 prices, then broke four things your agent depends on — rseroter · 2026-09-24
- ChatGPT Voice with tools and MCP impresses: interruptible, pulls local Mac files — athyuttamre · 2026-09-24
- Set max_tokens to 1 and Read Logprobs: Turn Any Hosted LLM Into a Classifier — keep_up_sharma · 2026-09-24
- Not every job needs the smartest AI model—good enough wins — ChrisUniverse · 2026-09-24