Reddit asks whether continued pretraining, SFT or RL works best on Qwen3.6-27B

No-Paper-557 · reddit · 2026-07-27

Qwen3.6-27B tuning thread compares pre-training, SFT and reinforcement post-training

A Reddit user asks whether anyone has directly compared continued pre-training, SFT/LoRA and reinforcement post-training on Qwen3.6-27B, especially when the goal is to add a capability without breaking the base model.

The post cites several recent papers suggesting different trade-offs around catastrophic forgetting:

The user asks for real before/after results, training parameters, and benchmarks on coding, reasoning, instruction following, tool use, long context, and general knowledge.

Original post →

More from Research

Research channel →