Reddit asks whether continued pretraining, SFT or RL works best on Qwen3.6-27B
No-Paper-557 · reddit · 2026-07-27
Qwen3.6-27B tuning thread compares pre-training, SFT and reinforcement post-training
A Reddit user asks whether anyone has directly compared continued pre-training, SFT/LoRA and reinforcement post-training on Qwen3.6-27B, especially when the goal is to add a capability without breaking the base model.
The post cites several recent papers suggesting different trade-offs around catastrophic forgetting:
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting argues SFT forgets more than reinforcement fine-tuning
- The Role of On-Policy Data in Mitigating Forgetting reports better retention with on-policy/RL training across Qwen and Llama models
- RL Forgets! Towards Continual Policy Optimization shows RL can still forget
- Fine-Tuning Without Forgetting via Loss-Adaptive Learning reports reduced forgetting with schedule changes, including Qwen3 experiments
The user asks for real before/after results, training parameters, and benchmarks on coding, reasoning, instruction following, tool use, long context, and general knowledge.
More from Research
- A 1951 mechanical tortoise is being used to explain today’s LLM scaling walls — mtizard · 2026-07-27
- Cheap storage makes SCD Type 2 look obsolete, says a Meta-style data engineer — Zachly · 2026-07-27
- Creed-Bench launches as a new eval for personal context — craighepburn · 2026-07-27
- Researchers observe models often need an a→b→c→d path before discovering a simple arithmetic trick — dejavucoder · 2026-07-27
- Claude Code is not reliable enough for long research projects without human supervision — _akpiper · 2026-07-27
- Alibaba should ship multiple Qwen sizes, says researcher focused on interpretability — traviscline · 2026-07-27