Hugging Face releases SmolLM3 mid-training checkpoint amid 200x efficiency debate
eliebakouch · x · 2026-08-20
SmolLM3 author eliebakouch announced the release of the pre-SFT mid-training checkpoint (the it-mid-training branch of SmolLM3-3B-checkpoints, Apache-2.0, 6.17GB), recommending future work start from it — expecting it to be beaten easily given it's a year old. He calls the "200x" claim somewhat overstated but nice work, and agrees mid-training tokens shouldn't be weighted the same as the final SFT dataset.
More from Models
- Harmony tags confuse ChatGPT, causing incorrect DeepSeek explanation — mitsuhiko · 2026-08-20
- Demo: How to setup MINIMAX H3 — Background_Shift_333 · 2026-08-20
- Unsloth Releases Qwen3.8-27B GGUFs with 10% Higher Accuracy — danielhanchen · 2026-08-20
- Claude 5.6 Sol Ultra with computer use called a qualitative leap like Opus 4.5 with MCP — curious_vii · 2026-08-20
- Gemini 3.7 Flash Test: 350 tok/s Speed, Mixed Coding Results — haider1 · 2026-08-20
- AntLing open-sources Ling-3.0 models using WSM to replace LR decay — AcanthisittaOk1699 · 2026-08-19