Two opening tokens lift base model MATH-500 from 42% to 78%, rivaling RL training
arankomatsuzaki · x · 2026-10-06
A paper by Sophie L. Wang, Amil Dravid, et al. (including Alexei Efros) shows base models' reasoning can be elicited without any training:
- Method: just prefill a few opening tokens. Adding ".\n\n Okay" raises Olmo-3-7B's MATH-500 pass@1 from 42% to 78%; " Alright," lifts Qwen3-14B from 72% to 87% — close to its RL-trained counterpart.
- Mechanism: these effective cues come from associations learned in pretraining data; RL makes such cues more likely, and fixing them recovers much of RL's gain over the base model.
- Further finding: by changing those data associations, even "Chicken" can elicit reasoning.
- Significance: the work offers both a minimal way to access base-model reasoning and a data-side explanation for what post-training adds, with results generalizing across model families. Code and paper are open.
Related event: Study: Two Opening Tokens Unlock Base Model Reasoning Without RL(4 posts)→
More from Models
- SuperGrok users hit Grok Bot usage limits fast, calling for a 1.5x bump — nima_owji · 2026-10-07
- jev Decision Index v0.3 arrives with private benchmark half and vision rankings — multimodalart · 2026-10-07
- NVIDIA's Nemotron Labs partners with Artificial Analysis on open-model evaluation — NVIDIAAI · 2026-10-07
- Mistral Large 4 generates a Japanese-inspired floating voxel island, sparking 'Is the EU back?' buzz — kevinkern · 2026-10-07
- Marin 535B-A23B open model training crosses halfway, Percy Liang shares learnings — ericjang11 · 2026-10-07
- llama.cpp ships Day-0 support for Google's EmbeddingGemma 2 — ggerganov · 2026-10-07