1,451 verified max-effort reasoning traces from DeepSeek V4.1 Flash released under MIT
Paramecium_caudatum_ · reddit · 2026-09-19
A new MIT-licensed dataset ships 1,451 answer-verified reasoning traces from DeepSeek V4.1 Flash at max reasoning effort — 11.6M reasoning tokens total, 8k avg per trace (max 119k) — designed for SFT/distillation of small reasoning models.
Pipeline: starting from 53,740 competition problems (Big-Math-RL-Verified), dedup and quality filtering left 49,387; hard-band selection (llama8b solve rate ≤0.2) yielded a stratified sample of 2,000. Generation took 9 minutes at 2,500 concurrency (208 tok/s). Pass@1 was 78.5%; only correct traces ship.
Verification: a sympy-based code grader matched 1,462/1,957 directly; the remaining 495 went to a locally-run Qwen3.5-4B judge that only rescues equivalent answers, never rubber-stamps — and flagged 3 grader bugs.
Decontamination: Qwen3-VL-Embedding-8B embeddings checked against MATH-500 (cos ≥0.94) and a 993-problem AIME reference (1983–2026), dropping 86 traces.
Caveats: answer-verified only (no step-level CoT audit), single teacher, approximate dedup, inherited gold answers. 87 traces exceed 30k thinking tokens, useful for long-horizon evals.
More from Research
- Cosmos DB VLDB paper shows better scheduling can be worth $100M+ a year — blaizedsouza · 2026-09-19
- Debate: Verification, Not Data, Is the Missing Leap for Useful AI — gerardsans · 2026-09-19
- Tiny tuned classifier beats Jev: GLiNER 2.5 hits 99.7% vs 83.6%, 8.8x faster locally — rickasaurus · 2026-09-19
- Researcher proposes open eval cards and public benchmark repository to fix AI evaluation trust — evijit · 2026-09-19
- How to Benchmark Enterprise AI Memory Beyond LoCoMo — blaizedsouza · 2026-09-19
- MIT's injectable magnetic nanoantennas kill 52% of drug-resistant brain cancer cells — RosalindPicard · 2026-09-19