1,451 verified max-effort reasoning traces from DeepSeek V4.1 Flash released under MIT

Paramecium_caudatum_ · reddit · 2026-09-19

A new MIT-licensed dataset ships 1,451 answer-verified reasoning traces from DeepSeek V4.1 Flash at max reasoning effort — 11.6M reasoning tokens total, 8k avg per trace (max 119k) — designed for SFT/distillation of small reasoning models.

Pipeline: starting from 53,740 competition problems (Big-Math-RL-Verified), dedup and quality filtering left 49,387; hard-band selection (llama8b solve rate ≤0.2) yielded a stratified sample of 2,000. Generation took 9 minutes at 2,500 concurrency (208 tok/s). Pass@1 was 78.5%; only correct traces ship.

Verification: a sympy-based code grader matched 1,462/1,957 directly; the remaining 495 went to a locally-run Qwen3.5-4B judge that only rescues equivalent answers, never rubber-stamps — and flagged 3 grader bugs.

Decontamination: Qwen3-VL-Embedding-8B embeddings checked against MATH-500 (cos ≥0.94) and a 993-problem AIME reference (1983–2026), dropping 86 traces.

Caveats: answer-verified only (no step-level CoT audit), single teacher, approximate dedup, inherited gold answers. 87 traces exceed 30k thinking tokens, useful for long-horizon evals.

Original post →

More from Research

Research channel →