OpenBMB open-sources training data: 400B tokens of code and 500K agent samples
zibuyu9 · x · 2026-09-10
OpenBMB released a suite of open-source training datasets behind MiniCPM5-2B, covering code, agent training, RL, and large-scale data refinement:
- UltraX: a function-calling approach to data refinement turning fine-grained edits into executable operations; open preview has 100B tokens across five refined web datasets.
- UltraData-Code: 400B tokens (L2) and 150B tokens (L3) across 11 programming languages.
- UltraData-SFT-Agent-2609: 500K agent samples spanning Search, Code Agent, General Agent, and tool use, with tasks like web search, software engineering, multi-turn memory, and database interaction.
- UltraData-RL-2609: 80K+ RL samples across math, coding, knowledge, and long-context.
More from Research
- ICLR 2027 opens Google's Gemini-powered Paper Assistant Tool to submitters for free pre-review feedback — iclr_conf · 2026-09-11
- Adversarial examples and geometry: an ECCV 2026 paper — ducha_aiki · 2026-09-11
- CMU Researchers Unveil TAHI, a Human-Agent Interaction Framework for Expert-Grade AI Output — EchoShao8899 · 2026-09-11
- Robotic fly achieves straight, level flight and recovers from wind gusts — moyix · 2026-09-11
- Two Millennium Problems solved within a month? Viral visualization gets a clearer remake — lishali88 · 2026-09-10
- AWS introduces AEM, a turn-level metric to isolate cascading agent errors — AWS ML Blog · 2026-09-10