OpenDCAI/DataFlow: open-source pipeline toolkit for pre-training data prep
Puzzleheaded_Box2842 · reddit · 2026-09-18
OpenDCAI/DataFlow is an open-source tool that organizes pre-training data preparation into reusable operators and pipelines. Each operator handles one focused task via a common interface while the pipeline chains steps and stores intermediates, combining rule-based filters, model-based evaluation, and LLM-based data synthesis. Typical pipelines cover whitespace/HTML cleanup, language and length checks, blocklist filtering, MinHash dedup, quality scoring, and structured sample generation like QA pairs or SFT data, with operators freely replaceable or reorderable.
More from Infra
- STRATUM: a 2.3km procedural city where CUDA kernels generate every asset on the GPU — techartist_ · 2026-09-18
- Teortaxes: once NS-solver-level training is mastered, remaining chip bottlenecks fall fast — teortaxesTex · 2026-09-18
- Open-source Bonsai Swarm runs a 27B model across volunteer GPU browser tabs via WebGPU — airesearch12 · 2026-09-18
- Open-source models already at SOTA — Anthropic/OpenAI edge is just 5GW compute, dev argues — ccerrato147 · 2026-09-18
- Broadcom's VMware Private AI Cloud targets enterprises bleeding money on public cloud AI workloads — DavidLinthicum · 2026-09-18
- 0.05% sampling to validate cache hits: developer marvels at compute saved across the system — DanielLockyer · 2026-09-18