PKU open-sources RayOrch, lineage-aware data-prep engine with up to 15.14x speedup
PekingUniversity · hf · 2026-09-28
- Peking University released RayOrch, an open-source programming model and distributed execution engine for foundation-model data preparation pipelines.
- Problem: pipelines expand heterogeneous documents/videos into parent-child record trees; existing systems either hide parallelism behind coarse-grained jobs or expose flat records, forcing apps to manage lineage and regrouping.
- RayOrch preserves parent-child relations throughout execution: programs declare ordered variable-cardinality expansions with matching gathers, validated by the compiler; the runtime tracks child membership, parents, ordinals, and terminal states; per-call FIFO ready queues batch children across parents; gathers reconstruct results from declared membership; typed parent-scoped failures suppress only that parent's undispatched siblings.
- Performance on NVIDIA H20 GPUs: 15.14x speedup scaling MinerU from 4 to 64 GPUs, 7.82x for a video pipeline from 8 to 64 GPUs; reduces end-to-end time by 13.1% vs Ray Data and 29.0% vs Daft on MinerU, and 16.0% vs Ray Data on Docling.
- Code: github.com/OpenDCAI/RayOrch
More from Infra
- FailureAtlas: most severe LLM gateway failures return HTTP 200 and silently corrupt your app — its_vayishu · 2026-09-28
- Used RTX 3090 prices creep toward $1,500 on eBay amid GPU shortage — sleight42 · 2026-09-28
- gufo inference doubles prefill speed vs llama.cpp forks for Qwen 3.8 Flash Next on Strix Halo — fallingdowndizzyvr · 2026-09-28
- Scraping p50 stabilized at 2s: keep your app and databases colocated — DanielLockyer · 2026-09-28
- Why credit quality matters most in compute: clusters are opaque, people are trackable — AccBalanced · 2026-09-28
- First buyable Gorgon Halo chip lands as AMD 495 with only a small bump over 395 — julianharris · 2026-09-28