RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation
The paper introduces RayOrch, a programming model and distributed execution engine designed for foundation-model data preparation pipelines. RayOrch preserves parent-child relations and execution order using per-call FIFO ready queues to batch children across parents. On NVIDIA H20 GPUs, RayOrch achieves 15.14 times speedup when scaling MinerU from 4 to 64 GPUs, 7.82 times speedup when scaling a video pipeline from 8 to 64 GPUs, and reduces end-to-end time compared to Ray Data and Daft.
RayOrch achieves 15.14 times speedup when scaling MinerU from 4 to 64 GPUs and 7.82 times speedup when scaling a video pipeline from 8 to 64 GPUs.
It reduces end to end time by 13.1 percent versus Ray Data and 29.0 percent versus Daft on MinerU, and by 16.0 percent versus Ray Data on Docling.