From PDF archives to a trainable model: an open-source fine-tuning data workflow

Puzzleheaded_Box2842 · reddit · 2026-09-30

Domain knowledge locked in PDFs doesn't enter a model by upload or naive RAG — it needs to become reliable training data first.

The proposed pipeline: extract text/tables/structure from PDFs; denoise, fix layouts, dedupe, chunk; generate domain-specific QA or instruction SFT samples; evaluate and filter the data; fine-tune via a training backend; then rerun the pipeline on new documents for controlled incremental updates.

Tooling: OpenDCAI/DataFlow covers data prep (PDF processing, cleaning, generation, eval, filtering, training-format conversion) and OpenDCAI/DataFlex serves as the configurable fine-tuning backend. "Dynamic training" here means a repeatable validate-then-update loop, not blind weight updates on every upload.

Original post →

More from Research

Research channel →