From PDF archives to a trainable model: an open-source fine-tuning data workflow
Puzzleheaded_Box2842 · reddit · 2026-09-30
Domain knowledge locked in PDFs doesn't enter a model by upload or naive RAG — it needs to become reliable training data first.
The proposed pipeline: extract text/tables/structure from PDFs; denoise, fix layouts, dedupe, chunk; generate domain-specific QA or instruction SFT samples; evaluate and filter the data; fine-tune via a training backend; then rerun the pipeline on new documents for controlled incremental updates.
Tooling: OpenDCAI/DataFlow covers data prep (PDF processing, cleaning, generation, eval, filtering, training-format conversion) and OpenDCAI/DataFlex serves as the configurable fine-tuning backend. "Dynamic training" here means a repeatable validate-then-update loop, not blind weight updates on every upload.
More from Research
- Anthropic open-sources jacobian-lens, decoding internal activations into readable tokens — QuixiAI · 2026-09-30
- LoRA adapters break on distilled video models, long post explains why — burkov · 2026-09-30
- LadderMan, zero-shot sim-to-real humanoid ladder climbing, wins CoRL 2026 Spotlight — yuewang314 · 2026-09-30
- EnerTune at SOSP'26 cuts LLM serving energy 1.4-2.3x vs SOTA systems — tianyin_xu · 2026-09-30
- NeurIPS workshop paper derives closed-form Hilbert metric for the SPD bicone of extended Gaussians — FrnkNlsn · 2026-09-30
- SOSP'26 paper proposes energy-conscious GPU sharing for inference serving, beyond utilization — tianyin_xu · 2026-09-30