Industrial-Instruction: Building Datasets from Tech Reports

Parsa Bakhtiari · hf · 2026-08-25

Addressing the difficulty of indexing heterogeneous industrial reports, this paper presents the Industrial-Instruction framework. Using 906 public Panasonic documents (7,525 pages), the authors created two open QA datasets (13.6k pairs each) via layout-aware extraction and semantic retrieval. Fine-tuning small open LLMs (<10B) boosted Set-Match Accuracy from 28.5% to 42.0%. The study compares data generated by Qwen3-30B-A3B-Instruct and Claude-Opus-4.6, finding Claude yields cleaner data and higher gains at roughly 100x the cost.

Original post →

More from Research

Research channel →