Industrial-Instruction: Building Datasets from Tech Reports
Parsa Bakhtiari · hf · 2026-08-25
Addressing the difficulty of indexing heterogeneous industrial reports, this paper presents the Industrial-Instruction framework. Using 906 public Panasonic documents (7,525 pages), the authors created two open QA datasets (13.6k pairs each) via layout-aware extraction and semantic retrieval. Fine-tuning small open LLMs (<10B) boosted Set-Match Accuracy from 28.5% to 42.0%. The study compares data generated by Qwen3-30B-A3B-Instruct and Claude-Opus-4.6, finding Claude yields cleaner data and higher gains at roughly 100x the cost.
More from Research
- Thomson Technical Report Outlines Blueprint for SovereignAI via Continual Learning — schwarzjn_ · 2026-08-25
- Delta-Mem Paper Reveals Lightweight Memory Mechanism — burny_tech · 2026-08-25
- GPT-5.6 Solves ~10% of 3,300 Open Math Problems — burny_tech · 2026-08-25
- Why Action Chunking Improves Robotic Control: New Paper by Sergey Levine — tarantulae · 2026-08-25
- Graph Rewrites Generate Labyrinths, Informing AI Models — mtizard · 2026-08-25
- FinGPT-Forecaster Open-Sourced for Stock Prediction — aigleeson · 2026-08-25