'Daizhige' Chinese classics dataset goes Markdown, fixing errors that polluted AI corpora
vista8 · x · 2026-10-07
A shared GitHub repo maintains the "Daizhige" collection of classical Chinese texts (Buddhist, Confucian, Daoist canons, Siku Quanshu and more; 523 stars), converting all raw TXT files to Markdown with YAML metadata and feeding a full-text search site.
Key contribution: systematic error correction. The original dataset — already used for AI training by many research institutions — contained conversion errors ("记忆体" should be "内存", "香港脚" should be "脚气", "利瓦伊" should be "李维"), pasted forum content, and leftover HTML/scripts, potentially polluting classical Chinese corpora. Actively maintained; valuable for researchers building Chinese pretraining corpora.
More from Research
- NVIDIA's VeriFine Co-Evolves Policy and Judge to Scale Self-Improvement in Embodied Reasoning — nvidia · 2026-10-07
- Single Image to Full 3D Scene: Adaptive Chunking Extends Object Generators to Outdoor Rome — Jiraphon Yenphraphai · 2026-10-07
- EmbodiedSmith: Recursive Self-Improvement Flywheel Scales Embodied Training Data in Simulation — Yikai Qin · 2026-10-07
- Video World Models Flunk Physics: Best Model Scores 57.76/100 on New 40-Task Benchmark — Mingju Gao · 2026-10-07
- Meta formalizes personal-agent mediated recommendation, releases MediateRec benchmark — meta · 2026-10-07
- Agents judge tool results useless 97-100% of the time yet rarely stop: NTU study — nanyang-technological-university-singapore · 2026-10-07