'Daizhige' Chinese classics dataset goes Markdown, fixing errors that polluted AI corpora

vista8 · x · 2026-10-07

A shared GitHub repo maintains the "Daizhige" collection of classical Chinese texts (Buddhist, Confucian, Daoist canons, Siku Quanshu and more; 523 stars), converting all raw TXT files to Markdown with YAML metadata and feeding a full-text search site.

Key contribution: systematic error correction. The original dataset — already used for AI training by many research institutions — contained conversion errors ("记忆体" should be "内存", "香港脚" should be "脚气", "利瓦伊" should be "李维"), pasted forum content, and leftover HTML/scripts, potentially polluting classical Chinese corpora. Actively maintained; valuable for researchers building Chinese pretraining corpora.

Original post →

More from Research

Research channel →