Dev releases 275k parsed Indian gov documents, plug-and-play for RAG
malliktwts · x · 2026-10-01
Developer @kingofknowwhere released an open dataset: 275k public Indian government documents parsed from 15k gov websites (322k official PDFs) across 128 topics. Raw and cleaned versions ship in CSV, Parquet, and SQLite; at least 50k pages are considered useful, and the data plugs into any RAG framework.
- No license or credits required, aimed at anyone building an Indian gov chatbot
- The author plans to scale to 20 million documents by month end, with the pipeline running autonomously on JarvisLabs
More from Research
- BlockSearch: a 0.6B in-context retriever rivals dense retrieval at million-token scale — CShorten30 · 2026-10-01
- Eval-cooperativeness alignment research wins Corrigibility Research Fund prize — dhadfieldmenell · 2026-10-01
- 0.8B model plus 9 LoRA adapters routes agent decisions 38x faster with +8.7 accuracy — Usual_Maximum7673 · 2026-10-01
- The first AI-discovered cancer drug reportedly just worked — theimposingshadow · 2026-10-01
- MedKIT Benchmark at NeurIPS 2026 Shows LLMs Can Recall Updated Facts But Fail to Use Them — zeynepakata · 2026-10-01
- SynthID Bio Watermarking Tested Only on AlphaFold 3 But Should Generalize, Author Says — davidstutz92 · 2026-10-01