Hadith NLP entered the LLM era; this review says the bottleneck is evidence infrastructure

Hadith computational science in the age of large language models: a critical narrative review

Md. Ashraful Haque, Riasat Islam

Greentech Apps Foundation, United Kingdom / Queen Mary University of London, London, United Kingdom

cs.CL, cs.AI

2026-06-18

32-study review: segmentation is fairly mature, narrator identity and transfer stay weak. LLMs changed scale and workflow; treat it as evidence infrastructure, not isolated scores.

What problem this solves

Hadith computation grew out of rule systems and small corpora. Azmi et al.'s 2019 survey mapped segmentation, narrator extraction, ontologies, retrieval, and authentication support. Transformers and LLMs then lifted publication counts. Recent reviews mostly count papers and tag methods. A different question is still open: which results travel, which are bound to one benchmark, and which failures still block scholarly use.

This paper is a critical narrative review, not a PRISMA census. Searches ran from late 2025 into early 2026 across Google Scholar, Scopus, ACL Anthology, SpringerLink, ScienceDirect, and arXiv. The working list started at 42 records and narrowed to 32. Appraisal dimensions are fixed in advance: canonical-only versus diverse corpora, cross-collection transfer, reproducibility, expert involvement, and whether the task is a proxy or a real scholarly bottleneck.

Method

There is no new model. The review splits the pipeline into three levels. Level 1 is isnad-matn segmentation, separating the chain of transmission from the report text. Level 2 covers narrator disambiguation, source verification, question answering, and authentication-related modeling. Level 3 covers knowledge graphs, corpus-scale enrichment, and research infrastructure.

Primary studies are scored on six axes rather than one league table. Altammami et al. give an end-to-end segmentation benchmark on the six canonical books, with some cross-book testing, still canonical-bound. Muther and Smith put boundary ambiguity in long classical texts into the annotation scheme. Mahmoud et al. turn narrator disambiguation into a trainable task that looks strong on artificial sanads and weaker on real ones. Mghari et al.'s Sanadset holds 650,000 narrations from 926 books: a diversity gain, not a transfer solution. Mubarak et al. make faithfulness and hallucination testable on Qur'an and Hadith QA. Asgari-Bidhendi et al. run an LLM pipeline over 1.2 million narrations with six domain experts scoring the work. That is an infrastructure jump. Reproducibility is marked no.

Results

Progress is uneven, and older methods were not simply replaced. Rule and hybrid pipelines still compete on stable segmentation. Transformers raise the ceiling on segmentation and retrieval. Knowledge graphs and retrieval grounding push the field toward inspectable evidence. LLMs make corpus-scale enrichment, multilingual access, and grounded evaluation plausible. Level 1 matured fastest because the task is clean and annotations reuse. Level 2 got broader without converging. Level 3 moved from local graphs to pipeline-scale enrichment.

Four structural gaps stay hard. Benchmark culture still concentrates on the six canonical Sunni collections, which are well edited and regular, so scores look good and transfer is overestimated. The commentarial layer, takhrij reasoning, and fiqh al-hadith are barely computational objects. Links from hadith to Qur'an, seerah, and biographical literature are mostly pairwise and task-specific; a multi-hop evidence graph is missing. The bridge into contemporary jurisprudence is framed as evidence support. Automated fatwa issuance is the wrong target.

LLMs changed three things. Scale: systems such as Rezwan make processing hundreds of thousands to millions of narrations imaginable. Workflow: OCR repair, segmentation, retrieval, checking, simplification, semantic tagging, and human review get chained. Proof: a high score on a tidy corpus is no longer enough. They did not solve narrator identity resolution, brittle preprocessing, or the gap between fluent generation and reliable scholarship.

Why it matters

Anyone working on religious text, classical low-resource language, or high-stakes NLP gets a stricter way to read this literature. Publication growth is not task maturity. Accuracy on synthetic chains does not translate into real narrator disambiguation. Expert-scored pipelines are not commensurate with classic segmentation F1. The authors recast success as evidence infrastructure: provenance, expert-in-the-loop review, and uncertainty reporting are part of system quality.

The product implication is equally concrete. Islamic QA and fatwa generation are already moving. A system that returns one narration without provenance, variants, or commentary is not serious jurisprudential support. Use these methods for enrichment, grounded retrieval, and assisted segmentation. Do not use them for unsupervised authentication or automatic rulings.

Limitations

Narrative reviews carry search and selection bias, and the authors say so: indexed English-heavy venues are favored, weakly indexed Arabic venues and gray literature drop out. Thirty-two refined records cannot support a statistical map of the whole field. Claims that the six books dominate, or that commentary is missing, are inferences from a representative sample. Scholar-perspective papers are mostly conceptual and rest on small expert pools. The authors refuse to treat any one paper as "the" Islamic scholarly view.

There is no new experiment, so the appraisal table cannot be independently re-run. Preprints stay in the narrative when they shape later infrastructure talk, which lets unreviewed claims into the main story. The research agenda on expert governance and fiqh-facing evidence support is clear. Shared benchmarks, open data paths, and evaluation protocols remain proposals.

Terms

Source

Related papers

All paper explainers