LegalPincite: Multi-level Legal IR Dataset Addresses Paragraph-Level Citation Gaps
Theresia Veronika Rampisela · hf · 2026-08-06
Most existing public legal Information Retrieval (IR) datasets lack paragraph-level citation annotations. Furthermore, some available datasets suffer from data leakage in query text and exclude non-citing paragraphs, creating unrealistic retrieval settings that inflate performance.
To address this, researchers constructed LegalPincite, a large-scale legal IR dataset based on Court of Justice of the European Union (CJEU) judgments. The dataset features:
- Masked case/paragraph queries with citation info removed
- A comprehensive corpus including all paragraphs
- Case- and paragraph-level ground-truth citations with partial human expert validation
This dataset supports the development and rigorous evaluation of legal IR methods across multiple query-document levels (case-to-case, paragraph-to-case, and paragraph-to-paragraph).
More from Research
- Nature Paper: AI Model Forecasts Personal Health Trajectories Up to 20 Years — EricTopol · 2026-08-06
- Continual Learning Bench: Simple Context Memory Beats Expensive Dedicated Systems — ajratner · 2026-08-06
- Goodfire AI's MAPS Explains 2.1 Million Genetic Variants Mechanistically — mathildepapillo · 2026-08-06
- New Theory Explains the Effectiveness of Stop-Gradient in Flow Models — kwangmoo_yi · 2026-08-06
- Fudan Researchers Show AI Models Can Autonomously Self-Replicate Like Worms — willknight · 2026-08-06
- Evaluating LLM Sycophancy: Which Models Hold Their Ground? — zero0_one1 · 2026-08-06