Claim That 8% of The Pile Is Christian Content Challenged: Keyword Method Miscounts Long Books
cis_female · x · 2026-09-18
An ICMI working paper by Tim Hwang estimated, via strict keyword classification on a 100,000-document sample of The Pile (825B tokens), that roughly 67 billion tokens (8.1%) are explicitly Christian content — 32x the Islamic share, 19x Buddhist, 45x Hindu, 93x Jewish, and equivalent to 15x English Wikipedia — proposing a "Christian Prior" hypothesis that frontier models' moral reasoning is by default shaped more by the Christian tradition than any other ethical framework.
However, @cisfemale challenges the methodology: Hwang counts a source as Christian if it hits a keyword threshold, and 90% of his "Christian" sources are just long books that mention Christianity in passing, not Christian books. If correct, the 8.1% estimate is significantly inflated, undermining the paper's quantitative case that the alignment community is neglecting a massive Christian prior in pretraining data.
More from Research
- Frank Nielsen publishes paper on two types of geometric Jensen–Shannon divergences — FrnkNlsn · 2026-09-18
- Periodic Labs details its stack: 4.1x Megatron throughput, frontier-beating science models on 1,300 H200s — hsu_byron · 2026-09-18
- Video DeltaNet hybrid attention speeds livestream video generation 14.5x on 8x B200 — Haocheng Xi · 2026-09-18
- JEPA-Anything: One Predictive Framework Spans Vision, Biology, Weather and More — Taoyong Cui · 2026-09-18
- Tencent's WeVisDoc tops OmniDocBench with 95.38 via two-stage data-centric training — tencent · 2026-09-18
- RetireOPD: self-retiring on-policy distillation lifts agent RL success rates by up to 18.8% — Yan Yu · 2026-09-18