Rare book dealers fear tech firms are destroying obscure editions to train AI

Dutch antiquarian booksellers allege that tech firms—including OpenAI, Anthropic, and an unnamed Chinese large language model company—are acquiring and physically destroying pre-19th-century rare books (e.g., 18th-century Dutch legal compendia, 17th-century Latin theological manuscripts) at 3–5× market price to extract high-resolution OCR scans for LLM training, triggering concerns over cultural heritage loss and AI data ethics.
Core Facts
According to a NL Times report published on June 25, 2026, longstanding Dutch antiquarian booksellers—including Antiquariaat De Vries & De Vries and Boekhandel Van Stockum—confirmed anomalous purchasing patterns: anonymous buyers (linked via transaction logs and logistics data to OpenAI, Anthropic, and an unnamed Chinese large language model company) have been acquiring specific obscure out-of-print volumes at 3–5× retail value, with contractual stipulations prohibiting resale and mandating submission of disassembled, single-page scans—after which the physical books are destroyed.
Key Evidence Chain
- Targeted titles show extreme specificity: concentrated in 16th–18th century non-English, non-mainstream printed works—e.g., Nederlandsch Wetboek (Amsterdam, 1742), a Dutch legal codex; Theologia Reformata (Leiden University Press, 1687), a Latin Reformed theology treatise—texts with <0.002% coverage in public Hugging Face corpora;
- Standardized physical destruction protocol: Booksellers report buyers supplied custom ‘deconstruction kits’ (including acid-free adhesive removers, electrostatic dust pads, cold-light flatbed scanners) and required ISO 14524-compliant TIFF scans (600 DPI, 16-bit grayscale); original paper was subsequently shredded to EN 15713:2022 compliance by contracted shredding facilities;
- Technical rationale is empirically grounded: arXiv paper OCR-Augmented Pretraining for Low-Resource Historical Languages (arXiv:2504.04412, Apr. 11, 2025) demonstrates that each 1% gain in OCR accuracy on 17th-century Latin manuscripts improves GPT-4o-mini’s F1 score on the Historical NLU Benchmark by 0.83 points—confirming marginal utility of scarce historical OCR data for low-resource language understanding.