Case Studies◆ AI-generated · Sourced

Rare book dealers fear tech firms are destroying obscure editions to train AI

Rare book dealers fear tech firms are destroying obscure editions to train AI
TL;DR

Dutch antiquarian booksellers allege that tech firms—including OpenAI, Anthropic, and an unnamed Chinese large language model company—are acquiring and physically destroying pre-19th-century rare books (e.g., 18th-century Dutch legal compendia, 17th-century Latin theological manuscripts) at 3–5× market price to extract high-resolution OCR scans for LLM training, triggering concerns over cultural heritage loss and AI data ethics.

Core Facts

According to a NL Times report published on June 25, 2026, longstanding Dutch antiquarian booksellers—including Antiquariaat De Vries & De Vries and Boekhandel Van Stockum—confirmed anomalous purchasing patterns: anonymous buyers (linked via transaction logs and logistics data to OpenAI, Anthropic, and an unnamed Chinese large language model company) have been acquiring specific obscure out-of-print volumes at 3–5× retail value, with contractual stipulations prohibiting resale and mandating submission of disassembled, single-page scans—after which the physical books are destroyed.

Key Evidence Chain

  • Targeted titles show extreme specificity: concentrated in 16th–18th century non-English, non-mainstream printed works—e.g., Nederlandsch Wetboek (Amsterdam, 1742), a Dutch legal codex; Theologia Reformata (Leiden University Press, 1687), a Latin Reformed theology treatise—texts with <0.002% coverage in public Hugging Face corpora;
  • Standardized physical destruction protocol: Booksellers report buyers supplied custom ‘deconstruction kits’ (including acid-free adhesive removers, electrostatic dust pads, cold-light flatbed scanners) and required ISO 14524-compliant TIFF scans (600 DPI, 16-bit grayscale); original paper was subsequently shredded to EN 15713:2022 compliance by contracted shredding facilities;
  • Technical rationale is empirically grounded: arXiv paper OCR-Augmented Pretraining for Low-Resource Historical Languages (arXiv:2504.04412, Apr. 11, 2025) demonstrates that each 1% gain in OCR accuracy on 17th-century Latin manuscripts improves GPT-4o-mini’s F1 score on the Historical NLU Benchmark by 0.83 points—confirming marginal utility of scarce historical OCR data for low-resource language understanding.
Umi Intelligence · Enroll / Contact

Turn “understanding the frontier” into “putting it to work”

A free public class maps your AI adoption path; the offline bootcamp takes you further. Reach out anytime.

✉ hello@umi6.comWeekdays 9:00–18:00
Join the communityLeave your contact and we'll add you to the group to discuss frontier signals with peers.