started · updated
AI companies use destructive scanning of rare books for training
Artificial intelligence companies are reportedly engaging in “destructive scanning” to acquire rare training data. As digital knowledge available online becomes saturated with AI-generated content, tech firms are increasingly seeking authentic, non-digitized sources to train large language models.
Reports indicate that companies are placing massive, erratic bulk orders at secondhand bookstores across the United Kingdom, Ireland, the United States, Canada, and mainland Europe. Unlike traditional collectors, these buyers request highly varied titles—ranging from niche historical biographies to specific foreign-language editions—without seeking bulk discounts.
Investigations, including one involving an Apple AirTag used to track a shipment, suggest that books are being purchased, scanned, and then destroyed or recycled. This practice, known as destructive scanning, allows companies to claim “fair use” under U.S. law because the original works are not being redistributed digitally. Previous court documents related to Anthropic have already highlighted similar practices where books were stripped of their spines and scanned before being recycled.