< Back to all clusters
[TECHNOLOGY] · 2 sources

started · updated

AI developers face scrutiny over book destruction and web crawling practices

Investigations have revealed that artificial intelligence developers are engaging in large-scale physical book scanning, sometimes resulting in the destruction of original copies. Documents from a lawsuit against Anthropic detailed “Project Panama,” a venture where the company reportedly purchased millions of books, removed their bindings to scan the pages, and shredded the originals. An investigation by 404 Media used an Apple AirTag to track a shipment of books to an Amazon-owned warehouse, where employees reported scanning books as a primary task. This practice has extended beyond bestsellers to include rare, antique, and out-of-print volumes.

Simultaneously, the landscape of web crawling for AI training is shifting. While much public attention and legal negotiation focus on Google and OpenAI, data from bot-defense vendor DataDome indicates that Meta’s crawlers now account for the majority of AI agent traffic. While publishers frequently negotiate licensing deals or implement blocks against Google and GPTBot, Meta’s web indexers have seen significant growth in requests without providing comparable referral traffic back to publishers. This highlights a growing tension between AI companies' need for massive datasets and the rights of content creators to control and monetize their work.

Entities

404 Media · Anthropic · DataDome · Google · Meta

Sources

about 1 month ago
about 1 month ago