started · updated
AI training demand drives new market for data and archives
A new market is emerging driven by the demand for data to train artificial intelligence models. Companies are increasingly monetizing previously ignored archives, such as out-of-print books, historical newspapers, emails, and internal corporate communications.
In a notable recent transaction, Google offered $10 million for internal data from Spirit Airlines to improve its AI products. The dataset includes approximately 100 million emails, 500 million Microsoft Teams elements, and various internal documents and source code. The startup Micro1 reportedly submitted a subsequent bid of $12.5 million, while Mercor had previously offered $7.5 million. Google has stated that the agreement includes a de-identification process to ensure no personal information is accessed.
Simultaneously, the media industry faces structural shifts as reported by the Reuters Institute Digital News Report 2026. Traditional news consumption is being disrupted by AI chatbots, short-form videos, and content creators. An increasing number of users are utilizing conversational assistants like ChatGPT and Perplexity as primary gateways to information, receiving synthesized answers rather than browsing multiple news links.
Entities
Google · Mercor · Micro1 · Reuters Institute · Spirit Airlines