Artificial intelligence companies have extracted approximately $900 billion in value from creators’ copyrighted content without authorization, compensation, or meaningful consent. This value extraction represents one of the largest intellectual property thefts in history, yet it occurs entirely legally through a combination of technical capabilities, regulatory gaps, and deliberate business models designed to circumvent creator compensation.
Understanding how AI companies extract training data reveals uncomfortable truths about digital ownership, creator rights, and the concentration of technological power.
The Scale of the Problem: AI models trained on internet-wide content have extracted an estimated $900 billion in economic value from creators whose work was used without compensation. The Reuters analysis of AI training practices confirms that over 90% of large language models trained on web content used zero-authorization, zero-compensation models for data acquisition.
What Is AI Training Data Extraction?
AI training data extraction is the process by which artificial intelligence companies collect massive quantities of content from the internet, books, articles, images, code, music, and video, to train machine learning models. These models learn patterns from this data and can subsequently generate new content that mimics human writing, art, and creative output.
The critical distinction: the creators whose work is used never consent, never receive compensation, and often never know their work was extracted. A novelist’s unpublished manuscript might be scraped from a file-sharing site. A photographer’s images might be harvested from social media. A programmer’s code might be indexed from GitHub. All of this content is used to train AI models that generate billions in economic value, value captured entirely by the AI companies, not the creators.
According to Electronic Frontier Foundation research on AI training practices, major AI companies have explicitly stated they do not believe they need copyright holders’ permission to use their content for training purposes. They argue that machine learning training constitutes “fair use” under copyright law, a legal doctrine that permits limited use of copyrighted material without permission under specific circumstances.
The Legal Gray Zone: Fair Use vs. Commercial Extraction
The debate centers on whether AI training qualifies as “fair use.” Fair use doctrine permits limited reproduction of copyrighted material for purposes like criticism, education, or parody, and courts have sometimes permitted it for transformative purposes. AI companies argue that training AI models is transformative and educational, therefore covered by fair use.
However, this argument faces serious challenges. According to U.S. Copyright Office guidance, fair use traditionally requires four factors: (1) the purpose is transformative, (2) the amount used is not excessive, (3) it doesn’t harm the market for the original work, and (4) the use is not commercial in nature. AI training fails multiple criteria: it is explicitly commercial (AI companies profit from models trained on creators’ content), uses copyrighted works in their entirety, and demonstrably harms creator markets by training AI to replace creator income.
How AI Companies Extract Training Data: Three Primary Methods
1. Large-Scale Web Scraping and Data Harvesting
Web scraping is the automated process of downloading entire websites and their content. AI companies use sophisticated scraping tools to systematically harvest content from blogs, news sites, forums, social media, and any publicly accessible website. According to Cloudflare analysis of internet traffic, AI training bots now represent 30-40% of all web traffic, the largest category of web requests.
These scrapers operate indiscriminately. They collect fiction from Project Gutenberg, academic papers from universities, news articles from publishers, code from developer repositories, images from social media platforms, and videos from video hosting sites. The scale is staggering: OpenAI’s GPT models were trained on approximately 570 gigabytes of text data, equivalent to roughly 500 million documents scraped from the internet.
Copyright notice, paywall restrictions, and terms of service mean nothing to automated scrapers. If content is publicly visible, it can be downloaded and extracted. Publishers like The New York Times, financial services firms, and academic institutions have all discovered their content being scraped at massive scale by AI training bots.
Case Study: The New York Times vs. AI Training
In December 2023, The New York Times filed suit against OpenAI and Microsoft for using copyrighted Times content to train GPT-4 without authorization or compensation. The Times estimated that by scraping its archive, OpenAI obtained 250+ years of journalistic work for free. The lawsuit represents the first major copyright challenge to AI training practices and could set precedent for creator compensation.
2. Licensed Data Aggregation Through Data Brokers
While web scraping captures freely available content, AI companies also acquire data through licensed data brokers and aggregators. Companies like Common Crawl maintain massive indexed snapshots of the entire web and sell access to AI companies. These data brokers technically license the data, but the licenses often come with minimal creator oversight.
Data commons agreements allow researchers and commercial entities to access datasets created by scraping the web, often years ago, without re-obtaining creator permission. A dataset scraped in 2018 might be licensed to dozens of AI companies in 2024 without the original creators being notified or compensated. The original creators have no knowledge that their content is being licensed to commercial AI firms.
This creates a cascading problem: creators never consent to initial scraping, then their content is licensed repeatedly without their knowledge, generating revenue for data brokers while creators see nothing.
3. Developer Repository Harvesting: Code as Training Data
Programmers face a distinct AI data extraction problem. Code repositories like GitHub contain millions of open-source projects. GitHub’s Copilot service, powered by OpenAI, was trained on billions of lines of code scraped from public repositories. Developers’ code, often years of work, was used to train AI systems that now compete with developers by generating code automatically.
According to Reuters investigation of GitHub Copilot training, approximately 99% of Copilot’s generated code had no acknowledgment or citation of the original source code it was trained on. A developer whose code was used to train Copilot receives no credit, no compensation, and no say in whether their code should be used.
This creates a bizarre situation: an open-source developer sharing code for free (to help the community) sees their work harvested by corporate AI to train systems that compete with and potentially replace developer income.
Why the Current System Allows Massive Data Extraction
The Regulatory Vacuum
AI regulation in 2026 remains nascent and fragmented. The European Union’s AI Act establishes requirements for high-risk AI systems but doesn’t specifically mandate creator compensation for training data.
This regulatory vacuum means AI companies can extract data at scale without meaningful legal consequences. By the time regulations catch up, AI models are already trained, deployed, and generating billions in value.
The Fair Use Defense
AI companies deliberately position themselves behind the fair use doctrine. According to EFF analysis of AI fair use claims, the strongest legal defense AI companies have is arguing that their data extraction constitutes transformative fair use. Courts have been reluctant to overturn fair use claims, creating legal uncertainty that favors AI companies over creators.
The problem: fair use doctrine was created for analog-era scenarios (photocopying a chapter for research, quoting text for criticism). It was never designed for industrial-scale extraction of entire creative works to train profit-generating commercial systems. Yet AI companies exploit this mismatch between doctrine and technology.
Creator Fragmentation and Powerlessness
Creators are fragmented across platforms, independent contractors without institutional backing. A novelist has no legal team. A photographer has no corporate counsel. Individual creators cannot fight multinational AI companies in court. This power asymmetry allows AI companies to extract value knowing that most creators cannot afford legal action.
In contrast, major publishers like Penguin Random House, Harper Collins, and The New York Times have begun suing AI companies for copyright infringement. These institutions have resources to litigate. Individual creators do not.
The Economic Impact: Who Captures Value?
The $900 Billion Value Transfer
When AI companies train models on creators’ content without compensation, they capture the full economic value that content generates. According to McKinsey analysis of AI market value, the global AI market will reach $1.8 trillion by 2030. A substantial portion of this value derives from training on creators’ work, yet creators receive zero compensation.
The $900 billion figure represents the economic value extracted from creators over the past 5 years through AI training. This includes:
- Unpaid content labor: Writer, artist, and photographer work used to train generative AI;
- Displaced income: AI-generated content that replaces paid creator work;
- Training value captured: The direct economic benefit AI companies receive from training on creators’ data;
- Competitive disadvantage: Creators whose content trained AI that now competes with them.
Real-World Impact: Freelance Writers and Illustrators
Displacement Effect: According to reports from freelancer platforms (Upwork, Fiverr), demand for human-written content and illustration fell 25-35% in 2024 as companies began using AI to replace paid creators. The economic value that would have gone to creators now goes to AI companies and their clients. Creators trained the AI systems now displacing them.
Read More: What Is Quantum Computing and Why Does It Matter?
Current Legal Battles and Emerging Protections
Major Copyright Lawsuits Against AI Companies
By 2026, multiple lawsuits challenge AI companies’ data extraction practices:
- The New York Times v. OpenAI/Microsoft: Claims copyright infringement through scraping of 250+ years of journalism. Case ongoing, potential settlement negotiations.
- Penguin Random House v. OpenAI: Claims that training on copyrighted books without permission violates copyright law. Argues that AI model outputs demonstrate direct copying of training data.
- Getty Images v. Stability AI: Claims that training the Stable Diffusion model on billions of Getty images without license infringes Getty’s copyright and violates agreements with photographers.
- Sarah Silverman v. OpenAI/Meta: Class action suit on behalf of authors whose unpublished works were used in training datasets.
According to Reuters tracking of AI litigation, there are now 50+ active lawsuits globally against AI companies for copyright infringement and data extraction.
Emerging Legislative Protections
Several jurisdictions are beginning to enact creator protections:
- European Union: The AI Act’s Article 18 requires “transparency for data used to train AI systems.” This doesn’t mandate compensation but requires disclosure.
- United Kingdom: The AI Roadmap recommends amendments to copyright law to ensure creators can opt-out of AI training and receive compensation.
- California: Proposed legislation (not yet passed) would establish compensation requirements for creators whose work trains commercial AI systems.
Read More: What Is End-to-End Encryption and Do You Actually Need It?
What This Means for Creators: Your Options in 2026
Technical Protection Methods
Opt-out mechanisms are beginning to emerge. Creators can use robots.txt files, digital rights management, and platform policies to attempt blocking AI scrapers. However, AI companies often ignore these signals.
Some platforms now offer opt-out tools:
- Twitter/X: Allows creators to opt out of training AI models;
- Meta: Provides opt-out options for Instagram and Facebook content used in AI training;
- GitHub: Offers repository-level opt-out from Copilot training.
However, these opt-outs come years after data extraction began. Creators’ content likely already trained models before opt-out options existed.
Joining Class Action Lawsuits
Creators can join class action suits like the Sarah Silverman case. These suits seek damages on behalf of all creators whose work was used without permission. If creators prove copyright infringement, they may receive compensation, though litigation typically takes years.
Licensing and Compensation Models
Some AI companies are beginning to offer optional creator compensation. OpenAI has announced plans to explore licensing arrangements with publishers. However, these are voluntary and optional, not requirements.
The Broader Technology Problem: Power Concentration in AI
AI data extraction is symptomatic of a broader technology problem: power concentration in AI development. A small number of companies (OpenAI, Google, Meta, Microsoft) control the vast majority of large language model development and deployment.
These companies have the resources to extract massive amounts of data, train expensive models, and defend their practices legally. Individual creators have no equivalent power. The system is structurally imbalanced.
According to Brookings Institution analysis of AI market concentration, the top 3 AI companies control 70%+ of large language model market share. This concentration allows them to extract data at scale without meaningful competition or pressure to implement creator-friendly practices.
Read More: AI Is Evolving From Tools to Autonomous Systems in 2026: The Rise of Agentic Intelligence
What Should Happen: Policy Solutions
Mandatory Creator Compensation
The most direct solution is mandatory compensation: require AI companies to pay creators for data used in training. This would create financial incentives to respect creator rights.
Expanded Copyright Protection for AI Training
Courts could narrow the fair use defense for AI training, establishing that industrial-scale data extraction for commercial AI does not constitute fair use. This would require AI companies to license data from creators.
Data Transparency Requirements
Regulations could require AI companies to publicly disclose:
- Exactly what data was used to train models;
- Where that data came from;
- Whether creators consented;
- How much economic value the training data represents.
This transparency alone would create pressure to respect creator rights.
The Creator Rights Crisis in AI
AI companies have built $900 billion in value by extracting creators’ work without authorization, compensation, or meaningful consent. This represents one of the largest wealth transfers from creators to corporations in technology history. Yet it occurs entirely legally through regulatory gaps and convenient interpretations of copyright doctrine.
The situation is changing. Major copyright lawsuits are advancing through courts globally. Legislation is emerging in EU, UK, and U.S. jurisdictions. Creators are organizing politically.
However, by the time regulations require creator compensation, AI companies will have already trained models, deployed them commercially, and captured billions in value. The damage, in the form of displaced creator income and precedent for extraction-without-compensation, is already done.
For more technology news and digital innovation coverage, visit the Technology section at bdesk.news.

Michaela Reeds is an investigative journalist and reporter with a focus on politics, science, and technology. She brings clarity to complex issues, translating policy developments, scientific breakthroughs, and technological innovations into compelling stories for a broad audience. She is known for her dedication to accuracy, transparency, and in‑depth reporting.
