Zum Inhalt springen

Democratizing Intellectual Property for the AI Age

Dieser Artikel ist auf Englisch.

TL;DR

AI companies are running out of training data, courts are split on fair use, and the entire intellectual property system is buckling under the weight of generative AI. This article explores two paths forward: complete IP abolition versus AI-specific exceptions. The evidence suggests neither extreme works. Instead, we need hybrid solutions that balance creator compensation with AI advancement. With data exhaustion projected by 2026-2032 and over 50 lawsuits piling up, this isn’t theoretical anymore. It’s happening now.


Introduction

Here’s a number that should terrify every AI company: 2026.

That’s when we might run out of human-generated text to train AI models. Not because people stopped writing. Because AI has already consumed nearly everything we’ve ever written on the public internet. Epoch AI estimates we’ve got maybe 300 trillion tokens of usable human text total. GPT-4 alone trained on 13 trillion. Do the math.

Meanwhile, the New York Times is suing OpenAI. Universal Music is negotiating equity stakes. Anthropic just settled for $1.5 billion. And over 50 copyright lawsuits are working their way through courts that are reaching wildly different conclusions about whether training AI on copyrighted work is even legal.

The question isn’t whether the intellectual property system will change. It’s how. And the stakes are massive. IP-intensive industries make up 47% of EU GDP (around €6.4 trillion) and employ 81 million people. [1] Yet five Nobel Prize-winning economists argue that copyright extensions provide „no additional incentive“ for creation while creating „devastating effects“ on cultural access. [2]

So what happens when an unstoppable force (AI’s need for training data) meets an immovable object (centuries of copyright law)? Let’s find out.


The Pre-Copyright World Wasn’t Great, Actually

Before we get too romantic about abolishing intellectual property, let’s talk about what the world looked like before copyright existed.

Prior to the Statute of Anne in 1710, creative production ran on patronage. Rich families and religious institutions paid creators salaries in exchange for work. Sounds fine until you realize this meant art served „narrow elite audiences.“ If a duke didn’t like your novel about peasant life, tough luck. No duke, no book.

The Statute of Anne did something radical: it vested rights in authors rather than publishers and created the concept of a public domain where works would „become freely available for reproduction“ after fixed terms. This wasn’t just legal innovation. It fundamentally changed who could create and for whom.

Full IP abolition would return us to something like the pre-copyright era, but in a radically different technological context. And the economic consequences would be wildly uneven across industries.

Pharmaceuticals Would Get Destroyed (And That Matters)

Patents are genuinely essential for drug development. The average cost to bring a new drug to market is $2.6 billion. Without patent protection to recoup that investment, a Carnegie Mellon survey found that 60% of drug inventions and 38% of chemical inventions would not have been developed. [3]

But here’s the twist: in most other industries, patents don’t matter nearly as much. The same research found that „the vast majority of inventions would have been developed without patent protection.“ Software firms report patents as „the least effective source of competitive advantage.“ [4]

The fashion industry is the perfect natural experiment. Operating under weak design protection, it generates $350 billion in US GDP [5] and $2 trillion globally [6] while maintaining robust innovation through what scholars call the „piracy paradox.“ Copying actually accelerates trend cycles and creates demand for new designs. [7]

So full abolition would crater pharmaceuticals while barely touching fashion. That’s not a balanced policy. That’s chaos with winners and losers determined by industry accident.

Consumers Would Win Short-Term, Maybe Lose Long-Term

Generic drug entry after patent expiration typically reduces prices by 38-48%, with prices dropping up to 80% when ten or more competitors enter. [8][9] That’s real money back in real people’s pockets.

The public domain has demonstrably enabled cultural flourishing. Disney built its empire on Alice in Wonderland, Pinocchio, and The Little Mermaid, all works whose copyrights had lapsed. But the superstar dynamics of creative markets mean most creators already earn almost nothing. In music streaming, artists receive approximately 16% of recording revenue while labels capture 64%. [10] Major labels earn roughly $1 million per hour from streaming platforms. [11]

Would full abolition make that better or worse? Honestly, it’s hard to see how individual creators get more leverage when their work becomes freely copyable by corporations with distribution advantages.

AI Gets Legal Clearance But Not More Data

Full abolition would remove all legal barriers to training data acquisition. Great news for AI labs‘ legal departments. But it wouldn’t solve their actual problem: there isn’t enough human-generated content to keep scaling models.

Epoch AI projects data exhaustion by 2026-2032. [12] Training on synthetic data (AI-generated content) risks „model collapse“ where performance degrades across generations. [13] Anthropic CEO Dario Amodei estimates a 10% chance we could run out of data to continue scaling. [14] Goldman Sachs‘ Chief Data Officer stated in October 2025 that „we’ve already run out of data.“ [15]

Removing copyright doesn’t magically create more training data. It just makes the existing finite pool legally accessible.


The Middle Path: AI-Specific Exceptions

So if full abolition is a mess, what about creating carve-outs specifically for AI training while maintaining other IP protections? This isn’t hypothetical. Multiple countries have already tried it.

Japan: The Most Permissive Regime

Japan’s Copyright Act Article 30-4 allows AI training „to the extent considered necessary“ for both commercial and non-commercial use, even on pirated material. [16] The only restriction is that uses cannot „unreasonably prejudice interests of copyright owner.“ Minister Nagoaka confirmed that AI companies can use „whatever they want“ regardless of source. [16]

This is basically the AI industry’s dream scenario. And it’s real, right now, in the world’s third-largest economy.

Europe: Conditional Access With Opt-Outs

The EU takes a more balanced approach through the DSM Directive, permitting text and data mining but allowing rights holders to opt out through „machine-readable means.“ [17] The EU AI Act, effective August 2025, requires general-purpose AI providers to publish „sufficiently detailed summaries“ of training content and respect opt-out reservations, with extraterritorial reach regardless of where training occurred. [18][19]

Translation: you can train on EU citizens‘ data even if you’re in California, but you need to disclose what you used and honor opt-outs.

United States: A Beautiful Mess

The US remains litigation-dependent, and courts are splitting in fascinating ways.

In June 2025, a Northern California court ruled in Bartz v. Anthropic that LLM training is „spectacularly transformative“ fair use when it extracts statistical patterns rather than expressive content. [20] But critically, the court held that using pirated copies „for ANY purpose, including training, destroys fair use defense.“

This suggests a potential resolution: lawful sourcing as a threshold requirement, enabling a market for training data licenses. You can transform copyrighted work for AI training, but you need to obtain it legally first. That’s a rule AI companies can actually work with.

Meanwhile, the Copyright Office’s May 2025 report concluded that AI training on copyrighted works constitutes prima facie infringement but that fair use analysis remains fact-specific. [21] They recommended against compulsory licensing based on stakeholder consensus, preferring voluntary market development.

Translation: figure it out yourselves, and good luck.


What the Evidence Actually Says About IP and Innovation

Let’s get uncomfortable for a minute. The academic evidence on whether intellectual property actually promotes innovation is way more contested than policy debates typically acknowledge.

A survey of empirical research reveals that patents are „the least important of the major appropriability mechanisms“ across most industries, with firms ranking lead time and learning curves higher. [22] Josh Lerner’s study of 60 countries from 1850-1999 found patent strengthening had no effect on innovation measured by British patenting, with domestic applications actually decreasing after strengthening. [23] Japanese reforms in 1988 that strengthened patents had „no discernible impact“ on R&D expenditures. [24] Petra Moser’s historical research shows countries without patents contributed substantial innovation at 19th-century World Fairs. [25][26]

If patents don’t drive innovation, what does?

The Open Source Miracle

The open source movement is the strongest evidence that innovation can flourish without traditional IP-based revenue models. A 2024 Harvard Business School study quantified open source’s demand-side value at $8.8 trillion, what it would cost firms to build equivalent software internally. [27] Open source code now appears in 96% of commercial codebases, [28] with just 5% of developers responsible for 96% of this value creation. [29]

The mechanisms enabling this aren’t charity. They’re reputation building, career advancement, direct utility from improvements, and corporate sponsorship of development that serves strategic interests.

Creative Commons has scaled to nearly 2 billion licensed works globally, [30] including all of Wikipedia’s 42.5 million articles. The Metropolitan Museum of Art released 375,000 digital works to the public domain via CC0 licenses, expanding rather than contracting cultural access.

But here’s the critical insight: open source and Creative Commons operate as supplements to the IP system, not replacements. Creators retain copyright and choose permissive licensing, often while maintaining proprietary revenue streams elsewhere. The GPL license, open source’s legal foundation, actually depends on copyright to enforce its sharing requirements.

This isn’t proof that IP abolition would work. It’s proof that hybrid models can work.


The Global Legal Landscape Is Fracturing

As of October 2025, over 51 copyright lawsuits have been filed against AI companies. [31][32] The New York Times case against OpenAI is proceeding to trial after a judge rejected dismissal motions. [33] Anthropic settled for $1.5 billion in September 2025, the first major AI training copyright settlement establishing a compensation benchmark. [34][35]

This jurisdictional fragmentation creates regulatory arbitrage opportunities. Training could occur in permissive Japan, deployment worldwide. The EU AI Act’s extraterritorial provisions attempt to address this, but enforcement across borders remains challenging. [18]

China’s State-Subsidized Advantage

China’s approach deserves attention given its AI ambitions. The 2023 GenAI Interim Measures mandate „legally sourced training data,“ [36] and forthcoming comprehensive AI legislation will further define the framework. But collective copyright management organizations remain underdeveloped, and the government is encouraging development of „common data resource databases“ to reduce procurement costs, potentially creating state-subsidized training data advantages. [37]

If China provides its AI companies with cheap, legally cleared training data while US companies pay market rates or fight lawsuits, that’s a competitive advantage measured in billions of dollars and years of development time.


Historical Precedents: What Actually Works

Pharmaceutical compulsory licensing offers the most developed precedent for forcing access to intellectual property while maintaining some creator compensation.

India’s Compulsory Licensing Experiment

The 2012 Natco-Bayer case granted India’s first compulsory license for Nexavar, a cancer drug priced at INR 2.8 lakh (around $3,350/month) that reached only 2% of patients. Natco’s generic was licensed at INR 8,800 (around $105/month), a 97% price reduction, with Bayer receiving a 6% royalty. [38]

India subsequently became the „pharmacy of the Global South,“ with pharmaceutical exports comprising 5.7% of total exports in 2022-23. That’s a real-world example of compulsory licensing enabling both access and industry development.

But the COVID-19 TRIPS waiver debate revealed the political economy limits of such reforms. India and South Africa’s October 2020 proposal for comprehensive IP suspension took nearly two years to resolve, and the final June 2022 waiver covered only vaccine patents, excluding therapeutics and diagnostics, by which point 12 billion doses had already been administered. [39] Critics called it „a mere diplomatic compromise with little practical consequence,“ as only 28% of people in low-income countries had received even one dose by March 2023.

Patent Pools for Standard-Essential Patents

Standards like USB, LTE, Wi-Fi, and 5G involve hundreds or thousands of patents from multiple holders, licensed collectively under FRAND (Fair, Reasonable, and Non-Discriminatory) terms. The Fair Standards Alliance comprises 45+ companies with aggregate turnover exceeding €2 trillion. [40]

This model reduces litigation and transaction costs while ensuring patent holders receive „appropriate remuneration.“ It’s a proven mechanism for managing complex IP in technology standards. Could something similar work for AI training data?


Practical Mechanisms for AI Training

Several concrete proposals have emerged, ranging from collective licensing to public commons.

The RSL Collective: ASCAP for the Internet

Launched September 2025, the Really Simple Licensing Collective adapts the ASCAP/BMI music model for digital content. Backed by Reddit, Yahoo, Medium, and Quora, it uses an XML-based protocol for machine-readable licensing. [41][42] Publishers add directives to robots.txt files; the collective negotiates terms and collects royalties. [42]

Reddit already receives an estimated $60 million annually from Google for training data. [43] The challenge: no major AI company has agreed to comply. [44]

Compulsory Licensing With Regulatory Rates

A December 2025 European Parliament study endorsed compulsory licensing with regulatory-set royalty rates as „welfare-superior,“ projecting $14 billion more annual welfare than alternatives. [45] ASCAP and other performance rights organizations strongly oppose this, citing historical „price-suppression“ under existing compulsory schemes. [45][46] The US Copyright Office also recommended against it. [21]

The debate mirrors century-old fights over music royalties. Unsurprisingly, rights holders prefer voluntary negotiation where they have leverage.

Public Data Trusts

Academics have proposed „scraping the internet as a digital commons“ and licensing to commercial developers for revenue percentages. [47][48] France’s national AI initiative pursues this direction with publicly accessible training datasets, while the EU’s ALT-EDIC consortium is building language resources with democratic oversight. [49]

This is the „nationalize the training data“ approach. Whether it would accelerate or slow AI development depends entirely on governance and access terms.

Transparency Requirements

The TRAIN Act would create subpoena mechanisms for copyright holders to discover whether their works were used, with non-compliance creating a „rebuttable presumption“ of infringement. [50] The EU AI Act already mandates „sufficiently detailed summaries“ of training content. [18][51]

Transparency commands the broadest political support because it doesn’t pick winners. It just makes the game visible.


The Path Forward: Navigating Real Tradeoffs

The evidence doesn’t support either maximalist position. Full IP abolition would deliver short-term access gains while potentially degrading long-term creative production, particularly in sectors like pharmaceuticals where patents genuinely drive investment. But current copyright terms (life plus 70 years) far exceed any plausible innovation incentive, as five Nobel laureate economists argued in their amicus brief against retroactive term extensions. [2]

For AI specifically, the practical question is less about theory than about building workable institutions.

The most promising path forward likely involves:

Transparency requirements enabling rights holders to know when their works are used, bringing AI companies to negotiation tables.

Collective licensing mechanisms reducing transaction costs that make individual licensing impossible at scale.

Technical attribution systems allowing compensation to flow to influential training sources.

Transition support for creators displaced by AI, whether through licensing revenues, universal basic income, or other mechanisms.

What remains genuinely uncertain is whether voluntary markets will develop adequate compensation or whether statutory intervention becomes necessary. [52] The Anthropic settlement establishes a $1.5 billion benchmark, [35][34] and major record labels are negotiating deals including equity stakes. [34][53] But these represent accommodation between large entities. Individual creators and small publishers have far less bargaining power.

The deeper question, whether AI-generated abundance devalues human creativity in ways copyright cannot address, may require policy responses beyond intellectual property entirely. As the EFF notes, „copyright is not a helpful framework for addressing concerns about automation reducing the value of labor.“ [54][55] The Authors Guild warns that without guardrails, „important, diverse voices“ will be „inevitably shut out.“ [56] Both may be right.


Conclusion: Beyond Binary Thinking

The democratization debate too often frames IP as binary (protection or abolition) when the empirical evidence points toward nuanced, sector-specific optimal policies. Patent protection matters enormously for pharmaceuticals and chemicals, marginally for most other industries. Copyright terms far exceed innovation incentives but fund established media companies‘ opposition to reform. Open source demonstrates that major innovation can occur without traditional IP revenue, while also depending on copyright’s legal infrastructure.

For generative AI, the immediate practical question is whether training on copyrighted works constitutes fair use. Courts are splitting: Bartz v. Anthropic endorsed „spectacularly transformative“ fair use while holding that pirated sources destroy the defense; [57] Thomson Reuters v. ROSS found AI use NOT transformative when serving similar purposes to originals. [58] The distinction may prove decisive. AI that creates genuinely new capabilities may qualify for fair use while AI that simply automates existing creative labor may not.

The data exhaustion timeline of 2026-2032 adds urgency. [12] If AI advancement requires more human-generated content than currently exists, the terms on which that content becomes available will shape the technology’s trajectory and who benefits from it.

The choice isn’t between innovation and creator welfare. It’s about designing institutions that achieve both. Historical precedents from patent pools to pharmaceutical licensing to Creative Commons demonstrate this is possible. Whether we achieve it depends on policy choices being made right now, in courtrooms and legislatures around the world.

We’re not going to solve this with a tweet-length take. But we might solve it with the humility to learn from what’s actually worked before.


References

[1] EUIPO – IPR-intensive industries and economic performance in the European Union

[2] University of Chicago Booth School – Seventeen economists‘ amicus brief on copyright extensions (Eldred v. Ashcroft)

[3] NBER – Carnegie Mellon survey on patent incentives

[4] Berkeley Patent Survey – High technology entrepreneurs and the patent system

[5] Fordham University Law Review – US fashion industry GDP

[6] Global fashion industry statistics – $2.4-2.5 trillion valuation

[7] The Piracy Paradox: Innovation and intellectual property in fashion design

[8] NBER – Patent expiration and pharmaceutical prices (38-48% price reductions)

[9] Effect of competition on generic drug prices (80% with 10+ competitors)

[10] Billboard – Music streaming royalty payments explained

[11] Music Business Worldwide – Major labels earn over $1 million per hour from streaming

[12] Epoch AI – Will we run out of data? Limits of LLM scaling (2026-2032)

[13] Model collapse explained – How synthetic training data breaks AI

[14] BigDATAwire – Are we running out of training data for GenAI? (Dario Amodei quote)

[15] Slashdot – AI has already run out of training data, Goldman’s data chief says

[16] Privacy World – Japan’s new draft guidelines on AI and copyright

[17] Kluwer Copyright Blog – The new copyright directive: Text and data mining (Articles 3 and 4)

[18] Clifford Chance – Copyright compliance under the EU AI Act for GPAI model providers

[19] IAPP – The EU AI Act and copyrights compliance

[20] Reed Smith – A new look at fair use: Anthropic, Meta, and copyright in AI training

[21] Skadden – Copyright Office weighs in on AI training and fair use (May 2025)

[22] NBER – Protecting their intellectual assets: Appropriability conditions and why U.S. manufacturing firms patent (or not)

[23] NBER – Patent protection and innovation over 150 years (Josh Lerner)

[24] NBER – Do stronger patents induce more innovation? Evidence from 1988 Japanese patent reforms

[25] NBER – How do patent laws influence innovation? Evidence from 19th-century World Fairs (Petra Moser)

[26] NBER – Petra Moser’s research on innovation without patents

[27] Sysdig – The hidden economy of open source software (Harvard Business School $8.8T study)

[28] Harvard Business School – The value of open source software

[29] Open source developer value creation – 5% create 96% of value

[30] Infojustice – Creative Commons state of the commons (nearly 2 billion works)

[31] The Authors Guild – Status of all 51 copyright lawsuits against AI (October 2025)

[32] Chat GPT Is Eating the World – Status of all 51 copyright lawsuits v. AI

[33] NPR – Judge allows New York Times copyright case against OpenAI to go forward

[34] Major record labels seek licensing fees and equity stakes in AI music companies

[35] NPR – Anthropic to pay authors $1.5 billion in settlement

[36] China Law Translate – Interim Measures for the management of generative AI services

[37] Artificial Intelligence 2025 – China (Chambers Global Practice Guide)

[38] KEI – India’s granting of compulsory license on Nexavar (Natco vs. Bayer)

[39] COVID-19 TRIPS waiver debate and outcomes – MSF Access

[40] Fair Standards Alliance – FRAND-compliant patent pools

[41] Slashdot – RSS co-creator launches new protocol for AI data licensing (RSL)

[42] TechCrunch – RSS co-creator launches new protocol for AI data licensing

[43] Reddit and AI data licensing revenue (estimated $60 million from Google)

[44] RSL Standard – Official website on AI company compliance

[45] PYMNTS – EU Parliament report endorses compulsory copyright licensing for AI training

[46] ASCAP opposition to compulsory licensing – Referenced in PYMNTS article

[47] ACM Digital Library – Reclaiming the digital commons: A public data trust for training data

[48] arXiv – Reclaiming the digital commons: A public data trust for training data

[49] ALT-EDIC – European Language Data Space

[50] Complete Music Update – New AI transparency proposal in US Congress (TRAIN Act)

[51] Hogan Lovells – The EU AI Act: Prohibited practices and AI literacy requirements

[52] Senator Welch – TRAIN Act bill to protect musicians, artists, and creators

[53] Tech Policy Press – How the emerging market for AI training data is eroding Big Tech’s fair use defense

[54] EFF – Have you tried turning it off and on again: Rethinking tech regulation and creative labor

[55] EFF – Copyright and AI: The cases and the consequences

[56] The Authors Guild – Understanding the AI class action lawsuits

[57] Reed Smith – Bartz v. Anthropic: Pirated sources destroy fair use defense

[58] BitLaw – Thomson Reuters v. ROSS Intelligence (AI transformative use analysis)

DSGVO Cookie Consent mit Real Cookie Banner