Hub
Opinion & Commentary
Copyright’s Hard Lesson for Generative AI
AI & Intellectual PropertyOpinion & Commentary

Copyright’s Hard Lesson for Generative AI

The law was built for copies; machine learning is forcing it to confront inference, scale and industrial opacity.

Society OS Research21 July 202612 min read

Key Insight: The future of AI and intellectual property will depend less on choosing between innovation and creators than on building workable rules for transparency, licensing and accountability at scale.

Copyright has moved from the margins to the centre

For years, intellectual property was treated as a technical domain: important to publishers, studios, collecting societies and specialist lawyers, but rarely decisive in wider economic strategy. Generative AI has changed that. Systems trained on vast corpora of text, images, audio and code have made copyright law newly political because they convert a long-settled legal architecture into a live question about who captures value in the knowledge economy.

The argument is often framed too crudely. One camp portrays copyright as an obstacle to progress, a relic of analogue distribution ill-suited to computation. The other casts AI developers as industrial-scale free riders, extracting value from human creativity without consent or compensation. Both positions contain some truth, and both are incomplete. Copyright has always been an imperfect tool for governing cultural production. Yet the arrival of generative AI has exposed something more fundamental: modern machine learning depends on acts of ingestion, transformation and output that do not fit neatly into legal categories designed around reproduction and distribution.

Generative AI has not merely created a copyright dispute; it has revealed a governance gap between legal concepts built for copies and technologies built on patterns.

This matters beyond the courtroom. If the rules remain ambiguous, markets for training data will stay underdeveloped, creators will face prolonged uncertainty, and the largest firms will enjoy the greatest freedom to operate simply because they can afford litigation. That is not a recipe for either innovation or fairness. It is a recipe for concentration.

The law was written for copies, not models

Copyright systems in Britain, America and Europe were designed to regulate identifiable acts: copying a book, performing a song, distributing a film, adapting a work. Those acts may now be embedded within machine learning pipelines that are technically diffuse and commercially opaque. Training a model typically involves scraping or otherwise acquiring vast quantities of material, processing it into machine-readable form, and adjusting statistical weights so the system can generate outputs in response to prompts. The legal question is deceptively simple: are those steps acts that copyright should control?

Different jurisdictions answer in different ways. In the United Kingdom, text and data mining exceptions are comparatively narrow, largely tied to non-commercial research, as set out by the Intellectual Property Office. In the European Union, the Directive on Copyright in the Digital Single Market created text and data mining exceptions under Articles 3 and 4, while preserving an opt-out mechanism for rights holders in some circumstances. In the United States, much of the immediate debate turns on fair use, a flexible doctrine that has historically accommodated certain forms of technological copying, including search indexing and digitisation for discovery.

But none of these frameworks was drafted with large-scale generative models in mind. The distinction between learning from a work and reproducing it begins to blur when model outputs can resemble training material, when memorisation is possible, and when the system itself becomes a commercial substitute for parts of the markets that sustained original production.

Fair use cannot carry the whole burden

American debates have focused heavily on whether AI training is “transformative” in the sense developed by US case law. There are precedents suggesting courts sometimes tolerate copying when the purpose differs from that of the original work and the social benefit is substantial. Search engines and certain digitisation projects benefited from that logic. It is tempting, then, to treat model training as simply another informational use.

That analogy is incomplete. Search helps users find existing works. Generative AI can produce outputs that compete with them. A model trained on illustration, journalism, music or software code may not merely index a market; it may enter it. That does not automatically make training unlawful. It does, however, weaken simplistic claims that all ingestion is functionally equivalent to reading.

Generative AI has not merely created a copyright dispute; it has revealed a governance gap between legal concepts built for copies and technologies built on patterns.

Moreover, fair use is a litigation doctrine, not a regulatory settlement. It is applied ex post, case by case, after expensive disputes. That may be tolerable for a library project or a search index. It is far less satisfactory as the default governance model for an industrial ecosystem involving millions of rights holders and global supply chains of data. A society that relies on court battles alone is effectively deciding that legal ambiguity is an acceptable subsidy for those with the deepest pockets.

Europe’s opt-out model is more principled than practical

The European Union deserves credit for recognising text and data mining as a distinct policy question rather than forcing every dispute through inherited doctrines. The DSM Directive’s framework tries to balance innovation and rights by allowing mining in some contexts while letting rights holders reserve their works in others. In theory, this is a pragmatic middle path.

In practice, opt-out systems are only as good as the technical and commercial infrastructure around them. Rights reservations must be machine-readable, discoverable and respected throughout highly fragmented data acquisition chains. Smaller creators and publishers may struggle to exercise their rights effectively. Meanwhile, model developers can face uncertainty over what was reserved, when, and under what standards. Without robust record-keeping and auditability, an opt-out can become a legal fiction: elegant in principle, weak in operation.

The deeper problem is asymmetry. Large platforms are better equipped to parse rights metadata, negotiate licences and defend their processes. Independent creators may technically possess control while lacking bargaining power. Regulation that assumes equal capability among market participants often ends by rewarding scale.

When rights management depends on perfect metadata and pristine audit trails, the law risks granting formal control to creators while delivering practical advantage to intermediaries.

Transparency is the missing layer

The most obvious flaw in current debates is that too much remains hidden. Rights holders often do not know whether their work was used, in what quantities, under which legal theory, or with what safeguards against memorisation and regurgitation. Developers, for their part, argue that disclosing datasets in detail may reveal trade secrets, expose personal data issues, or prove technically difficult because datasets are assembled from multiple sources over time.

Yet without transparency, there can be no mature market in licensing and no credible enforcement of exceptions. This is why recent policy moves matter. The European Union’s AI Act includes transparency obligations for providers of general-purpose AI models, and the final text links those obligations to copyright compliance, including making publicly available a sufficiently detailed summary of training content. One can debate whether the requirements are adequate; one cannot seriously argue that opacity is a sustainable equilibrium.

Transparency need not mean publishing every file. It can mean standardised provenance records, independent auditing, secure registries and disclosure categories that make rights legible without compromising security or privacy. Financial markets learned long ago that disclosure is not the enemy of enterprise. It is a precondition for trust. AI will need the same lesson.

Licensing is not a panacea, but it is necessary

Some technologists treat licensing as economically impossible: too many works, too many owners, too much friction. That objection is overstated. Collective licensing systems already exist in music, publishing and broadcasting because markets routinely develop intermediaries when rights are complex. The harder question is not whether licensing is possible, but where it is most useful and what exactly is being licensed.

When rights management depends on perfect metadata and pristine audit trails, the law risks granting formal control to creators while delivering practical advantage to intermediaries.

There are at least three distinct markets emerging. The first is licensing for training data, where publishers, archives, image libraries and other rights holders can authorise use of corpora under agreed terms. The second is licensing for retrieval and access, particularly where AI systems deliver or summarise protected material for end users. The third is licensing tied to output risk, such as permissions for style-sensitive or high-fidelity uses in commercial production.

None of these markets will cover every use case. Some exceptions should remain. Research, preservation, security testing and public-interest analysis all warrant room to operate. But the notion that the only alternatives are unrestricted scraping or universal prior permission is a false binary. Well-designed licensing regimes can reduce uncertainty, create revenue streams and encourage cleaner data practices.

They can also improve competition. If lawful access to high-quality data becomes more standardised, smaller firms and research institutions gain a clearer path to entry than one defined by ambiguous scraping and retrospective lawsuits.

Creators’ concerns are not nostalgic protectionism

It is fashionable in some circles to dismiss authors, artists and performers as incumbents resisting technological change. That is lazy analysis. Creative workers are not objecting to competition in the abstract; they are objecting to an economic arrangement in which their work may be appropriated as an input to systems that dilute demand for future work while offering little visibility, compensation or recourse.

These concerns are not confined to famous names. Mid-list authors, freelance illustrators, voice actors and independent journalists often occupy already fragile labour markets. Generative AI increases productivity for some users, but it also shifts bargaining power towards those who control computation, distribution and data aggregation. Copyright cannot solve every labour-market problem, but it remains one of the few institutional mechanisms that ties payment to use in cultural industries.

The policy challenge, then, is not to fossilise old business models. It is to ensure that productivity gains do not come from simply externalising costs on those least able to absorb them. Societies that fail to protect the economic base of cultural production may discover, belatedly, that abundant synthetic content and sustainable human creativity are not the same thing.

The essential question is not whether machines can remix culture, but whether the market can still reward the people whose work gives the machine something worth learning from.

Output harms are different from training harms

Too much commentary collapses all IP questions into a single issue. In fact, training and output require different tools. Training disputes concern ingestion, copying and market substitution at the input stage. Output disputes concern resemblance, attribution, consumer confusion and unfair competition at the point of use.

That distinction matters because a model could be trained lawfully yet still generate infringing outputs in particular cases. Equally, a contested training process does not mean every output is unlawful. Policymakers should resist one-size-fits-all remedies. Better output controls, provenance signals and red-teaming for memorisation can mitigate some risks without resolving the licensing debate. Conversely, a training licence does not excuse the release of tools that systematically imitate living artists in commercially damaging ways or reproduce protected material too closely.

Trade mark law, passing off, performers’ rights, database rights and contract law also enter the picture. The future legal framework will be plural, not singular. Copyright remains central, but it is not sovereign over every harm associated with generative systems.

The essential question is not whether machines can remix culture, but whether the market can still reward the people whose work gives the machine something worth learning from.

The public domain must not become collateral damage

In the rush to protect creators, there is a risk of over-correction. The public domain is one of modern society’s greatest cultural assets. So are open licences, open scientific corpora and other shared knowledge resources. Rules that make data use so burdensome that only wealthy actors can comply would paradoxically entrench the very concentration critics fear.

That is why the best settlement is neither maximalist control nor laissez-faire appropriation. It is a tiered system. Public-domain materials should remain freely usable. Clearly licensed open materials should remain available under their terms. Non-commercial research should benefit from robust exceptions, subject to privacy and security safeguards. Commercial use of protected works at scale should require either a clear exception with enforceable conditions or a licence. The crucial point is clarity.

Clarity has economic value. It lowers transaction costs, encourages investment and reduces the advantage of legal brinkmanship. Ambiguity, by contrast, favours those able to move first, settle later and treat compliance as a strategic variable.

What a workable settlement would look like

A sensible regime would start with four principles. First, disclosure: developers of general-purpose models should provide meaningful information about training sources, rights-reservation compliance and safeguards against memorisation. Secondly, licensability: rights holders should have practical ways to authorise large-scale use, individually or collectively, without prohibitive transaction costs. Thirdly, enforceability: regulators and courts need auditable records, not merely assurances. Fourthly, proportionality: obligations should reflect the scale and commercial significance of the model, with carve-outs for genuine research and public-interest uses.

That settlement would not please purists. Some rights holders would want broader control; some developers would want broader freedom. But policy is not a seminar in absolutes. It is a mechanism for reducing social conflict while preserving productive activity. The objective should be to make lawful conduct easier than unlawful conduct, and transparent conduct cheaper than opaque conduct.

There is also a geopolitical dimension. Jurisdictions that create credible, navigable rules for data use and licensing will be better placed to attract investment in both creative industries and AI research. The contest is not simply over who builds the largest models. It is over who builds the most governable ecosystem around them.

The argument is really about economic constitutionalism

Beneath the technical disputes sits a larger question: who sets the terms on which knowledge is transformed into economic power? Copyright law, for all its imperfections, is part of society’s answer to that question. It allocates bargaining power, shapes incentives and signals what kinds of extraction are acceptable.

Generative AI has exposed a structural imbalance. Those who create expressive works are numerous and fragmented. Those who can ingest them at planetary scale are comparatively few and well capitalised. Left alone, markets of that kind tend towards take-it-or-leave-it outcomes. Law cannot erase the asymmetry, but it can civilise it.

The wisest response is neither panic nor permissiveness. It is institutional design. The choice is not between protecting art and advancing science. It is between a future in which AI develops through negotiated legitimacy and one in which it grows by exploiting ambiguity until courts, regulators and creators force a reckoning. The former is slower, less theatrical and far more likely to endure.

Copyright’s hard lesson for generative AI is straightforward. Innovation is most durable when it pays attention to provenance, permission and power. The systems reshaping culture will not earn public legitimacy merely because they are impressive. They will earn it when the rules around them are intelligible, enforceable and fair.

Sources & Further Reading

  1. 1.
  2. 2.
  3. 3.
  4. 4.
  5. 5.
  6. 6.
  7. 7.
  8. 8.
  9. 9.
  10. 10.
AIIntellectual PropertyCopyrightGenerative AILicensingRegulationCreative Economy
The engine behind the Signal

Where this connects to Society OS

The Sovereign Intelligence Hub is the free, open front door of Society OS — the sovereign operating system that turns the ideas you just read into working governance. Where this piece names a problem, Society OS is building the machinery to solve it: AI agents that act with your authority, trust you can verify, and compliance that runs as code.

The 42-Protocol Stack

The governance engine beneath every article — led by the Sovereign Trinity: Human-Twin-Agent identity, HEARTrank trust, and WISE Contracts that execute law, not just code.

F-ACT — the open agent standard

The vendor-neutral framework for governing AI agents before they act: Authority, Scope, Data, Audit, Revocation — free to read, cite and implement.

The Sovereign Platform

Put it to work: govern a fleet of AI agents with verifiable authority, tamper-evident evidence, and compliance-as-code across your whole operation.

Explore membershipRead the F-ACT standard

Continue Reading

More from the Sovereign Intelligence Hub

Who Owns Intelligence When Machines Learn From Everyone
AI & Intellectual Property

Who Owns Intelligence When Machines Learn From Everyone

14 min

How AI Forced Intellectual Property Into a New Era
AI & Intellectual Property

How AI Forced Intellectual Property Into a New Era

14 min

Who Owns Intelligence When Machines Learn From Everything
AI & Intellectual Property

Who Owns Intelligence When Machines Learn From Everything

14 min

Intellectual property after the model frontier
AI & Intellectual Property

Intellectual property after the model frontier

18 min read

Standards Patents Became the New Border Checkpoints
AI & Intellectual Property

Standards Patents Became the New Border Checkpoints

11 min read

Who Owns Intelligence When Machines Learn From Everything
AI & Intellectual Property

Who Owns Intelligence When Machines Learn From Everything

13 min

Never miss a signal

Weekly intelligence, no noise

The Sovereign Intelligence Hub — Society OS

© 1989–2026 Society OS Pty Ltd. All rights reserved.