Intellectual property has become AI’s constitutional question
Artificial intelligence is often discussed as a contest over computing power, energy consumption and market structure. Yet beneath those visible struggles lies a quieter and more durable issue: intellectual property. As generative systems move from laboratories into publishing, software, design, science and entertainment, they challenge the legal frameworks that determine who may copy, transform, commercialise and claim authorship over knowledge and culture.
The difficulty is structural. Modern AI systems are built by ingesting immense volumes of text, images, audio, video and code, much of it protected by copyright or adjacent rights. They then produce outputs that may echo styles, patterns or fragments from that material without reproducing it in any straightforward human sense. The result is an awkward fit with doctrines designed for identifiable authors, bounded works and legible acts of copying.
The argument over AI and IP is really an argument about where value is created in the production of knowledge.
That is why the disputes now unfolding in courts and policy forums matter beyond the technology sector. They will influence the economics of media, the viability of open scientific exchange, the balance between competition and concentration, and the terms on which creative labour is recognised in digital markets.
The training-data dispute sits at the centre
The most contentious issue is whether using protected works to train AI models constitutes infringement, fair dealing, fair use, text-and-data mining, or some new category requiring bespoke treatment. In the United States, several lawsuits brought by authors, artists and news organisations turn on whether model developers made unauthorised copies during training and whether resulting systems create derivative uses that compete with original markets. In Britain and the European Union, the debate has often centred on statutory exceptions for text and data mining, and whether rightsholders should be able to reserve their works from such uses.
Legally, training does involve copying at some stage, if only to ingest and process material. The harder question is whether that copying should be treated as an intermediate technical step akin to indexing and search, or as the extraction of expressive and economic value on a scale that demands permission and remuneration. Courts have not settled the matter. Policy-makers have not, either.
One reason the issue is so difficult is that AI training is neither purely archival nor purely substitutive. It does not usually aim to republish the source works themselves. But it may nevertheless exploit them to build systems that can serve adjacent or competing functions. For creators, that can look less like research and more like unlicensed industrial appropriation.
Copyright doctrine was not built for probabilistic systems
Copyright law traditionally asks recognisable questions: was a protected work copied, was a substantial part taken, is the new use transformative, and does it harm a legitimate market? Generative AI complicates each of these tests. Models do not store or retrieve source material in the manner of a conventional database, yet they are undeniably shaped by it. Outputs are often novel in composition, yet they may be statistically dependent on the expressive choices of millions of prior authors. Harm may not arise from one infringing copy, but from a system-wide erosion of licensing markets.
The argument over AI and IP is really an argument about where value is created in the production of knowledge.
The US Copyright Office has already drawn one consequential line. In guidance and review decisions, it has said that works containing AI-generated material are eligible for protection only to the extent of human authorship, not for machine-generated elements standing alone. That principle preserves copyright’s commitment to human creativity, but it does not solve the practical question of how much human selection, arrangement or editing is enough in a workflow where machine generation does much of the first draft.
The problem is not merely doctrinal. It is evidential. Claimants may struggle to prove that a specific output reproduces protected expression from a specific source. Developers, meanwhile, may find it difficult or commercially inconvenient to document exactly what data entered training sets, under what terms, and with what technical safeguards. Litigation is therefore likely to revolve as much around transparency and records as around abstract legal theory.
Authorship is becoming a spectrum rather than a binary
The familiar image of a lone author was always incomplete. Creative work has long involved editors, assistants, software tools, archives and collaborative processes. AI extends that trend and makes the line-drawing exercise unavoidable. If a user crafts a sophisticated prompt, curates outputs, iterates repeatedly and substantially edits the result, there is a plausible claim to authorship in the final work. If the user merely requests a finished image or article and accepts the first version, the claim is weaker.
That distinction matters because copyright does not protect effort in the abstract; it protects original expression attributable to a human author. The emerging legal consensus is therefore practical rather than philosophical: AI may assist creativity, but legal rights will attach principally where a person can show sufficient control over expressive choices in the resulting work.
In most legal systems, creativity is not being mechanised so much as disaggregated into stages, each with different claims to ownership.
This may produce a more layered market in rights. One party may own the copyright in underlying source material, another may assert contractual control over model outputs, and a third may hold rights in the human-authored arrangement or post-production. For publishers, studios and software firms, the challenge will be less about metaphysics than chain-of-title management.
Licensing is the most obvious solution, but not a simple one
In principle, broad licensing could reduce conflict. If model developers licensed books, articles, images, music and code for training, creators would be paid and legal uncertainty would diminish. In practice, however, licensing at internet scale is difficult. Many works have fragmented ownership, uncertain metadata or unclear territorial status. Collective licensing is feasible in some sectors, especially music and publishing, but much harder for the sprawling and globally distributed universe of online content.
There is also a strategic problem. Large incumbents may be able to absorb licensing costs and negotiate private deals, while smaller entrants and open research communities cannot. If the answer to AI and IP is simply more transaction-heavy licensing, the likely effect is to reinforce concentration. A legal regime intended to protect creators could end up fortifying the largest organisations at both ends of the chain.
That suggests a policy trade-off. Strong permission requirements may support rightsholders and encourage data provenance, but they can also raise barriers to entry and impede research. Broader exceptions may preserve experimentation and competition, but they risk hollowing out creative markets if compensation mechanisms are absent or ineffective.
Europe is testing transparency as a regulatory compromise
In most legal systems, creativity is not being mechanised so much as disaggregated into stages, each with different claims to ownership.
The European Union has approached the problem through a mix of copyright rules and platform regulation. Its Copyright Directive created text-and-data mining exceptions, including a broader exception subject to rightsholders’ ability to opt out in some circumstances. More recently, the AI Act introduced transparency obligations for certain general-purpose AI systems, including requirements related to publishing sufficiently detailed summaries of training data and complying with EU copyright law.
This does not fully answer whether training on protected material is lawful. But it recognises a basic institutional reality: legal rights are difficult to enforce if relevant information remains opaque. Transparency can therefore serve as a midpoint between unrestricted extraction and rigid ex ante licensing. It helps creators understand whether their works were likely used; it helps regulators assess compliance; and it may encourage the development of machine-readable reservation systems for rights management.
Still, transparency has limits. Summaries of training data may be too general to support meaningful enforcement. Detailed disclosure may expose commercially sensitive information or prove technically burdensome. And for globally deployed systems, divergence between jurisdictions could create a patchwork in which compliance standards vary widely across markets.
Patents pose a different, but related, challenge
Public debate focuses heavily on copyright, yet patents are also under pressure. AI is increasingly used in drug discovery, materials science, chip design and engineering optimisation. This raises at least two questions. First, can AI-assisted inventions satisfy existing standards of novelty, inventive step and sufficiency? Second, who counts as the inventor when a machine system contributes materially to the inventive process?
Courts in several jurisdictions have rejected attempts to list an AI system as an inventor, maintaining that inventorship remains a legal status reserved for natural persons. That position avoids immediate disruption, but it leaves unresolved how much AI contribution is compatible with naming a human inventor in good faith. If AI systems become integral to identifying patentable compounds, designs or methods, patent offices will face increasing pressure to clarify disclosure expectations and inventorship standards.
There is another concern. AI may lower the cost of generating large numbers of patent applications or prior-art searches, potentially worsening thickets in already crowded fields. Here, too, IP law faces a quantity problem: systems that amplify the production and processing of knowledge can overwhelm institutions designed for slower, more legible forms of innovation.
Trade marks and personality rights are becoming entangled with synthetic media
Generative systems do not only raise questions about ownership of works; they also unsettle the governance of identity and brand. Synthetic images, voices and videos can imitate the distinctive indicia of public figures, performers and commercial organisations. Some disputes will fall under trade-mark law where signs are used in commerce in a way that confuses consumers. Others will implicate passing off, unfair competition, publicity rights or data protection, depending on the jurisdiction.
The difficulty is that synthetic media can create harm without conventional counterfeiting. An AI-generated voice that evokes a performer may not use a registered trade mark, yet it may still appropriate valuable identity signals. A generated advertisement in a recognisable style may not reproduce a logo, yet it may trade on established reputation. Legal systems may therefore need to rely on a wider toolkit than classic IP alone.
This matters commercially because trust is itself an economic asset. As synthetic content becomes cheaper to produce, verification and provenance will become more valuable. That could shift attention from ex post litigation to ex ante authentication systems, contractual safeguards and disclosure standards.
Open knowledge and creative labour need not be enemies
The healthiest settlement will be one that preserves broad access to knowledge while making exploitation legible enough for creators to bargain.
One unhelpful tendency in this debate is to frame all restrictions as anti-innovation and all openness as anti-creator. In reality, knowledge economies depend on both access and reward. Scientific progress has long relied on shared corpora, citation, reuse and cumulative learning. Cultural production has likewise always involved influence, quotation and adaptation. The challenge is not to freeze these practices, but to preserve them while preventing extraction from becoming one-sided.
There are several policy paths. Governments could support extended collective licensing schemes in sectors where markets are mature enough to administer them. They could clarify research exceptions for non-commercial or public-interest uses while imposing stricter rules for commercial deployment at scale. They could require stronger provenance standards for training data and outputs. They could also invest in public datasets and public-interest AI infrastructure to reduce dependence on opaque private archives.
The healthiest settlement will be one that preserves broad access to knowledge while making exploitation legible enough for creators to bargain.
No single instrument will suit every field. Software, journalism, film, music, scientific literature and personal data each raise distinct issues. The task for law is not to discover one master rule, but to build interoperable regimes that reflect sectoral differences without creating unmanageable complexity.
The real contest is over bargaining power
Much of the rhetoric surrounding AI and IP invokes grand principles: freedom to learn, the rights of authors, the public domain, the future of innovation. These principles matter, but the operational question is bargaining power. Who has the information, leverage and legal resources to shape the terms on which data is used and value is shared? At present, those advantages are unevenly distributed.
Individual creators typically lack visibility into model training and the means to litigate complex claims. Smaller publishers and archives may have rights but not negotiating strength. Large technology groups, meanwhile, can combine technical opacity, contractual scale and cross-border deployment. Unless regulation improves transparency and collective bargaining options, the formal existence of rights may mean less than the practical ability to enforce them.
This is why the future of AI and IP will be determined not only in appellate judgments, but in standard-setting, licensing institutions, auditing practices and disclosure rules. Governance mechanisms that reduce information asymmetry may matter as much as any single legal doctrine.
What a durable settlement might look like
A workable long-term framework is beginning to come into view, even if no jurisdiction has fully articulated it. First, human authorship will remain the basis of copyright protection, with AI-treated works protected only where human creative contribution is demonstrable. Secondly, training on copyrighted works will probably not be treated as either wholly forbidden or wholly free. Instead, outcomes are likely to vary by context: more permissive for research and analysis, more constrained where uses are commercial, large-scale and market-substituting. Thirdly, transparency obligations will expand, because enforceable rights require traceability.
Fourthly, collective and standardised licensing mechanisms are likely to grow where transaction costs can be reduced. Fifthly, courts and patent offices will continue to insist on human inventorship while tightening expectations around disclosure of AI’s role in inventive processes. And finally, adjacent legal regimes, from competition law to consumer protection, will increasingly supplement IP where synthetic content affects market fairness and trust.
The broader lesson is that AI has not made intellectual property obsolete. It has revealed which assumptions in twentieth-century IP law were contingent on the pace and scale of human production. When systems can absorb millions of works, generate endless variants and circulate globally at negligible cost, the old categories do not disappear, but they must be reinterpreted. The societies that manage this well will be those that treat IP not as a brake on technical progress nor as a sacred absolute, but as a tool for allocating incentives, accountability and legitimacy in an economy where intelligence is increasingly collective, computational and contested.




