The long prehistory
Intellectual property law was not designed for generative AI. Copyright emerged to govern copying and distribution of expressive works; patent law to reward inventions; trade mark law to distinguish goods and services. For most of their modern history, these regimes assumed a relatively stable chain of human activity: a person made something, another person copied or commercialised it, and the law assigned rights and liabilities accordingly.
Machine learning disrupted that sequence. Modern AI systems are trained on vast corpora of text, images, audio and code, much of it protected by copyright or shielded by contract. They can also produce outputs that resemble, transform or compete with those inputs. As a result, the legal dispute is no longer limited to who copied a finished work. It increasingly centres on whether ingesting works for training is itself an actionable use, whether machine-generated material can be protected, and who bears responsibility when outputs infringe.
Those questions did not arrive all at once. They emerged through a series of decisions, consultations and lawsuits that gradually exposed the limits of twentieth-century IP doctrine. The timeline below traces the main inflection points.
2014 and the UK text-and-data-mining exception
One of the earliest significant policy markers came in Britain. In 2014, the UK introduced a copyright exception permitting copies for text and data analysis for non-commercial research. At the time, the reform was framed around computational research rather than generative AI. Yet it mattered because it acknowledged a distinction that has since become central: copying for machine analysis is not quite the same as copying for human consumption.
The UK exception was narrow. It applied only where a person had lawful access to the work, and only for non-commercial research. Even so, it established a precedent for treating data mining as a socially useful activity that copyright should not unnecessarily impede. Similar debates later intensified across Europe and beyond as AI developers argued that training requires broad access to information, while rightsholders countered that mass ingestion extracts value from creative labour without permission.
What looked like a modest research reform now appears as an early sign of a larger fault line. Once AI moved from the laboratory to the market, the non-commercial limitation became the very point of contention.
2016 to 2019 and Europe redraws the map
The European Union’s Digital Single Market Directive, adopted in 2019 after years of debate, brought text and data mining to the centre of copyright policy. It created two distinct exceptions: one mandatory exception for research organisations and cultural heritage institutions, and another broader exception for text and data mining unless rightsholders explicitly reserve their rights.
This was an attempt at balance. Europe recognised that machine analysis of large datasets could support innovation and scientific discovery. But it also sought to preserve bargaining power for publishers, image libraries and other rightsholders by allowing an opt-out for commercial uses. That architecture now sits at the heart of contemporary disputes over AI training in Europe.
AI has pushed copyright law away from simple acts of duplication and towards a more difficult question: when does analysis become appropriation?
The directive did not settle the issue. It effectively turned control over machine-readable reservation of rights into a practical and technical problem. If rightsholders can opt out, how must they signal that choice? If developers scrape the open web at scale, how should those signals be respected? These are legal questions, but they are also infrastructural ones. They concern metadata, platform design and record-keeping as much as doctrine.
AI has pushed copyright law away from simple acts of duplication and towards a more difficult question: when does analysis become appropriation?
Human authorship becomes the first hard boundary
While policymakers were grappling with training, courts and copyright offices were revisiting a more basic question: can a work generated by a machine attract copyright protection at all? In the United States, the Copyright Office repeatedly affirmed that copyright protects “the fruits of intellectual labour” founded in human creativity, not works produced absent human authorship. That position drew on older case law but gained fresh salience as generative systems improved.
A notable turning point came in litigation over an image created by an AI system and submitted for registration without a human author. In 2023, a federal court in Washington upheld the Copyright Office’s refusal to register the work, endorsing the human authorship requirement. The ruling did not deny that AI can be used in creative processes. Rather, it marked a threshold: copyright may attach where a human exercises sufficient creative control, but not where a machine is presented as the author of record.
This distinction is likely to endure, though it remains fact-sensitive. Editing prompts, curating outputs and making expressive selections may support protection for some aspects of a work. Merely initiating an automated process may not. In practical terms, the law is edging towards a layered view of authorship rather than a binary one.
Patents and the limits of machine inventorship
Patent law confronted a parallel controversy. A series of applications around the world sought to list an AI system as the inventor. Courts and patent offices in the United Kingdom, the United States and elsewhere rejected those applications, generally on the ground that existing statutes require an inventor to be a natural person.
The UK Supreme Court’s 2023 decision in Thaler v Comptroller-General of Patents, Designs and Trade Marks made the point plainly. Under the Patents Act 1977, an inventor must be a person. Similar conclusions had already emerged in the United States and before the European Patent Office. The immediate result was formal rather than philosophical: current law does not recognise machine inventors.
But the broader implications are significant. Patent systems are meant to reward and disclose inventive activity. If AI tools increasingly contribute to problem-solving, then the legal system will face a practical question: how much human contribution is enough for inventorship? So far, the answer has been to preserve the human-centred model. Yet the pressure is not going away, especially in fields where algorithmic design or molecular discovery becomes deeply entwined with research workflows.
2022 and the first wave of generative AI copyright litigation
The public release of powerful image and text generators transformed an academic and policy debate into a commercial and cultural confrontation. By 2022 and early 2023, artists, photo agencies, authors and software developers were bringing lawsuits in the United States that alleged copyright infringement, trade mark violations and unfair competition linked to training datasets and generated outputs.
These cases vary in strength and theory, but together they crystallise the core legal issues. Is the copying involved in training protected by fair use? Can plaintiffs show substantial similarity between outputs and specific works in a training set? Does the use of watermarked or licensed content alter the analysis? And where the output mimics a creator’s style, is style itself protectable under copyright, or only particular expressions?
Much of this litigation remains unresolved. That uncertainty matters. In technology markets, legal ambiguity often functions as a temporary operating environment. Firms proceed, investors price risk imperfectly, and rightsholders test which claims can survive early motions. The first wave of cases is less important for any single verdict than for the evidentiary and doctrinal framework it is beginning to establish.
Getty Images, artists and the economics of training data
The fight over AI and IP is not only about ownership of outputs. It is about bargaining power over the inputs that make machine intelligence possible.
Among the most closely watched disputes have been claims brought by image rightsholders and groups of artists. These lawsuits have sharpened attention on a basic economic question: if training depends on large quantities of copyrighted material, should there be licensing markets for that use?
Rightsholders argue that AI developers have benefited from a vast unremunerated transfer of value. Developers typically respond that training is transformative, that models do not store or republish works in any straightforward sense, and that broad licensing mandates could entrench incumbents while choking experimentation. Both arguments have force. Copyright has long tolerated some intermediate copying in the service of new technologies. But scale matters, and AI training is unprecedented in scale.
The fight over AI and IP is not only about ownership of outputs. It is about bargaining power over the inputs that make machine intelligence possible.
The likely medium-term outcome is not a universal legal principle so much as a patchwork. Some sectors may move towards licensing and collective deals; some uses may be upheld as lawful without permission; some markets may rely more heavily on opt-outs, technical controls or synthetic data. IP law may set outer boundaries, but commercial ordering will do much of the detailed work.
Authors sue and the text corpus comes under scrutiny
Text-based AI generated a different but related challenge. For authors and publishers, the concern is not merely that books and articles have been ingested for training. It is also that systems can produce summaries, imitations or substitutes that compete with the originals. Lawsuits brought by writers have therefore become a test of whether literary and journalistic works receive meaningful protection when used as raw material for language models.
These cases are likely to turn on classic copyright concepts applied to unfamiliar facts. Courts may ask whether training copies are transient or enduring, whether the use is transformative, whether outputs reproduce protected expression, and what effect the technology has on actual or potential markets. They may also confront evidentiary asymmetries. Developers know more than claimants about what datasets were used, what filtering occurred, and how outputs are generated.
That asymmetry is one reason transparency has become an IP issue. Disclosure requirements, dataset documentation and record-keeping could influence not only regulatory oversight but also the practical enforceability of private rights. A right that cannot be audited is often a right that cannot be meaningfully exercised.
2023 and official guidance on AI-generated works
Regulators and administrative bodies have tried to reduce uncertainty even where legislatures have not acted. In 2023, the US Copyright Office issued guidance clarifying that works containing AI-generated material may be registered only to the extent of the human-authored elements. Applicants must disclose the inclusion of AI-generated content and identify the human contribution.
This guidance is modest but consequential. It rejects both extremes: neither a blanket denial of protection for all AI-assisted works, nor automatic protection for outputs generated by software. Instead it adopts a granular approach focused on human creative control. That approach is administratively cumbersome, but it reflects a deeper reality. Creative production is becoming hybrid. Law is therefore being asked to separate human expression from machine execution within a single artefact.
The same logic is appearing elsewhere, including in policy discussions about moral rights, performer rights and database protection. In each area, the central task is not to decide whether AI is present. It is to determine which legally relevant acts remain attributable to people.
2024 and the EU AI Act brings transparency into focus
The most important policy shift is towards traceability: who used what, under which legal basis, and with what opportunity for challenge.
The European Union’s AI Act is not an IP statute, but it has important IP consequences. The final text includes transparency obligations for certain general-purpose AI models, including a requirement to make publicly available a sufficiently detailed summary about the content used for training. It also requires policies to comply with EU copyright law.
This is a notable development because it links AI governance to IP compliance without attempting to rewrite copyright itself. In effect, the EU is using horizontal AI regulation to improve visibility into training practices. That may help rightsholders assess whether their works were used and whether opt-outs were respected under the DSM framework.
Whether the mechanism works will depend on implementation. A summary that is too vague will offer little practical value; one that is too specific may raise confidentiality and security objections. Even so, the direction of travel is clear. The governance of AI is moving beyond abstract principles towards obligations that can alter evidentiary balance and transaction costs in IP disputes.
The most important policy shift is towards traceability: who used what, under which legal basis, and with what opportunity for challenge.
What the courts are really deciding now
It is tempting to see the current confrontation as a referendum on whether AI and copyright can coexist. That is too crude. Courts are deciding a narrower but still consequential set of questions. How should fair use or fair dealing apply to training? What counts as sufficient human authorship? Can existing remedies handle probabilistic systems whose outputs are not deterministic copies? And how should liability be allocated across model developers, deployers and end users?
These are not merely technical matters. They will shape industrial structure. A permissive interpretation of training could favour scale, because large models benefit most from broad access to data. A restrictive interpretation could strengthen licensing markets and perhaps advantage firms with proprietary archives or the means to strike deals. Intermediate outcomes may encourage new forms of rights management, provenance tracking and collective bargaining.
In that sense, AI & intellectual property is becoming a contest over market design. The law will not decide every commercial arrangement, but it will determine who comes to the table with leverage.
The next phase
The next chapter is unlikely to be defined by a single landmark ruling. More probably it will be shaped by convergence across several fronts: case law on training and outputs, administrative practice on registration, contractual standards for data access, and technical systems for attribution and opt-out signalling. Different jurisdictions will move at different speeds, and some divergence is inevitable.
Still, a few broad conclusions already stand out. First, human authorship remains the anchor of copyright and inventorship, even as AI tools become more capable. Secondly, the lawful use of training data has become the decisive frontier of IP conflict. Thirdly, transparency is emerging as the bridge between abstract rights and practical enforcement.
That makes this moment less revolutionary than it first appears. Intellectual property law is not being replaced. It is being redistributed across the AI pipeline: from creation and publication to ingestion, training, generation and traceability. The institutions of the old system still matter. But they are being asked to govern a world in which copying is not always consumption, creation is often collaborative with machines, and value is captured long before any final work reaches the public.
For judges, legislators and creators alike, the challenge is no longer simply to defend rights against infringement. It is to define what lawful participation in the knowledge economy looks like when machines learn from culture at industrial scale.




