Hub
Who Owns Intelligence When Machines Learn From Everything
AI & Intellectual Property

Who Owns Intelligence When Machines Learn From Everything

Generative AI is forcing copyright, patent and trade-mark law to answer questions it was never designed to ask.

Society OS Research24 July 202613 min read

Key Insight: The central intellectual-property challenge in AI is not whether old rules disappear, but which existing rights can still be enforced when creation, copying and attribution become probabilistic and massively distributed.

Artificial intelligence has turned intellectual property into infrastructure

For years, intellectual property was often treated as a specialist concern: material to publishers, pharmaceutical companies, film studios and patent litigators, but peripheral to wider debates about economic governance. Artificial intelligence has changed that. Systems trained on enormous collections of text, images, audio, video and software code depend on access to expressive works at a scale that copyright law never anticipated. At the same time, those systems produce outputs that resemble writing, design, music, source code and scientific discovery, bringing them into direct contact with nearly every branch of intellectual-property law.

The result is not a single dispute over ownership. It is a cluster of interlocking questions. Can copyrighted works be used to train models without permission? When model outputs resemble protected material, what counts as infringement? Can an AI-generated work receive copyright protection at all? If a machine contributes to an invention, who is the inventor? When synthetic content mimics a person’s voice or style, are copyright rules enough, or are other rights needed? These are no longer abstract puzzles. They shape the economics of creative industries, the development of research tools and the balance between private control and public knowledge.

AI has not abolished intellectual property. It has exposed how much of it depends on assumptions about human authorship, identifiable copying and traceable provenance.

The most important point is that AI does not simply challenge legal categories from the outside. It reveals their hidden architecture. Copyright is built around human authorship and expressive originality. Patent law is organised around inventorship, novelty and disclosure. Trade-mark law is concerned with source identification and consumer confusion. Database rights, trade secrets and performers’ rights each protect narrower forms of value. Generative AI interacts with all of them at once, often in ways that blur their boundaries.

Copyright sits at the centre of the dispute

The first and most contested issue is training. Modern foundation models are built by analysing vast amounts of existing material. In practical terms, that process usually involves copying data into storage, transforming it into machine-readable formats, extracting statistical relationships and then retaining some version of that information in model weights or associated systems. Rights-holders argue that such acts implicate copyright because protected works are reproduced and exploited without authorisation. Developers and many researchers counter that training is more akin to analysis than substitution: a model does not preserve works in ordinary human-readable form, but learns patterns from them.

Whether that argument succeeds depends heavily on jurisdiction. In the United States, much attention has focused on fair use, a flexible doctrine that weighs the purpose of use, the nature of the work, the amount used and the effect on the market. In the European Union and the United Kingdom, the relevant questions often involve text-and-data-mining exceptions, lawful access and whether rightsholders can reserve their rights. The legal terrain therefore differs sharply across advanced economies, even before courts rule on the specific facts of AI training.

The policy tension is obvious. If every use of copyrighted material for training required negotiated licences, only the largest actors might be able to assemble datasets at scale. Yet if training is treated as broadly exempt, creative workers and publishers may see their work absorbed into systems that compete with them without payment or attribution. Copyright law is thus being asked to decide not merely what is lawful, but what sort of market structure should govern knowledge-intensive industries.

Fair use and text-and-data mining are doing unusually heavy work

The burden placed on legal exceptions is striking. In the United States, fair use has historically accommodated activities such as search indexing, intermediate copying in software and certain forms of transformation. Courts have often been receptive where copying enables a new informational function rather than serving as a market substitute. That history matters because model training can be presented as a computational analysis of works rather than consumption of them as works. But the analogy is imperfect. Search engines point users back to sources; generative systems can produce substitute outputs in the same expressive market.

AI has not abolished intellectual property. It has exposed how much of it depends on assumptions about human authorship, identifiable copying and traceable provenance.

In the European Union, the Copyright in the Digital Single Market Directive created text-and-data-mining exceptions, including one for research organisations and cultural-heritage institutions and another, broader provision that permits mining unless rights are expressly reserved. That framework is more explicit than American fair use but also more administratively demanding. It raises operational questions about machine-readable reservations of rights, the technical means of respecting them and the consequences for models trained on mixed-status corpora. The United Kingdom introduced a text-and-data-mining exception for non-commercial research, but proposals to broaden it met strong opposition from creative sectors and have not been adopted.

What appears, at first, to be a narrow matter of legal interpretation is really a problem of institutional design. Exceptions built for search, research and digital archiving are now being stretched to cover the industrial production of synthetic content. Courts and legislators may prove reluctant either to outlaw training outright or to grant it near-total immunity. The likely outcome is a more fragmented settlement: some uses tolerated, some licensed, some restricted and many still uncertain.

Output liability is harder than input liability

Training disputes are only half the story. Even if the ingestion of works is lawful, outputs may still infringe. An image generator may produce a composition closely resembling a protected illustration. A language model may emit passages that reproduce parts of books, articles or source code. A music system may generate recordings that sound strikingly similar to copyrighted tracks. The legal question then turns from training to reproduction, substantial similarity and, in some jurisdictions, derivative works.

This is harder than it sounds because generative systems do not behave predictably. Most outputs are not close copies of any single work; they are synthetic recombinations shaped by prompts, system constraints and probabilistic sampling. Yet some models can memorise and regurgitate training data, especially where examples are repeated or highly distinctive. That makes infringement assessment a matter of both doctrine and engineering. How often do models reproduce protected expression? Under what conditions? Can safeguards materially reduce the risk? What records exist to trace the provenance of an output?

The law is comfortable with copying it can point to. AI often produces resemblance without an obvious chain of custody, which makes enforcement as much a forensic exercise as a legal one.

For claimants, the practical challenge is evidence. It is not enough to suspect that a model was trained on a particular work or that an output feels derivative. Courts typically require proof of access and actionable similarity, though standards vary. That shifts attention to audits, benchmark testing and discovery procedures. It also encourages technical measures such as filtering, deduplication and red-team testing, not because they solve the legal problem entirely, but because they shape how responsibility will be judged.

Authorship remains stubbornly human

Another legal front concerns whether AI-generated outputs can attract copyright protection at all. Here, many authorities have taken a relatively clear line: copyright protects works of human authorship. In the United States, the Copyright Office has repeatedly stated that purely machine-generated material is not registrable, though human selection, coordination and arrangement may be. Courts have echoed the principle that copyright requires human creative input. Similar instincts can be found elsewhere, even if statutory wording differs.

This does not mean AI-assisted works fall into a legal void. In practice, many outputs are the product of collaboration between human users and software systems. A designer may iteratively refine prompts, select among outputs, edit compositions and integrate them into a larger project. A writer may use a model to generate alternative phrasing before heavily revising the text. The legal question becomes how much human control and originality is enough. That is a familiar copyright problem in a new technical setting.

There is also a policy reason for caution. If law granted strong copyright to fully machine-generated material, it could create an immense volume of low-cost exclusive rights, potentially clogging cultural markets with little connection to human creativity. By insisting on human authorship, legal systems preserve a threshold that is both normative and practical. They recognise that copyright is not simply a reward for output, but an institution linked to personhood, creative labour and public accountability.

Patent law is confronting invention without an inventor

The law is comfortable with copying it can point to. AI often produces resemblance without an obvious chain of custody, which makes enforcement as much a forensic exercise as a legal one.

If copyright asks who authored an expressive work, patent law asks who invented a technical solution. AI complicates this too. Machine-learning systems are increasingly used in drug discovery, materials science, engineering optimisation and the generation of candidate inventions. In some cases, AI operates as a sophisticated tool under close human direction. In others, it appears to identify combinations or pathways that no individual researcher explicitly foresaw.

Courts and patent offices have so far resisted recognising AI systems as inventors. High-profile litigation involving patent applications that named an AI system as the inventor ended with courts in several jurisdictions concluding that existing law requires a natural person. The reasoning is partly textual and partly conceptual: inventorship carries duties of disclosure and attribution that legal systems assume can be borne by humans.

Yet the practical issue is not going away. If a research team uses AI extensively to generate and screen possibilities, who among the humans qualifies as inventor? The person who framed the problem, the person who selected the data, the person who evaluated the result or the person who recognised that a machine-generated output was patentable? Patent doctrine already struggles with collaborative science; AI intensifies the ambiguity. It may also increase pressure on disclosure rules, as patent examiners and courts seek greater transparency about how claimed inventions were produced.

Trade marks, passing off and the economics of synthetic identity

Not every AI-related intellectual-property dispute is about copyright or patents. Trade-mark law and related doctrines such as passing off are becoming more relevant as models generate brand-like signs, imitate packaging styles or produce content that appears to originate from a trusted source. The issue is not merely unauthorised copying of logos. It is consumer confusion in an environment where synthetic media can cheaply replicate the signals on which reputation depends.

This matters especially in advertising, search interfaces and automated agents. If AI systems recommend products, compose commercial messages or generate visual assets on demand, they can inadvertently reproduce protected marks or create misleading associations. Existing trade-mark law is well equipped to address confusion about commercial origin, but it is less suited to the diffuse and automated ways such confusion may now arise. Liability may sit with the user, the deployer or some intermediary, depending on the facts.

There is a broader economic point here. Intellectual property has always been partly about reducing information costs. Trade marks help consumers identify source; copyright structures markets for expression; patents disclose and reward invention. AI lowers the cost of imitation and increases the volume of plausible-but-unverified content. That makes the verification function of IP more valuable, even as the underlying rights become harder to police.

Transparency is becoming the hinge of enforcement

A recurring theme across these debates is opacity. Rightsholders often do not know whether their works were included in training data. Users may not know why a model produced a particular output or whether it carries legal risk. Regulators struggle to assess compliance without access to technical documentation, dataset records and internal testing. In response, transparency has become a central policy demand.

That demand appears in several forms: disclosure of training-data sources, mechanisms for rights reservation, watermarking or provenance metadata for outputs, record-keeping about model development and clearer notices about the role of AI in creative and inventive processes. None is a silver bullet. Full dataset disclosure may expose trade secrets or prove impractical where data sources are numerous and dynamic. Watermarking can be evaded or generate false confidence. Provenance standards depend on broad adoption. But without some increase in visibility, legal rights become theoretical.

In the AI era, the enforceability of intellectual property depends increasingly on visibility: what was used, how it was transformed and where synthetic outputs came from.

Transparency is therefore not just a compliance burden. It is a precondition for functioning markets. Licensing cannot scale if parties cannot identify relevant uses. Exceptions cannot be calibrated if policymakers lack evidence about harms and benefits. Litigation cannot reliably sort lawful from unlawful conduct if the underlying systems remain black boxes.

In the AI era, the enforceability of intellectual property depends increasingly on visibility: what was used, how it was transformed and where synthetic outputs came from.

Licensing will expand, but not as a universal solution

Many industry participants hope licensing will provide an orderly settlement. In some sectors it will. Publishers, image libraries, music-rights organisations and specialist data providers all have incentives to create markets for AI training and output use. Licensing can reduce uncertainty, compensate rightsholders and differentiate between premium, verified and unrestricted datasets. It may become especially important in enterprise settings where legal risk is priced carefully.

But licensing also has limits. Copyright ownership is often fragmented. Many works have multiple rightsholders; others have unclear provenance. Transaction costs can be prohibitive, especially for historical archives, user-generated content and web-scale corpora. A licensing-first model may also reinforce concentration by favouring actors with deep pockets and established legal infrastructure. From a competition perspective, that could be as significant as the copyright outcome itself.

For that reason, the future is unlikely to be a simple choice between unrestricted training and comprehensive licensing. More probable is a mixed ecology: compulsory or collective arrangements in some domains, bilateral licensing in others, public-interest exceptions for research, tougher rules for high-risk or highly substitutive uses and persistent litigation at the boundaries. Intellectual-property law will not converge neatly. It will stratify.

Courts can clarify doctrine, but legislatures will shape the market

Judges will decide important cases about infringement, fair use and registrability. Yet the larger settlement is legislative and regulatory. Lawmakers must decide how to balance innovation, competition and creative labour; whether to require opt-outs or opt-ins; how to handle public-sector and scientific uses; and what transparency obligations to impose on model developers and deployers. These choices affect who can participate in AI development and on what terms.

There is also a geopolitical dimension. Jurisdictions that make training easier may accelerate domestic AI development. Jurisdictions that strengthen rightsholder control may favour established creative sectors and encourage licensing markets. Neither approach is costless. A permissive regime may erode incentives for creators if substitution effects are severe. A restrictive regime may entrench incumbents and slow experimentation. The policy challenge is not to choose creativity over computation, but to prevent one from extracting value from the other without sustainable rules.

That is why AI and intellectual property should be understood as a question of political economy. The dispute is not merely about whether copying occurred. It is about how value flows through digital ecosystems: from creators to platforms, from archives to models, from public knowledge to private systems and back again. Legal doctrine matters because it allocates bargaining power within those flows.

The likely destination is a more conditional IP order

The long-term outcome will probably be neither the collapse of intellectual property nor its effortless extension into every corner of machine learning. Instead, AI is pushing legal systems towards a more conditional order. Rights will remain, but their scope will depend more heavily on context: whether a use is transformative or substitutive, whether access was lawful, whether rights were reserved, whether outputs are traceably derived, whether human creativity is meaningfully present and whether transparency obligations have been met.

This may disappoint those seeking a single grand principle. But it reflects the reality of the technology. AI is not one activity. It is a stack of activities: scraping, storing, training, fine-tuning, prompting, generating, ranking, editing and distributing. Each interacts with different rights and different policy objectives. Treating them all the same would be administratively convenient and economically crude.

The deeper significance of the current moment is therefore institutional. Intellectual-property law is being asked to govern systems that learn from culture as a raw material and then re-enter culture as competitors, tools and intermediaries. Whether that leads to a healthier knowledge economy depends less on slogans about disruption than on the patient work of specifying rights, exceptions, evidence and obligations. The age of AI will not end intellectual property. It will make its design choices impossible to ignore.

Sources & Further Reading

  1. 1.
  2. 2.
  3. 3.
  4. 4.
  5. 5.
  6. 6.
  7. 7.
  8. 8.
  9. 9.
  10. 10.
artificial intelligenceintellectual propertycopyrightpatentsfair usetext and data miningauthorshipregulation
The engine behind the Signal

Where this connects to Society OS

The Sovereign Intelligence Hub is the free, open front door of Society OS — the sovereign operating system that turns the ideas you just read into working governance. Where this piece names a problem, Society OS is building the machinery to solve it: AI agents that act with your authority, trust you can verify, and compliance that runs as code.

The 42-Protocol Stack

The governance engine beneath every article — led by the Sovereign Trinity: Human-Twin-Agent identity, HEARTrank trust, and WISE Contracts that execute law, not just code.

F-ACT — the open agent standard

The vendor-neutral framework for governing AI agents before they act: Authority, Scope, Data, Audit, Revocation — free to read, cite and implement.

The Sovereign Platform

Put it to work: govern a fleet of AI agents with verifiable authority, tamper-evident evidence, and compliance-as-code across your whole operation.

Explore membershipRead the F-ACT standard

Continue Reading

More from the Sovereign Intelligence Hub

Intellectual property after the model frontier
AI & Intellectual Property

Intellectual property after the model frontier

18 min read

Who Owns What When AI Helps Create
AI & Intellectual Property

Who Owns What When AI Helps Create

14 min

Standards Patents Became the New Border Checkpoints
AI & Intellectual Property

Standards Patents Became the New Border Checkpoints

11 min read

Why the next patent wars will be fought over training rights, not algorithms
AI & Intellectual Property

Why the next patent wars will be fought over training rights, not algorithms

11 min read

The Quiet Power of Publication as a Patent Defence
AI & Intellectual Property

The Quiet Power of Publication as a Patent Defence

11 min read

Standards Patents Are Becoming Industrial Policy by Other Means
AI & Intellectual Property

Standards Patents Are Becoming Industrial Policy by Other Means

11 min read

Never miss a signal

Weekly intelligence, no noise

The Sovereign Intelligence Hub — Society OS

© 1989–2026 Society OS Pty Ltd. All rights reserved.