For years, the standard argument in intellectual property ran along familiar lines. Should software patents be narrower. Are abstract claims choking follow-on innovation. Do patent thickets favour incumbents over laboratories, universities and smaller firms. Those questions still matter. But by mid-2026 they no longer capture the pressure point. The strategic contest is moving upstream, towards the legal conditions under which machines are allowed to read, classify and learn from the technical record.
This is not only an argument about generative AI. It is an argument about industrial research itself. Materials science, chemistry, medical diagnostics, semiconductors and advanced manufacturing increasingly depend on models trained on papers, patents, standards, regulatory filings, lab notebooks, maintenance logs and other structured or semi-structured corpora. In that setting, control over training inputs is becoming a surrogate for control over invention itself.
From patentability to trainability
The old software-patent battles were concerned with the output of computation: the claimed method, system or product. The new conflict concerns the input side: what can be lawfully ingested, mined and recombined in order to generate hypotheses, optimise design choices or automate prior-art analysis. Patent law alone does not answer that question. Instead, a stack of adjacent rights does: copyright in texts and images, sui generis database rights in Europe, trade-secret protection for proprietary technical corpora, contract restrictions in platform terms, and practical access controls that turn legal uncertainty into a market barrier.
This matters because most machine learning used in R&D does not require a breakthrough algorithm. It requires lawful, reliable access to a body of technical knowledge at sufficient scale and quality. When that access is fenced off, smaller research actors may find that they are formally free to invent but practically unable to train the systems that modern invention increasingly requires.
Control over training inputs is becoming a surrogate for control over invention itself.
The public patent record was built for disclosure, not enclosure
Patents are often defended on a simple bargain: temporary exclusivity in exchange for disclosure. Society tolerates the monopoly because the invention enters the public record in a structured form that others can study, improve around and eventually use. In principle, machine learning should strengthen that bargain. A corpus of millions of patent documents, classifications, citations and prosecution histories is one of the richest maps of technical progress ever assembled.
Yet the legal and practical architecture around that record is more brittle than the bargain suggests. Patent texts may be public, but the surrounding data infrastructure is uneven. Full-text availability, bulk download permissions, formatting, metadata quality and usage terms vary across jurisdictions and repositories. If the patent system is serious about disclosure as a social exchange, the relevant question in 2026 is no longer merely whether a human can read a document. It is whether a machine can lawfully and efficiently learn from it.
Europe’s text-and-data-mining settlement is only a partial answer
Control over training inputs is becoming a surrogate for control over invention itself.
The European Union has moved further than many jurisdictions in recognising text and data mining as a legitimate activity through the Digital Single Market Directive. But its compromise remains awkward. Research organisations and cultural heritage institutions benefit from a mandatory exception in some contexts, while a broader exception for text and data mining can be reserved against by rightholders. In practice, that creates a tiered landscape. Some actors enjoy relative certainty; many others must navigate reservations, licences and the possibility that contract terms will narrow what the law appears to permit.
For independent innovators, this is an institutional, not merely doctrinal, problem. A right that exists only for well-lawyered institutions or for actors able to negotiate bespoke licences is not a neutral research rule. It is a sorting mechanism. It decides which organisations can turn the published record into a training substrate and which must remain dependent on intermediaries.
The database right is a quiet but powerful gatekeeper
Outside specialist circles, Europe’s sui generis database right receives less attention than patents or copyright. It deserves more. Technical knowledge is increasingly valuable not because a single document is exclusive, but because millions of records have been selected, curated and normalised into usable datasets. Database protection can therefore become a quiet gatekeeper over access to industrial knowledge, especially where the legal status of extraction and re-utilisation is uncertain for large-scale computational uses.
One need not claim that database right is universally overbearing to see its strategic role. If patent texts, scientific abstracts, standards information, maintenance records or legal annotations are wrapped into proprietary databases, the economic centre of gravity shifts from inventing to controlling the searchable, trainable corpus. The result is not a classic patent monopoly. It is a knowledge-infrastructure monopoly assembled from adjacent rights.
Trade secrets are expanding into the training layer
At the same time, trade-secret law is doing more work than many patent debates acknowledge. The EU Trade Secrets Directive, like analogous regimes elsewhere, protects commercially valuable information kept secret through reasonable steps. In a world of data-intensive research, those steps often cover curation pipelines, lab outputs, synthetic datasets, feedback labels, model-evaluation sets and domain-specific ontologies. None of this is unusual. Firms have long guarded process knowledge. What is new is the migration of secrecy from the production line into the learning pipeline itself.
That creates a subtle asymmetry. A smaller actor may publish an invention through the patent system, thereby contributing to the public corpus. A larger actor, by contrast, can supplement the same public corpus with vast proprietary training and evaluation data held as trade secrets. The former discloses into common space; the latter learns from common space and private reservoirs simultaneously. The legal architecture is therefore not simply rewarding invention. It is rewarding the ability to accumulate non-public complements around disclosed invention.
Contract is the new enclosure
If patent lawyers once focused on claims, prosecution and infringement, they now have to pay equal attention to terms of use, API conditions, anti-scraping clauses, click-through restrictions and bulk-access licences. The new enclosure is contractual before it is technical. Many of the most consequential limits on training are not written in statute but embedded in access agreements that define who may download, parse, index or reuse information at scale.
The new enclosure is contractual before it is technical.
This should trouble anyone who takes disclosure seriously. A public-facing repository that is legally or operationally inaccessible for computational analysis is only partially public. The distinction matters because machine reading is no longer a speculative add-on. In patent searching, freedom-to-operate analysis, claim drafting and R&D scouting, machine processing is becoming routine infrastructure. When contract terms selectively suppress such uses, disclosure ceases to be a general public good and becomes a conditional privilege.
The new enclosure is contractual before it is technical.
Independent innovators need a different defensive strategy
Against this backdrop, defensive patent strategy needs revising. The conventional playbook emphasises narrow but credible claims, careful continuation practice, selected publication and, where useful, open licensing. Those tools remain relevant. But they are not enough when the competitive bottleneck is machine-readable access to technical knowledge.
Defensive filing now has to think about machine readability as well as claim scope. That means drafting specifications, abstracts, sequence listings, examples and terminology in ways that improve discoverability and interoperability. It means paying attention to structured disclosures, standard identifiers and open metadata. It also means deciding which supporting materials should be published outside the patent itself under licences that clearly permit computational reuse, even where the core invention remains protected. The goal is not maximal openness in every case. It is to prevent one’s own disclosures from becoming dead text in a world where useful knowledge is increasingly parsed by machines before it is read by people.
Open licensing should move from ideology to infrastructure
Open licensing debates often become moral dramas: openness versus ownership, commons versus enclosure. That framing is too blunt for 2026. The more practical question is how to design licences and publication practices that preserve reciprocal access to technical knowledge without requiring innovators to abandon all proprietary claims. In many sectors, the answer will involve layered strategies: patents on specific applications, open release of non-core datasets or taxonomies, and explicit permissions for text and data mining over selected materials.
Such arrangements are not a cure-all. They can be gamed, ignored or outcompeted by actors with deeper data reserves. But they can alter bargaining power. A well-designed open licence for technical documentation, benchmarking data or interoperability schemas can make a smaller innovator legible to collaborators and harder to isolate within a closed ecosystem. Openness, in this narrower sense, is less a moral statement than an anti-extractive design choice.
Patent pools and standards teach a useful lesson
There is a parallel here with the long, imperfect history of standards and patent pools. The OECD and competition authorities have repeatedly noted that pools can reduce transaction costs and clear fragmented rights, but only under governance that avoids exclusion and collusion. The same principle should inform data and training access. Shared technical corpora, reference datasets and machine-readable repositories can widen participation in innovation, but only if they are governed as common infrastructure rather than as bottlenecks dressed up as public resources.
Defensive filing now has to think about machine readability as well as claim scope.
The lesson is not that every corpus should be compulsory-licensed or publicly funded. It is that fragmented rights over indispensable inputs can create anti-competitive effects even when no single patent appears abusive. Competition policy, information policy and IP policy therefore need to be read together. A patent system that discloses generously but tolerates enclosure of the means to analyse disclosures is only half functioning.
Courts will struggle because the categories were built for another era
Judges and regulators are being asked to solve this problem with categories inherited from older disputes. Copyright doctrine was not built with industrial-scale model training primarily in mind. Database right was not designed as a grand policy lever for computational R&D. Trade-secret law does not distinguish neatly between legitimate protection of know-how and strategic opacity around foundational data resources. Patent law, meanwhile, still tends to assess inventive contribution at the level of the claimed output rather than the accessibility of the learning environment that made the output possible.
That mismatch helps explain the uneasy litigation around AI training and content use, visible in newspaper and other copyright disputes as well as quieter private negotiations. Courts can decide individual controversies. They are less well placed to build a coherent architecture for scientific trainability across sectors. The risk is path dependence: a patchwork of settlements and precedents that gradually privileges those already controlling the richest corpora.
What this means for the politics of invention
The political economy of patents is usually described as a struggle between exclusion and diffusion. In the emerging training economy, the more revealing divide is between formal rights and functional capability. Many actors may retain the formal freedom to read patents, papers and manuals. Far fewer will possess the legal certainty, technical access and compute-adjacent infrastructure required to transform those materials into models that accelerate invention.
This shifts the practical meaning of independence. An independent innovator is not only someone who owns their claims or avoids predatory licensing terms. It is someone who can build, or reliably access, a lawful learning stack: discoverable public disclosures, machine-readable corpora, rights-clear datasets, interoperable metadata and publication norms that do not punish computational use. Without that stack, patent ownership risks becoming ceremonial. One may hold title to an invention while lacking equal access to the knowledge systems that determine the next one.
The patent bargain needs a machine-age update
None of this requires dismantling patents. On the contrary, it requires taking their justificatory logic more seriously. If patents are granted in exchange for public disclosure, then the public side of the bargain must be evaluated under contemporary conditions. In a machine-mediated research system, disclosure should mean more than nominal visibility. It should include practical usability for search, analysis and training, subject to safeguards for privacy, security and genuinely confidential information.
That implies a modest but important intellectual shift. The central question for IP policy is no longer only how much exclusivity inventors deserve. It is also whether the legal architecture around disclosed knowledge preserves a competitive capacity to learn. The future patent wars, in other words, may not turn on who owns the cleverest algorithm. They may turn on who is permitted to learn from the accumulated record of human technique, and on what terms.
That is an unusual place for patent politics to arrive. Yet it follows directly from the system’s original promise. A society that asks inventors to disclose should not quietly permit the disclosed world to become untrainable except for those able to purchase access to every surrounding layer of rights. If that happens, extraction will not look like classic infringement. It will look like lawful asymmetry, organised through databases, secrecy and contract, until independent innovation finds that the public record is public in theory and rationed in practice.


