Open knowledge has entered an infrastructural phase
For two decades, the case for open knowledge was framed in expansive terms: transparency would improve accountability, shared research would accelerate discovery, and public access to information would broaden participation. Those claims still matter. But the debate has matured. The central question is no longer whether openness is desirable in principle; it is whether societies can build durable institutions around it.
That shift matters because open knowledge is now embedded in critical systems. Government statistics underpin fiscal and social policy. Scientific repositories support everything from epidemiology to climate modelling. Open mapping and geospatial data feed logistics, disaster response and urban planning. Public-interest archives and encyclopaedic resources increasingly serve as default reference layers for both citizens and machines. In other words, the commons is no longer peripheral. It is operational.
Yet operational systems have operating costs. Servers, curation, metadata, legal review, security, version control and community moderation are not incidental extras. They are the mechanisms through which openness becomes useful. The data brief emerging from current evidence is stark: many institutions have become better at releasing information than at sustaining it.
Open data creates value only when it remains legible, discoverable and governable long after first publication.
Publication is up, but usability is uneven
One reason open-data policy can appear more successful than it is lies in a familiar measurement problem. Counting portals, files and downloads is easier than assessing whether data can be integrated across systems or interpreted without specialist mediation. International assessments have repeatedly found improvements in disclosure and availability, while also highlighting gaps in quality, timeliness and interoperability.
The OECD's work on digital government and data policy has stressed that value comes not simply from opening datasets, but from their strategic reuse across institutions and sectors. Similarly, the European Commission's framework around high-value datasets emphasises machine readability, standard formats and APIs because these features reduce the transaction costs of reuse. Without them, nominally open data often functions more like an archive than an active public asset.
The same pattern appears in research. Open-access mandates and repository growth have substantially widened access to publications and, increasingly, to underlying data. But access does not guarantee reproducibility or broad reuse. Documentation standards vary, file formats decay, and incentives for careful curation remain weak. A spreadsheet deposited to satisfy a grant condition is not the same thing as a maintained, well-described dataset that others can confidently build upon.
The result is an asymmetry between release and reliability. Policy often rewards the first act of publication; users depend on the slower work of maintenance.
The funding model is the weakest link
Most commons infrastructure has a lopsided financial profile. The initial launch can often be supported through grants, special programmes or political momentum. Ongoing stewardship is harder to finance because it produces diffuse, system-wide benefits rather than a neat, attributable return. This is a classic public-goods problem.
UNESCO's Recommendation on Open Science recognises this directly, calling for long-term investment in the infrastructures, skills and governance that make knowledge sharing possible. The recommendation is notable because it treats openness not as a one-off compliance exercise but as an ecosystem requiring durable support. That framing aligns with practical experience from libraries, archives and public repositories, where the unglamorous costs of preservation and curation frequently determine whether a resource remains usable after the launch announcement has faded.
Open data creates value only when it remains legible, discoverable and governable long after first publication.
The problem is not only the quantity of funding but its structure. Short cycles encourage projects to prioritise novelty over robustness. Core maintenance, migration to new standards, and user support are routinely underfunded. In software, this would be recognised as technical debt. In data commons, it is better understood as institutional debt: a backlog of stewardship obligations that accumulates quietly until trust begins to erode.
This helps explain why seemingly mature open resources can be fragile. Users see a stable interface; maintainers see rising storage bills, outdated schemas, moderation burdens and uncertain staffing. The economics are real even when the access price is zero.
Legal openness is necessary, but not sufficient
Licensing has been one of the major achievements of the open knowledge movement. Standard public licences reduced uncertainty, clarified reuse rights and enabled cross-border collaboration. They remain indispensable. But legal openness alone does not resolve the harder questions emerging around privacy, data protection, community rights and the asymmetric capacity to exploit shared resources.
This is especially evident for data about people, places and ecosystems. Open release may be lawful yet still create downstream risks, including re-identification, discriminatory inference or extractive use by actors with far greater analytical power than the communities represented in the data. The World Bank and other development institutions have increasingly emphasised responsible data governance for precisely this reason: availability and protection must be designed together rather than treated as sequential concerns.
There is also a tension between universal openness and legitimate restrictions tied to Indigenous data governance, public safety or sensitive environmental information. The point is not that the commons should shrink. It is that a mature commons needs sharper boundary-setting and clearer stewardship norms. In many domains, the future lies in governed access, data trusts, secure research environments and tiered permissions, not in a simplistic binary between closed and open.
The strongest commons are not the least governed; they are the most clearly governed.
Standards do more economic work than most policy debates admit
When policymakers discuss openness, they often focus on licensing and publication obligations. Yet the largest efficiency gains may come from less visible choices: metadata schemas, persistent identifiers, shared vocabularies and interoperable formats. Standards reduce the labour required to combine information from multiple sources. They also improve resilience by making datasets less dependent on the tacit knowledge of a small group of maintainers.
The FAIR principles, first articulated in Scientific Data in 2016, became influential precisely because they translated an abstract commitment to sharing into operational qualities: findable, accessible, interoperable and reusable. FAIR does not demand that every dataset be openly downloadable. Rather, it insists that data should be organised in ways that support discovery, interpretation and appropriate reuse. That distinction has been important for fields balancing openness with confidentiality.
In economic terms, standards lower search costs, switching costs and coordination costs. They make it easier for smaller institutions, civic groups and researchers in lower-resource settings to participate. Conversely, poor metadata and idiosyncratic formatting impose a hidden tax on reuse that favours larger organisations with enough capacity to clean and reconcile messy data at scale.
This is why standards policy is distributional policy. It affects who can actually benefit from ostensibly open resources.
Research data shows both the promise and the friction
Nowhere is the open-knowledge agenda more consequential than in science. Publicly funded research has moved decisively towards greater openness in publications, methods and data. Initiatives from funders and multilateral bodies have pushed institutions to treat data sharing as part of research integrity, not merely dissemination. The benefits are substantial: faster verification, wider collaboration and more cumulative scholarship.
The strongest commons are not the least governed; they are the most clearly governed.
But the research sector also illustrates the frictions. Data management plans are now common, yet compliance varies in quality. Repositories differ in preservation capacity. Sensitive data in health and social science often require controlled access. And the labour of preparing data for future use is still unevenly recognised in hiring and promotion.
The lesson is not that open science has stalled. It is that the next gains are likely to come from professionalising the support layers around it: data stewards, repository managers, common identifiers, clearer citation norms and long-term infrastructure funding. The European Open Science Cloud and related initiatives point in this direction by trying to create federated systems rather than isolated silos. Their significance lies less in any single platform than in the institutional model they imply: shared standards, shared governance and shared maintenance burdens.
That model matters beyond academia. It offers a template for any domain where the value of shared information compounds over time.
Public-sector data is most valuable when boring things work well
Government open-data programmes are often judged by headline releases: company registers, procurement data, transport feeds or environmental records. These are important, and in some cases transformative. Yet users often derive the greatest practical value from datasets that are reliable rather than dramatic: consistent geographic codes, timely updates, clear methodological notes and stable URLs.
The Open Data Charter has long argued that openness should be embedded by design in public administration. That principle sounds procedural, but its implications are strategic. If data publication is bolted on at the end of an administrative process, quality problems multiply. If data is created with reuse in mind from the start, the costs of sharing fall and the benefits rise.
This is one reason statistical agencies remain among the most important institutions in the knowledge commons. Their authority stems not merely from release, but from methodology, continuity and documentation. In an era of synthetic media and proliferating low-quality information, those traits become more valuable, not less. Open knowledge requires trusted producers as well as open channels.
For public administrations, then, the agenda is surprisingly unglamorous: stronger records management, common identifiers, procurement rules that preserve data portability, and publication workflows built into routine operations. The boring things are what make openness durable.
The global gap is shifting from access to capacity
Open knowledge was once discussed largely as a problem of access: too many journals behind paywalls, too much public information trapped in inaccessible formats, too many barriers to reuse. Those obstacles remain. But a second divide has become clearer. Even where resources are nominally open, the capacity to curate, combine and exploit them is highly uneven.
This matters internationally. Lower-income countries may face acute constraints in storage, bandwidth, skilled staffing, legal support and preservation infrastructure. Local institutions can end up contributing data to global systems without receiving proportional analytical or economic benefits in return. UNESCO's open-science framework and broader debates on data equity both point to the same concern: openness without capability can reproduce dependency rather than reduce it.
Capacity is not just a technical matter. It includes governance literacy, negotiating power, and the ability to set terms around local priorities. A healthy commons should not simply maximise extraction from all contributors. It should broaden the range of actors who can shape standards, steward resources and derive value from them.
A data commons that others can mine but few can govern is open in form and unequal in practice.
Artificial intelligence raises the stakes for provenance
A data commons that others can mine but few can govern is open in form and unequal in practice.
The rise of data-hungry machine learning has given open knowledge a new strategic significance. Publicly accessible text, images, code and metadata now serve not only human readers but computational systems trained to summarise, classify and generate. That has sharpened old questions about attribution, consent and sustainability while adding new ones about provenance and traceability.
For the commons, the opportunity is clear: better discoverability, richer tools for analysis and wider circulation of public knowledge. The risk is equally clear: resources assembled through volunteer labour or public funding may be consumed at industrial scale without commensurate support flowing back to the institutions and communities that maintain them. At the same time, unreliable or weakly documented inputs can propagate errors through downstream systems much faster than in a purely human research cycle.
This makes provenance infrastructure more important. Persistent identifiers, version histories, citation practices and transparent terms of reuse are no longer niche administrative features. They are the basis for accountability in an increasingly automated information environment. If machine-scale reuse becomes a normal condition of openness, then the commons will need machine-readable governance as much as machine-readable data.
What a more durable commons would look like
The practical agenda emerging from current evidence is less about grand declarations than institutional design. First, public and philanthropic funders need to treat data stewardship as core infrastructure, not discretionary overhead. Multi-year funding for repositories, archives and standards bodies is more valuable than a proliferation of short-lived pilots.
Secondly, success metrics should move beyond counts of released datasets or repository deposits. More meaningful indicators would include update regularity, metadata completeness, citation rates, interoperability, documented reuse and preservation performance. These are harder to measure, but they track actual public value more closely.
Thirdly, governance must become more explicit. That means clearer licensing where openness is appropriate, but also clearer rules for restricted access, privacy protection, community oversight and responsible downstream use. The aim should be legitimate reuse, not indiscriminate exposure.
Fourthly, institutions should invest in stewardship professions. Librarians, archivists, research software engineers, data curators and records managers are often treated as supporting actors. In reality, they are part of the production function of knowledge itself.
Finally, standards coordination deserves far more attention. The most valuable commons are those that can interoperate across borders, sectors and time. That requires persistent identifiers, common vocabularies and stable institutional forums in which those conventions can evolve.
From openness to stewardship
The strongest argument for open knowledge has never been that information wants to be free. It is that societies function better when important knowledge can be scrutinised, combined and reused beyond narrow institutional boundaries. That argument remains persuasive. But it now needs a companion principle: what is shared must also be maintained.
The future of the data commons will therefore hinge on stewardship. Not stewardship as paternal gatekeeping, but as a disciplined commitment to preservation, documentation, standards and fair governance. Openness lowers barriers to entry; stewardship lowers barriers to meaningful use. One without the other produces disappointment.
For policymakers, funders and institutions, the implication is plain. The age of symbolic openness is ending. What comes next is an era in which the quality of common knowledge infrastructure will shape scientific capacity, democratic accountability and economic participation. The commons is no longer simply a publishing question. It is a state capacity question, a research capacity question and, increasingly, a strategic one.
If that sounds less romantic than the early rhetoric of digital liberation, it is. But it is also more useful. Mature public infrastructure is rarely built on slogans. It is built on budgets, standards and trust.



