What is a data commons?
The phrase data commons is often used loosely to describe any large pool of shared information. In practice, it means something more specific: data that is made accessible for collective use under rules designed to preserve public value over time. That distinguishes a commons from a one-off data dump, a proprietary platform or a narrow bilateral sharing agreement. A commons is not simply an archive. It is a social and institutional arrangement.
The same is true of open knowledge more broadly. Open access research, open educational resources, open-source software and public-interest datasets all belong to a wider ecosystem in which knowledge is treated less as a scarce commodity and more as a shared foundation for learning, innovation and accountability. The internet made dissemination cheaper. It did not solve the harder problem of stewardship.
That distinction matters. Information can be technically open and still functionally closed if it is hard to find, poorly documented, legally ambiguous or accessible only to those with substantial computing resources. Conversely, carefully governed shared datasets can enable public benefits even when some access conditions are needed to protect privacy or community rights.
Open data becomes a commons only when access is matched by stewardship, accountability and the capacity to use it.
From openness to infrastructure
For much of the past two decades, openness was framed as a corrective to secrecy. Governments launched open-data portals. Universities expanded open-access publishing. International organisations encouraged data-sharing to improve development outcomes and scientific reproducibility. These efforts mattered, and many still do. The World Bank’s data catalogue, the European Union’s open data framework and the OECD’s work on public-sector information each helped establish openness as a norm rather than an exception.
Yet the first wave of open-data enthusiasm sometimes treated publication itself as success. Too often, datasets were released without sustained funding, metadata standards, update schedules or user support. Some portals became cluttered repositories of stale spreadsheets. Others provided valuable transparency but little practical utility for researchers, journalists or local communities.
That experience has shifted the debate. The central question is no longer whether information should be open where possible; it is how to build institutions that make shared knowledge reliable, inclusive and durable. In that sense, open knowledge is increasingly understood as infrastructure. Like roads, libraries or clean water systems, it requires maintenance, rules and public investment.
Why commons create public value
When they work well, data commons reduce duplication and expand the range of people who can participate in inquiry. Public health researchers can compare outcomes across jurisdictions. Urban planners can map transport gaps and environmental risks. Investigative journalists can trace procurement patterns. Civic groups can monitor air quality or budget spending. Teachers and students can build on a body of freely accessible research rather than encounter paywalls at every turn.
Scientific collaboration offers some of the clearest examples. The Human Genome Project established early norms around rapid data release that shaped later thinking about open science. More recently, international moves towards open research data have been reinforced by bodies such as UNESCO and the OECD, which argue that broad access to scientific knowledge can accelerate discovery and improve reproducibility. The logic is straightforward: when knowledge can be inspected, reused and challenged, collective learning improves.
Open data becomes a commons only when access is matched by stewardship, accountability and the capacity to use it.
Public administration can also benefit. Open contracting data, for instance, can make procurement systems easier to scrutinise and compare. The Open Contracting Partnership and the World Bank have documented how standardised disclosure can strengthen oversight and competition. Here, openness is not a philosophical gesture. It is a practical means of reducing information asymmetries.
The hidden labour behind openness
One reason the commons metaphor is useful is that it draws attention to maintenance. Shared resources do not look after themselves. Datasets need cleaning, labelling, version control and documentation. Legal terms need to be clear. Interfaces need to be usable. Sensitive material may require access tiers, de-identification or secure research environments. Communities contributing data may need representation in governance bodies and clear routes for redress.
Much of this work is undervalued because it sits between disciplines and budgets. Academic prestige often rewards novel findings more than curation. Government departments may finance initial publication but not long-term stewardship. Philanthropic funding can help establish repositories but not always sustain them. The result is a paradox: societies increasingly depend on shared information systems while often treating their upkeep as an afterthought.
This hidden labour includes social as well as technical work. Translating a dataset into public value requires intermediary institutions: libraries, universities, standards bodies, investigative newsrooms, community organisations and public-interest technologists. Without such intermediaries, formally open resources tend to be captured by those already rich in expertise, compute and organisational capacity.
The most important resource in a knowledge commons is not data alone, but the institutions that make data legible, trustworthy and usable.
Governance is the real battleground
Any commons raises a basic political question: who sets the rules? In data governance, this includes decisions about collection, consent, licensing, access, benefit-sharing, security and acceptable uses. These are not merely administrative details. They shape who can participate and who bears the risks.
The instinctive answer is often maximal openness. But absolute openness is not always compatible with fairness. Indigenous data, health records, geolocation traces and administrative microdata may have immense research value while still requiring careful restrictions. The aim should not be openness at any cost, but the greatest degree of openness consistent with rights, safety and legitimate community control.
That is why the language of governance has become more prominent in international debates. UNESCO’s Recommendation on Open Science emphasises inclusiveness and equity alongside access. The Organisation for Economic Co-operation and Development has similarly stressed trustworthy data-sharing frameworks. In parallel, work on Indigenous data sovereignty, including the CARE Principles developed by the Global Indigenous Data Alliance, argues that communities should have a say not just in whether data is shared, but in how benefits and authority are distributed.
These frameworks point to a mature view of the commons. Shared resources need rules, and good rules are not the enemy of openness. They are what make openness sustainable.
Privacy, power and the limits of anonymisation
The most important resource in a knowledge commons is not data alone, but the institutions that make data legible, trustworthy and usable.
One persistent misconception is that privacy can be solved simply by removing names. In reality, re-identification risks can remain high when multiple datasets are linked or when location, demographic and behavioural records create unique patterns. The rise of machine learning and large-scale data integration has only sharpened this concern.
Regulators and researchers have therefore moved beyond a binary distinction between open and closed. Between those poles lies a spectrum of controlled access models: data enclaves, secure labs, trusted research environments and synthetic datasets, each with trade-offs. The United Kingdom’s Office for National Statistics, for example, has developed secure access mechanisms for sensitive public data, reflecting a wider recognition that public-value research may require guarded rather than unrestricted sharing.
Power matters as much as privacy. Open datasets are often harvested and recombined by well-resourced actors far more easily than by citizens or small organisations. This can create extractive dynamics in which communities supply data while others capture the economic or analytical value. A data commons worthy of the name must confront these asymmetries, not conceal them behind the rhetoric of openness.
Standards make sharing possible
Commons fail when information cannot travel. Technical and semantic standards are therefore central to the open-knowledge ecosystem. Common metadata, persistent identifiers, interoperable formats and shared vocabularies allow datasets and publications to be discovered, linked and compared. Without such standards, openness fragments into isolated silos.
The FAIR principles, published in Scientific Data in 2016, have become especially influential. FAIR stands for Findable, Accessible, Interoperable and Reusable. Importantly, FAIR does not mean everything must be fully open. It means that data should be structured so that authorised users can discover and use it properly. This distinction helps reconcile openness with legitimate access controls.
Standards also lower the cost of participation. The more predictable the formats and licensing terms, the easier it is for civic groups, smaller research teams and local governments to contribute and benefit. In that sense, standards are not bureaucratic add-ons. They are an inclusion mechanism.
Who gets left out?
Open knowledge is often presented as inherently democratic. Sometimes it is. But the distribution of benefits can be highly uneven. Large research institutions in wealthier countries typically have better connectivity, staffing, language resources and analytic capacity than universities, municipalities or civil-society groups operating with fewer resources. A dataset released in principle to the whole world may in practice be most usable by a narrow elite.
This problem is not confined to the global level. Within countries, communities affected by environmental hazards, policing or housing inequality may lack the tools to interpret the very data collected about them. Accessibility therefore depends on translation, training and context, not simply legal permission. Public libraries, schools and local intermediaries can play a decisive role here, yet they are rarely centred in national data strategies.
Language is another barrier. Much open research remains overwhelmingly Anglophone. Metadata standards may travel globally, but knowledge does not become genuinely open if it is inaccessible to those excluded by language, disability or digital infrastructure. The politics of the commons begins with access, but it does not end there.
A commons that only the well-resourced can use is not a commons in any meaningful civic sense.
A commons that only the well-resourced can use is not a commons in any meaningful civic sense.
Open science and the future of research
Few areas illustrate the promise and tensions of open knowledge as clearly as science. Funders and universities increasingly support open access publishing, preprints, data repositories and reproducibility practices. The European Commission’s open science policies and UNESCO’s global framework both reflect a growing consensus that publicly funded research should be as accessible as possible.
Still, implementation remains uneven. Publishing costs can shift from readers to authors, creating new inequalities. Data-sharing mandates can burden researchers who lack technical support. Sensitive disciplines face particular dilemmas around consent and misuse. And metrics still reward speed and novelty more than replication, documentation or curation.
A more resilient model would treat research outputs as part of a knowledge commons that includes methods, code, protocols and negative findings, not just finished papers. That requires investment in repositories, persistent identifiers, governance mechanisms and incentives that recognise maintenance as scholarly work. The point is not simply to make science visible. It is to make it cumulative.
What governments should do
If open knowledge is infrastructure, public policy should reflect that fact. Governments can begin by prioritising a smaller number of high-value datasets and repositories rather than measuring success by volume alone. Quality, documentation and update reliability matter more than sheer quantity. Public-sector information should be released in machine-readable formats with clear licences and robust metadata wherever feasible.
Stewardship needs stable funding. Long-term repositories, national libraries, archives and statistical agencies require budgets for maintenance, security and user support, not just launch events. Procurement rules can encourage open standards and interoperability so that public information does not become trapped in closed technical systems.
Governments should also distinguish between openness and trustworthiness. Sensitive data may need tiered access, independent oversight and sanctions for misuse. Community participation should be built into governance, especially where data concerns Indigenous peoples, marginalised groups or local environmental resources. In such cases, consultation after collection is too late; legitimacy has to be designed in from the start.
What a healthy commons looks like
A healthy open knowledge ecosystem has several recognisable traits. Its resources are easy to discover. Their provenance is clear. Their formats are interoperable. Their licences are understandable. Their governance is transparent. Their stewards are funded. Their users are broader than a technical priesthood. And their benefits are not wholly detached from the communities from which data is drawn.
None of this is glamorous. But the history of successful public infrastructure rarely is. Libraries became indispensable not because books existed, but because institutions learned how to catalogue, preserve and circulate them at scale. Data commons demand a comparable shift in ambition: from release to stewardship, from ideology to institution-building, from technical possibility to civic design.
The stakes are rising because digital societies increasingly govern themselves through information. Decisions about transport, health, climate adaptation, education and research funding all depend on the quality and accessibility of shared knowledge. Where those systems are robust, accountability and innovation become easier. Where they are brittle or extractive, trust erodes.
Open knowledge, then, is not a niche concern for librarians, data scientists or legal specialists. It is part of the constitutional fabric of the digital age. The real test is not whether societies can publish more data. It is whether they can build commons that are worthy of the public.



