The long prehistory of openness
The idea of a knowledge commons predates the internet by centuries, but its modern form begins with a simple political proposition: information held in the public interest should, wherever possible, be accessible to the public. In practice, that notion developed unevenly. Libraries, archives, scientific societies and statistical offices all expanded access in the 19th and 20th centuries, yet much public information remained difficult to obtain, expensive to reproduce or governed by restrictive legal rules.
A decisive early shift came with freedom-of-information regimes. Sweden’s constitutional tradition is often cited as the earliest example, but in the contemporary era the landmark was the United States Freedom of Information Act of 1966, later strengthened by amendments. Its significance went beyond journalism or litigation. It established a durable principle that public records were not inherently the property of bureaucracies. Comparable laws would spread across democracies over the following decades, creating a legal foundation for the later open-data movement.
Open knowledge, then, did not begin as a technical trend. It began as an institutional argument about who gets to see, use and question information created in the course of public life.
1989–1991: the web makes sharing scalable
The next inflection point was architectural. In 1989 Tim Berners-Lee proposed the World Wide Web at CERN, and by 1991 the first website was online. The importance of the web to open knowledge was not merely that it put information online. It introduced a low-friction model for publishing, linking and retrieving documents across institutions and borders. Knowledge that had once been siloed in libraries, agencies or specialist networks could now be referenced and recombined at much larger scale.
CERN’s later decision to place the web software into the public domain, announced in 1993, mattered as much as the invention itself. It ensured that the web would develop as a common layer rather than a proprietary network controlled by a single gatekeeper. That choice shaped everything that followed: open standards bodies, public repositories, collaborative encyclopaedias and government portals all depended on a shared, interoperable web.
Open knowledge became transformative only when publication ceased to be scarce and linking became routine.
The web did not automatically produce openness. It simply made openness technically feasible on a scale that had not previously existed.
1990s science and software lay the groundwork
During the 1990s, two adjacent movements gave the future data commons both a philosophy and a practical toolkit. The first was open-source software. The GNU project had begun earlier, but the decade saw free and open-source methods gain broader institutional traction, demonstrating that distributed collaboration could produce robust public digital goods. The second was the expansion of open scientific repositories, especially in fields such as physics, where preprint culture began to flourish online.
The arXiv repository, founded in 1991, became a model of what networked scholarly communication could look like when dissemination was treated as a public good rather than a scarce commodity. Researchers could circulate findings quickly, establish precedence and widen access beyond the holdings of elite libraries. Although open access publishing remained contested, the repository model showed that digital networks could support a commons for knowledge production, not merely knowledge consumption.
These developments also familiarised institutions with licences, metadata and version control. Such tools might sound technical, but they are central to any commons. Shared resources are only genuinely reusable when users understand the terms of access, provenance and modification.
Open knowledge became transformative only when publication ceased to be scarce and linking became routine.
2001–2002: open licensing creates a legal grammar
If the web solved distribution, licensing helped solve permission. One of the defining problems of the early internet was that digital copying was easy but lawful reuse was often unclear. In 2001 a new generation of standardised public licences began to address that gap. The following year saw the formal launch of widely used open licences designed to let creators signal, in machine-readable and human-readable terms, how others could reuse their work.
This mattered enormously for the growth of a commons. Without clear permissions, digital abundance can still produce legal uncertainty. Open licences created a scalable legal grammar for sharing text, images, educational resources, research outputs and, eventually, datasets. They helped shift the default question from “may I use this?” to “under what conditions may I use this?”
The significance of this change is easy to underestimate. Commons are not sustained by goodwill alone. They require institutions that reduce transaction costs. Standard licences did precisely that, giving schools, museums, universities, researchers and civic groups a common framework for release and reuse.
2003–2007: open access and open knowledge become organised movements
The early 2000s brought a more explicit politics of openness. In scholarly publishing, the Budapest Open Access Initiative of 2002 articulated the case for free online access to research literature, arguing that digital networks had made a new settlement possible between authors, readers and publishers. This was followed by further declarations and mandates that pushed universities and funders towards more open dissemination of publicly funded research.
Alongside this, advocates of open knowledge began to define openness beyond academia. The argument expanded from journal articles to cultural materials, educational resources, government information and data. The Open Definition, first developed in the mid-2000s, helped codify a stricter conception of openness: access was not enough; reuse, redistribution and interoperability also mattered.
By this point, the intellectual architecture of the modern commons was becoming clear. A resource counted as open not merely because it was visible, but because it could legally and technically circulate.
A dataset behind restrictive terms is visible; it is not necessarily part of a commons.
2007–2009: public-sector data enters the frame
The phrase “open data” gained political force when governments began treating administrative and statistical datasets as assets for public release rather than internal records. A crucial intellectual marker came in 2007, when a group of advocates articulated a set of principles for open government data, stressing completeness, timeliness, machine-readability, non-discrimination and licence-free reuse.
The following years saw public administrations begin to operationalise the idea. In 2009 the United States launched Data.gov, one of the first large national open-data portals. The same year the United Kingdom’s data portal followed, signalling that open government data had become a mainstream instrument of administrative reform rather than a niche cause.
These portals varied in quality, and many early releases were patchy, poorly documented or politically selective. Yet their importance was symbolic as much as practical. They reframed government information as reusable infrastructure. Civil servants, journalists, researchers and software developers could begin drawing from common pools of transport data, procurement records, geospatial information and performance metrics.
The open-data agenda also connected transparency to innovation. That linkage was sometimes overstated, but it helped broaden support. Information released for accountability could also support services, analysis and new forms of public participation.
A dataset behind restrictive terms is visible; it is not necessarily part of a commons.
2010–2013: linked data, portals and the rise of machine-readability
Once data publication became a policy objective, attention shifted from access to usability. A scanned report posted online might satisfy a narrow interpretation of disclosure, but it does little for a genuine commons. The early 2010s therefore saw growing emphasis on structured formats, APIs, metadata standards and linked data.
Tim Berners-Lee’s five-star model for open data became influential because it captured a practical hierarchy: put data on the web, make it machine-readable, use open formats, assign identifiers and link datasets to other datasets. The model did not solve every problem, but it gave institutions a shared language for improving quality.
International bodies reinforced this trajectory. The World Bank, OECD and others expanded data platforms and publication standards. At the same time, city governments embraced open-data portals, often motivated by transport, planning and civic technology. The local scale mattered. It showed that a commons could be built not only from grand national databases but from mundane administrative records made consistently available.
Still, machine-readability brought a warning. Releasing data in bulk is not the same as making it meaningful. Documentation, context and stewardship remained indispensable.
2013–2016: global charters and the politics of default openness
By the middle of the decade, open data had moved from experimental policy to an international governance norm. The G8 Open Data Charter in 2013 set out principles including “open by default”, quality, usability and comparability. In 2015 a broader International Open Data Charter extended the framework beyond a small club of states. The language of default openness became especially influential because it reversed the burden of proof. Instead of asking why data should be released, institutions were pushed to justify why it should not.
This period also exposed the limits of the first wave. Many governments published portals with fanfare but weak maintenance. Datasets went stale. Funding was episodic. The legal status of some releases remained ambiguous. And publication often favoured less sensitive material while omitting records that would genuinely illuminate power, spending or performance.
Even so, the normative shift was real. Openness had become a benchmark of competent public administration, not merely a concession to activists. The commons was entering the language of statecraft.
The hardest part of open data is rarely release; it is sustained stewardship.
2016–2020: research crises, reproducibility and FAIR data
As open data matured in government, science confronted its own pressures. Concerns about reproducibility, access to underlying research materials and the opacity of some publication practices pushed institutions towards stronger data-sharing norms. In 2016 the FAIR Guiding Principles for scientific data management and stewardship crystallised a powerful idea: data should be findable, accessible, interoperable and reusable.
FAIR did not mean all data must be open without restriction. That was an important distinction, especially in fields involving personal, clinical or sensitive information. Rather, FAIR recognised that stewardship involves graduated access, good metadata and clear governance as well as public release. In effect, the scientific community refined the commons model by acknowledging that openness and responsibility must be designed together.
The hardest part of open data is rarely release; it is sustained stewardship.
This nuance proved valuable beyond science. It offered a way to think about data commons that are not simply public dumps, but managed ecosystems balancing access, ethics and utility.
2020–2022: pandemic data tests the commons
The covid-19 pandemic was an unusually severe stress test for open knowledge systems. Public dashboards, preprint servers, genomic databases and international statistical reporting all became central to understanding a fast-moving global emergency. The benefits of openness were plain: rapid sequencing and data-sharing accelerated scientific collaboration; public dashboards helped citizens and policymakers track trends; and open research practices shortened the time between discovery and dissemination.
Yet the shortcomings were equally visible. Definitions varied across jurisdictions. Data pipelines were uneven. Public communication often struggled to keep pace with uncertainty. Some releases were timely but not comparable; others were comparable but delayed. The episode underlined a hard truth: information commons cannot be improvised in a crisis. They depend on prior investment in standards, institutions and trust.
The pandemic also sharpened debate over who can access high-value datasets, on what terms, and with what protections. Open knowledge proved essential, but so did curation, quality control and privacy safeguards.
2022 onwards: digital public goods, data spaces and governance
In the present decade, the debate has moved beyond release towards governance. Policymakers increasingly ask how shared data resources should be stewarded across sectors and borders. European initiatives around common data spaces, multilateral interest in digital public goods, and renewed focus on public-interest technology all point in this direction.
This marks an evolution in thinking. The first era of open data was preoccupied with publication. The current era is more concerned with institutional design: who maintains datasets, who sets standards, how harms are mitigated, how communities participate, and how value generated from common resources is distributed. These questions are harder than launching portals, but they are also more consequential.
Open knowledge now sits at the intersection of competition policy, scientific collaboration, AI development, democratic accountability and cultural preservation. Training data for machine-learning systems, public-sector algorithms, digitised heritage collections and cross-border health research all raise the same underlying issue: whether society treats high-value information as fenced assets or as governed commons with public obligations.
What the timeline suggests
Looking back, the history of open knowledge is neither a simple march towards transparency nor a tidy story of technological liberation. It is a layered construction. Freedom-of-information laws created rights; the web provided infrastructure; open licences supplied legal clarity; repositories and portals demonstrated practical models; and standards made reuse feasible. At each stage, institutions had to decide whether information would be hoarded, monetised, disclosed reluctantly or shared as a public resource.
The lesson is that commons do not emerge naturally from digitisation. In some respects, digital systems make enclosure easier: access can be metered, formats can be proprietary and platforms can intermediate visibility. The endurance of open knowledge therefore depends on deliberate governance. Public funding, archival capacity, legal permissions, technical standards and civic oversight all matter.
That is why open knowledge is best understood not as an ideological accessory but as infrastructure. It supports scrutiny, scientific progress, better administration and broader participation in public life. But like all infrastructure, it can decay if maintenance is neglected or if incentives turn against openness.
The timeline also suggests a more modest, more durable ambition than some early rhetoric implied. Open data will not by itself fix democratic deficits or guarantee innovation. What it can do is widen the pool of actors able to inspect, combine and build upon socially valuable information. In a digital society, that is not a peripheral benefit. It is part of how a public sphere remains legible to itself.



