The meaning of alignment is widening
In the early years of machine-learning ethics, alignment was often framed in narrow technical terms: can a system optimise for what its designers intended rather than for a badly specified proxy? That framing emerged from real problems. Recommendation systems can maximise engagement by promoting outrage; prediction systems can exploit spurious correlations; large language models can produce plausible falsehoods when trained to imitate text rather than to represent truth. In each case, the machine appears to follow instructions while violating the purpose behind them.
Yet the public argument has moved. As general-purpose models spread across administration, education, software, research and security, alignment can no longer be understood solely as a property of a model. It is increasingly a property of a socio-technical system: the training data, the reward functions, the deployment setting, the human overseers, the procurement rules, the legal standards and the institutions that determine who bears risk. The technical question—how to steer a model—now sits inside a larger constitutional one: who gets to define the destination?
Alignment is shifting from a question of model obedience to a question of institutional legitimacy.
This is not an argument against technical research. On the contrary, work on robustness, interpretability, uncertainty estimation and oversight remains indispensable. But it is an argument against treating those tools as sufficient. The more capable the system, the more consequential the surrounding governance becomes. A model aligned with its operator’s goals may still be misaligned with public interest.
Why technical fixes alone cannot settle ethical conflict
A persistent fantasy in technology policy is that contested social values can be converted into engineering objectives and solved through optimisation. That is tempting because it promises clarity: write a specification, train a model, test against benchmarks, iterate. But many ethical disputes are not failures of precision. They are genuine conflicts between legitimate aims.
Consider fairness. A substantial academic literature has shown that different statistical definitions of fairness can be mutually incompatible in real-world settings. The United States National Institute of Standards and Technology has stressed that bias in AI is not only computational but systemic and human, shaped by historical inequities and institutional context. No amount of parameter tuning can remove the need for judgement about which harms matter most, for whom and under what conditions.
The same is true of privacy, safety, free expression and access. A content-moderation system may reduce abuse while suppressing lawful speech. A medical triage model may improve efficiency while reducing clinician discretion. A fraud-detection system may lower losses while imposing burdens on already marginalised groups. These are not bugs to be eliminated once and for all. They are trade-offs that must be governed, justified and periodically revised.
This is one reason the Organisation for Economic Co-operation and Development and UNESCO have both framed trustworthy AI around principles such as accountability, transparency and human rights rather than around a single technical metric. Alignment in the civic sense requires procedures for handling disagreement, not merely better loss functions.
From individual users to collective stakes
The vocabulary of alignment often assumes a dyad: a user issues a prompt, a model responds, and success means the output matches the user’s legitimate intent. That picture is increasingly incomplete. Many deployments affect people who are not the user and who may never know a system has shaped the decision. Credit scoring, welfare administration, hiring, predictive policing and border control all involve third parties whose rights and opportunities can be altered without direct participation.
Alignment is shifting from a question of model obedience to a question of institutional legitimacy.
Collective effects go further. Generative models can change labour markets by automating tasks faster than institutions can retrain workers. They can alter information ecosystems by flooding search, social media and public consultation processes with synthetic text and images. They can influence scientific practice by accelerating discovery while also scaling error, irreproducibility or fabricated references. In each case the relevant unit of ethical analysis is not the isolated transaction but the system-wide consequence.
The European Union’s AI Act reflects this broader view by distinguishing uses according to risk and by imposing obligations that vary with context. Whether one agrees with all its provisions or not, the underlying point is sound: alignment cannot be judged abstractly. The same model may be benign in one setting and intolerable in another. Governance must therefore focus not only on capability but on function, domain and impact.
Power is the missing variable
Many alignment discussions are, at heart, discussions about control. Who can deploy powerful models? Who sets thresholds for acceptable risk? Who audits the evidence? Who can contest a harmful decision? When these questions are ignored, ethics becomes decorative.
The political philosopher’s insight here is simple: values do not operate in a vacuum. They are interpreted by organisations with incentives. A firm judged on speed to market may underinvest in evaluation. A public agency facing budget pressure may automate discretion too quickly. A military bureaucracy may prioritise strategic advantage over explainability. Even well-intentioned actors can create misalignment when institutional rewards point elsewhere.
The Bletchley Declaration, signed by a wide range of governments in 2023, recognised that frontier AI could generate serious, even catastrophic, harms and called for international co-operation on safety. Its importance lay less in any single policy proposal than in acknowledging that the highest-stakes systems raise governance questions that no developer can settle alone. Where capabilities have systemic implications, oversight must match scale.
A system can be perfectly aligned with its owner’s incentives and dangerously misaligned with society’s interests.
This is especially urgent where asymmetries of information are large. Those affected by AI systems often cannot inspect training data, evaluate model updates or understand the provenance of outputs. Without external scrutiny, the power to define alignment remains concentrated in the hands of deployers. Democratic governance, by contrast, requires avenues for redress, transparency proportionate to risk and institutions capable of independent evaluation.
Interpretability matters, but so do due process and contestability
Interpretability research aims to understand why a model produced a particular output, or how internal representations correspond to concepts. This work is valuable. It can reveal hidden failure modes, improve safety and support debugging. But interpretability is not the same as accountability.
A citizen denied a service by an automated system may care less about a neuron-level explanation than about whether the decision was lawful, reviewable and open to appeal. A worker screened out by an algorithmic hiring system needs to know whether criteria were fair and whether corrections are possible. In these contexts, due process can matter more than model transparency in the narrow technical sense.
The White House blueprint for an AI Bill of Rights, though not binding law, captured this distinction by emphasising notice, explanation and human alternatives. Likewise, the Council of Europe’s Framework Convention on Artificial Intelligence and Human Rights, Democracy and the Rule of Law ties AI governance to established civic protections. These frameworks suggest that alignment should be measured partly by whether people can contest machine-mediated decisions and whether institutions can be held answerable for them.
A system can be perfectly aligned with its owner’s incentives and dangerously misaligned with society’s interests.
That has practical consequences. Documentation, audit trails, incident reporting and clear lines of responsibility are not bureaucratic afterthoughts. They are alignment mechanisms. They ensure that failures can be traced, challenged and corrected rather than absorbed into opaque systems.
The frontier risk debate should not eclipse present harms
Public debate on AI ethics often swings between two poles. One focuses on immediate harms: discrimination, surveillance, labour displacement, misinformation and concentration of power. The other focuses on frontier risks from highly capable systems: loss of control, strategic instability and large-scale misuse. These camps are frequently portrayed as rivals. They need not be.
The two agendas share underlying concerns about reliability, oversight and institutional preparedness. A system that hallucinates legal citations today may illuminate how models fail under pressure tomorrow. Weak governance in consumer applications may foreshadow larger failures in critical infrastructure or defence. Conversely, frontier safety work can produce methods—such as red-teaming, evaluations and anomaly detection—that improve current deployments.
The mistake is to assume that one horizon cancels the other. The interim report of the International Scientific Report on the Safety of Advanced AI, prepared under an international process involving leading researchers, stresses uncertainty while urging evidence-based governance. Such caution is sensible. Serious policy should resist both complacency and theatrical alarm. Ethical attention must be distributed across timelines, with safeguards proportionate to capability and consequence.
The dispute between present harms and future risks is often overstated; weak institutions make both more likely.
Open models, closed models and the politics of diffusion
No alignment debate is complete without considering how access to powerful models should be governed. Wider diffusion can promote research, competition and public scrutiny. It can also enable misuse, from fraud and cyber offence to biosecurity-relevant assistance, depending on capability thresholds and accompanying controls. Restrictive access can reduce some risks while concentrating power and limiting independent verification.
This is not a binary contest between openness and secrecy. Different components—weights, training data, system cards, evaluation methods, interfaces and deployment permissions—can be governed in different ways. The policy question is how to distribute capability without losing the capacity to manage harms. Much depends on context: scientific reproducibility may call for more openness; high-risk deployment may justify stronger restrictions; public-sector procurement may require independent access for auditors even where full public release is inappropriate.
The wider lesson is that alignment is shaped by political economy. Diffusion determines who can inspect, adapt, contest and benefit from AI systems. A regime that privileges speed and concentration may look efficient at first, yet prove brittle if trust erodes or external scrutiny is blocked. Durable alignment requires governance arrangements that balance innovation, safety and legitimacy rather than sacrificing one to the other.
International co-ordination is necessary, but hard
Artificial intelligence is developed through global supply chains of chips, data, research talent and cloud infrastructure, yet governed through fragmented national systems. That mismatch complicates alignment. Standards developed in one jurisdiction may be ignored in another. Firms may relocate risky activities. Governments may see safety rules through the lens of industrial strategy or geopolitics.
The dispute between present harms and future risks is often overstated; weak institutions make both more likely.
Still, co-ordination is possible. The Hiroshima Process launched by the G7, the work of the OECD on AI principles, and multilateral discussions around advanced AI safety all indicate an emerging architecture of soft law, standards and information sharing. These instruments are imperfect. But they matter because they can create common expectations around testing, incident disclosure, provenance, red-teaming and high-risk uses.
What seems unlikely, at least in the near term, is a single global authority for AI. More plausible is a layered model: domestic regulation for local accountability, international standards for interoperability, and targeted agreements for especially sensitive areas such as autonomous weapons, cyber operations and biosecurity-related capabilities. Alignment under such conditions will remain uneven. The task is to reduce the most dangerous gaps while preserving room for democratic variation.
What good governance would look like
If alignment is becoming a governance problem, what follows? First, evaluation must be continuous rather than episodic. Models change after deployment, and so do their uses. Independent testing, post-market monitoring and mandatory reporting of serious incidents are more credible than one-off certification.
Secondly, risk should be allocated to those best placed to manage it. Developers, deployers and procurers each control different levers. Assigning all responsibility to end users is untenable where systems are opaque or unavoidable. Public procurement rules can be especially powerful, requiring evidence of safety, auditability and rights protection before systems enter education, welfare or policing.
Thirdly, high-risk domains need stronger participatory mechanisms. Civil-society groups, affected communities, domain experts and labour representatives often identify failure modes that technical teams miss. Consultation will not remove conflict, but it can surface assumptions before they harden into infrastructure.
Fourthly, governance must preserve human capability rather than merely inserting nominal human oversight. A human reviewer presented with machine outputs at scale may become a rubber stamp. Effective oversight requires time, training, authority and the option to reject automation.
Finally, institutions need strategic patience. It is better to phase deployment in sensitive areas than to rush systems into settings where errors are costly and reversibility is low. The history of digital governance suggests that once flawed infrastructures become normal, reform is slow and politically difficult.
The real alignment question
For years, the defining thought experiment in alignment asked whether highly capable systems would faithfully pursue human goals. That remains a serious inquiry. But in practice, the more immediate question is more political and more ordinary: which humans, acting through which institutions, under which constraints, are entitled to set those goals?
That question has no purely technical answer. It requires law, standards, public reasoning and administrative capacity. It also requires intellectual honesty about uncertainty. Some harms are measurable; others are diffuse. Some risks are immediate; others emerge only when systems become deeply embedded. Ethical governance must therefore be iterative, evidence-based and open to revision.
There is a temptation, especially in fast-moving fields, to treat governance as a brake and alignment as a software feature. Both views are too narrow. Governance is the mechanism by which societies decide what kinds of optimisation are acceptable, where automation belongs, and what red lines should not be crossed. Alignment, in that fuller sense, is not only about making machines do what we mean. It is about ensuring that the systems we build remain answerable to the public worlds they reshape.





