Alignment used to sound like a problem for philosophers and computer scientists working at opposite ends of the same corridor. One side asked how to formalise values; the other asked how to optimise systems without unleashing behaviour no one intended. By mid-2026, that framing feels incomplete. The question still matters, but it no longer captures where much of the real-world pressure lies. The practical challenge is not simply to make a model share human aims. It is to make claims about its behaviour legible enough for institutions to inspect, dispute and constrain.
This shift matters because advanced systems are now deployed across settings where acceptable behaviour is not reducible to a single utility function. Employment screening, public services, medicine, education, policing support and consumer interfaces all involve layered norms: fairness, privacy, proportionality, due process, professional duties and local law. A model can appear aligned in a laboratory sense while still acting illegitimately in an institutional one. The gap between those two senses of alignment is becoming the decisive fault line.
From moral philosophy to administrative proof
Classic alignment debates often begin with a grand ambition: if only engineers could specify human preferences more accurately, systems would reliably pursue what people actually want. That ambition has produced useful research, from reward modelling to constitutional prompting and reinforcement learning from human feedback. Yet every one of these methods hides a practical question. Whose judgements trained the system, under what conditions, and how would an outsider know when those judgements cease to apply?
The result is an important inversion. Instead of assuming that legitimacy flows from a sufficiently sophisticated objective, regulators and auditors increasingly ask for evidence of process. How were risks identified. What tests were performed. Which failure modes were anticipated. How are incidents recorded. Can affected people contest decisions or outputs. Alignment, in this sense, is becoming less like the discovery of a moral theorem and more like the production of an audit trail.
The operative question is no longer whether values can be encoded once and for all, but whether claims about behaviour can be examined and challenged.
Why behaviour beats intention
Institutions do not regulate intentions; they regulate conduct. This sounds obvious, but it marks a genuine departure from much technical discourse. A developer may sincerely aim to build a helpful model, and a model may score highly on helpfulness benchmarks, while still generating discriminatory recommendations, unsafe medical suggestions or manipulative conversational cues. Public legitimacy depends less on the proclaimed goal than on observable performance under conditions of stress, ambiguity and conflict.
That is why standards and risk frameworks have become so influential. The NIST AI Risk Management Framework emphasises governance, mapping, measurement and management rather than any single definition of ethical AI. The EU AI Act likewise treats many obligations as matters of documentation, testing, human oversight, record-keeping and post-market monitoring. These are not substitutes for moral reasoning. They are admissions that in complex societies, the first requirement for acceptable AI is not perfect virtue but accountable behaviour.
The politics of benchmarked virtue
The operative question is no longer whether values can be encoded once and for all, but whether claims about behaviour can be examined and challenged.
There is a temptation to treat evaluation scores as ethical facts. If a model is less toxic, more truthful, less biased or more robust than its predecessor, one might conclude that it is better aligned. Sometimes that is true. But benchmarks do political work as well as technical work. They define what counts as harm, which populations matter, what language varieties are visible, and which trade-offs deserve optimisation. A system can satisfy a benchmark and still violate a social expectation that was never measured.
This is not a minor defect. It is a structural feature of evaluation. Most benchmarks are selective simplifications created under time pressure and data constraints. They are often useful, sometimes indispensable, and never neutral. If an organisation says a system is aligned because it passed an internal suite of tests, the correct response is not disbelief but a further question: aligned to which operationalised norm, and who had standing to shape it.
Alignment failures are often procedural failures
Many notorious failures are not cases of a machine pursuing an alien objective with superhuman cunning. They are cases of ordinary institutional weakness. Training data are poorly documented. Domain assumptions drift. Human reviewers are overburdened. Escalation channels are unclear. Incidents are not shared across teams. Vendors and deployers each assume the other owns the residual risk. In those circumstances, even a technically capable model will behave in ways that are predictably unacceptable.
The language of procedure can sound bloodless, but that is precisely its value. It directs attention to repeatable controls. Was there a red-team exercise with adversarial inputs. Were there tests for minority-language performance. Were user complaints analysed for systematic patterns. Did procurement contracts allocate duties for monitoring and remediation. If the answer is no, then speaking loftily about aligned values obscures the fact that basic governance never took place.
Contestability is the missing norm
A well-aligned system is often imagined as one that gives the right answer. In democratic settings, however, a crucial property is whether the answer can be questioned. Contestability is not simply a legal afterthought. It is a core ethical requirement where systems influence rights, opportunities or reputations. People need routes to challenge an output, trigger review and present contextual information that the model could not have known.
This is particularly important because many harms emerge from cumulative small errors rather than spectacular malfunctions. A misleading eligibility suggestion, an overconfident summary in a case file, or a subtle pattern of deference toward some users and impatience toward others may be individually deniable and collectively corrosive. Systems need not only safeguards against catastrophic misuse; they also need institutional designs that let ordinary people expose ordinary unfairness.
A system can satisfy a benchmark and still violate a social expectation that was never measured.
Human oversight is not a magic stamp
A system can satisfy a benchmark and still violate a social expectation that was never measured.
Much policy language still relies on the reassuring phrase human in the loop. The difficulty is that humans in the loop can be ceremonial, rushed or structurally unable to disagree. Research on automation bias and human-machine interaction has long shown that oversight can degrade into rubber-stamping when operators lack time, expertise or authority. In those settings, adding a human checkpoint may increase liability management more than real control.
Good oversight therefore has design conditions. Reviewers need intelligible reasons for escalation, realistic workloads, training in model limits, and the power to override outputs without penalty. They need records that distinguish model suggestions from human decisions. They need feedback channels through which discovered errors reshape future deployment. Otherwise oversight serves as theatre: ethically comforting, operationally thin.
Plural values do not converge by themselves
The old alignment dream assumes that underneath disagreement lies a stable core of shared human values waiting to be captured. There is some truth in this. Most societies do reject arbitrary violence, deception and cruelty. Yet many real disputes are not misunderstandings on the way to consensus. They are durable political disagreements about equality, dignity, speech, risk, merit and the proper role of the state. No amount of scale or parameter tuning removes that fact.
For that reason, alignment in plural societies cannot mean silent convergence on one hidden moral order. It must mean explicit procedures for handling disagreement. Different applications will require different tolerances for error, different forms of explanation and different governance structures. A tutoring system, a clinical assistant and a border-control tool should not inherit the same ethical assumptions simply because they share an underlying model architecture.
The supply chain problem
Another reason alignment is turning into an audit question is that modern AI is assembled through layered supply chains. Foundation models, fine-tuning datasets, safety filters, application interfaces, retrieval systems and domain-specific instructions may come from different actors at different times. Responsibility is therefore distributed. A harmful output may result from interactions among components rather than from one obvious design decision.
This complicates both ethics and law. If a deployer relies on third-party evaluations, who verifies that they remain valid after adaptation to a new context. If a public authority procures a system trained elsewhere, how far can it inspect the assumptions embedded upstream. If post-deployment monitoring detects drift, which party bears the duty to correct it. Alignment here is inseparable from traceability. Without provenance and records, responsibility dissolves into the seams.
Evidence will matter more than promises
As AI systems become ordinary infrastructure, declarations of principle will carry less weight than demonstrated controls. That does not mean ethics codes are useless. It means they are increasingly judged by whether they generate observable practices: dataset documentation, incident reporting, independent testing, user redress, rollback procedures and periodic review. The vocabulary of trust is giving way to the machinery of assurance.
The deepest alignment problem may not be teaching machines human values, but deciding which humans, which institutions and which procedures count.
This is, in one sense, a healthy development. Grand claims about beneficent intent are easy to issue and hard to verify. By contrast, audit logs, model cards, risk registers and adverse-incident processes can be scrutinised, however imperfectly. They create a basis for external challenge. They also force organisations to state where they believe risks are tolerable. That is ethically valuable because many harms thrive in vagueness.
The danger of performative compliance
Still, one should not romanticise auditing. Bureaucratic proof can become ritual. Checklists may be completed mechanically. Documentation can be crafted to satisfy examiners while obscuring uncertainty. External auditors may lack access, expertise or incentives to probe deeply. A sophisticated organisation can appear exemplary on paper and remain brittle in practice. The history of finance, aviation and data protection offers ample warnings.
The answer is not to abandon assurance but to understand its limits. Good governance requires dynamic scrutiny, not static certification. It should include incident-led learning, adversarial testing, independent research access where lawful, and mechanisms for updating controls when social expectations shift. Ethical alignment is not a badge earned once. It is a continuing argument conducted through evidence.
Where philosophy still bites
If alignment is becoming administrative, philosophy has not disappeared. It has moved upstream and outward. Decisions about what to measure, what to disclose, who can contest outputs and what counts as acceptable risk are unavoidably normative. The choice between minimising aggregate error and protecting worst-case minorities is moral before it is technical. So is the question of when a system should remain undeployed because no governance arrangement can make its use legitimate.
This is where the field often understates the depth of the problem. The deepest alignment problem may not be teaching machines human values, but deciding which humans, which institutions and which procedures count. Technical systems inherit authority from social orders that are themselves contested. Auditability does not solve that contest. It merely makes the terms of disagreement more visible.
A more sober ambition for the next phase
By mid-2026, the most serious view of alignment is therefore less utopian and more constitutional. It asks not how to build a machine that is morally perfect, but how to place fallible systems inside arrangements that make abuse harder, error more detectable and correction more legitimate. This ambition is smaller than solving morality in code. It is also more compatible with how modern societies actually govern dangerous, useful technologies.
That may disappoint those who hoped for a clean technical breakthrough. Yet it has one large advantage: it treats alignment as something that must survive contact with institutions, incentives and dissent. In the long run, that is likely to be the standard that matters. Systems will be called aligned not when they claim to embody humanity, but when they can be shown, repeatedly and under challenge, to stay within bounds that a plural society has reason to accept.
The search for moral machines is not ending. It is being absorbed into the older human project of building accountable institutions.



