Hub
Analysis
The ethics gap is moving from models to institutions
Ethics & AlignmentAnalysis

The ethics gap is moving from models to institutions

By mid-2026, the hardest alignment failures look less like rogue objectives inside machines and more like unresolved conflicts among the organisations that deploy them.

Society OS Research10 July 202611 min read read

Key Insight: The frontier problem in AI alignment is increasingly constitutional rather than purely computational.

For a decade, the vocabulary of AI alignment has been dominated by an engineering image: a system has an objective, the objective is misspecified, and the machine pursues the wrong thing with unnerving competence. That image remains useful for technical safety. But as AI systems move into hiring, insurance, policing support, border management, education, welfare administration and clinical settings, a different picture is becoming harder to ignore. The central failures are often not those of a machine escaping human intent, but of institutions failing to specify, justify and revise the intentions they encode.

That distinction matters. A recommendation model that amplifies sensational content, a triage tool that reorders patients, or a screening system that narrows a shortlist may be behaving exactly as instructed. The difficulty is that the instruction itself is socially underdefined. Which values count: efficiency, equal treatment, proportionality, privacy, safety, solidarity, liberty, fiscal restraint, child protection, national security, professional autonomy, due process. In most real settings, these values are plural, partly incompatible and distributed across groups that do not agree on their ranking. Alignment is becoming a problem of public reason before it is a problem of machine behaviour.

From inner objectives to outer authority

The classic alignment story asks whether a model’s learned objectives diverge from the goals intended by its designers. That remains an important research agenda, especially for highly capable general-purpose systems. Yet deployment has exposed another layer of risk: organisations frequently lack legitimate procedures for deciding what the system should optimise in the first place. A public authority seeking to reduce fraud may impair access to lawful benefits. A hospital seeking throughput may shift burdens on to already disadvantaged patients. A school system seeking consistency may flatten legitimate pedagogical discretion.

In such cases, the machine is not misaligned with its operator. It may be perfectly aligned with a narrow managerial proxy. The deeper question is whether the operator is aligned with the public, the law and the affected communities. What fails in practice is often not optimisation but authorisation.

The return of constitutional questions

Seen this way, AI ethics begins to resemble constitutional design. Constitutions do not solve moral disagreement by discovering a single true objective function. They manage disagreement through rights, competences, procedures, review and appeal. They decide who may decide, on what grounds, with what evidence and subject to what constraints. Those are increasingly the decisive questions in automated governance.

The AI Act in the European Union, the OECD AI Principles and NIST’s AI Risk Management Framework all point, in different idioms, towards this institutional turn. None assumes that ethics can be reduced to a technical metric. Instead they emphasise risk management, human oversight, documentation, accountability and impact on fundamental rights. The most mature policy frameworks are, in effect, less interested in whether a model is intelligent than in whether an organisation can be trusted to use it under conditions of justified authority.

Without contestability, alignment is merely compliance with the preferences of the powerful.

Why value pluralism defeats tidy optimisation

Alignment is becoming a problem of public reason before it is a problem of machine behaviour.

The aspiration to align AI with “human values” always concealed a complication. There is no single, stable and operational set of human values available for upload. Political communities disagree, and rightly so, about the good life, acceptable risk, fairness between groups, acceptable trade-offs between liberty and security, and the limits of bureaucratic discretion. Even within one institution, executives, frontline professionals, regulators and citizens may hold conflicting but defensible views.

Technical methods such as reinforcement learning from human feedback can produce systems that better match annotator preferences or local norms of helpfulness and harmlessness. They cannot by themselves resolve normative conflict over public rules. A welfare office cannot ethically outsource the meaning of fairness to a labelling workforce. Nor can a court system reduce due process to a convenience score. In contested domains, alignment requires legitimate procedures for handling disagreement, not merely better preference aggregation.

The neglected role of administrative law

One underappreciated source of guidance lies outside computer science altogether. Administrative law has long dealt with decisions made at scale under conditions of uncertainty, constrained resources and unequal power. Its concerns are strikingly contemporary: reason-giving, non-arbitrariness, proportionality, reviewability, record-keeping and the right to be heard. These are not decorative safeguards attached after the real work of optimisation. They are part of what makes a decision acceptable in the first place.

For AI systems used in consequential settings, administrative principles offer a more realistic alignment target than vague appeals to beneficence. A system should not only be accurate. It should support traceable reasons, preserve avenues for challenge, and fit within legal competences. When the relevant values are contested, the ability to explain, contest and amend a decision may matter more ethically than marginal gains in predictive performance.

Three levels of alignment, only one of them technical

It helps to distinguish three levels. First is model alignment: whether outputs are robust, safe and broadly responsive to intended instructions. Second is organisational alignment: whether the institution deploying the model has chosen objectives, thresholds and escalation paths that are lawful and justified. Third is civic alignment: whether the broader social settlement around that deployment respects rights, democratic legitimacy and distributive fairness.

Most current debate remains concentrated on the first level because it is legible to engineers and investors. But many high-stakes harms emerge at the second and third. An impeccable model can be inserted into a broken process. Conversely, a modest model inside a well-designed institution with audit trails, appeal rights and human review may produce more legitimate outcomes than a more capable system deployed under opaque incentives.

Measurement is not settlement

The governance literature has made substantial progress on audits, benchmarks and impact assessments. These are necessary, but there is a temptation to treat measurement as resolution. If bias is quantified, if transparency documentation is completed, if a risk score is assigned, it can appear that the ethical work has been done. It has not. Metrics help reveal a conflict; they do not determine how the conflict should be settled.

What fails in practice is often not optimisation but authorisation.

Consider equal treatment and equal outcomes. A system can be adjusted to meet one statistical fairness criterion and thereby worsen another. The choice among criteria is not a hidden engineering parameter waiting to be tuned by evidence alone. It is a normative decision with legal and political implications. Similar tensions appear between privacy and fraud detection, between false positives and missed harms, and between standardisation and professional discretion. Ethics enters not after the dashboard, but at the point where institutions decide which losses are bearable and for whom.

Public sector deployment is the real stress test

These issues are sharpest in the state, because public bodies exercise coercive or gatekeeping power. They allocate entitlements, inspect compliance, issue sanctions and mediate access to essential services. In those domains, an aligned system is not merely one that follows instructions faithfully. It is one embedded in a process that citizens can regard as procedurally fair even when outcomes are unfavourable.

That requires more than a human in the loop as a ceremonial safeguard. Human oversight can become a rubber stamp if staff are overloaded, over-reliant on model outputs or institutionally discouraged from dissent. The relevant standard is not nominal human involvement but meaningful human agency: the ability, resources and authority to override, investigate and justify departures from automated recommendations.

What fails in practice is often not optimisation but authorisation.

Children, patients and migrants expose the limits of generic ethics

One reason broad ethical principles can feel unsatisfying is that vulnerability is domain-specific. Systems affecting children raise developmental and dependency concerns recognised in international human rights instruments. Clinical systems implicate informed consent, professional judgement and patient safety. Migration systems combine asymmetries of information, language barriers and high consequences for error. A generic commitment to fairness does not say enough about any of these settings.

The implication is not moral relativism. It is institutional specificity. Alignment should be specified against the duties of the particular domain: duties of care in health, heightened safeguards for children, strict legality and review rights in migration, academic freedom and pedagogical judgment in education. The same model behaviour may be acceptable in one context and indefensible in another because the surrounding obligations differ.

The audit function must become adversarial

If institutional alignment is the frontier, then oversight cannot remain a box-ticking exercise performed solely by the deploying organisation. Internal risk teams are useful, but they inhabit the same incentive structure as the deployment itself. More credible governance requires structured challenge: independent auditors, regulators with technical capacity, civil society scrutiny, and channels through which affected people can supply evidence of harm.

Without contestability, alignment is merely compliance with the preferences of the powerful.

This is where impact assessments, as discussed in legal scholarship and emerging regulation, matter most. Their value lies not in producing a perfect forecast, which is impossible, but in forcing organisations to articulate assumptions before deployment and to revisit them after contact with reality. The serious version of an impact assessment is less like a press release than a litigable administrative record.

Why alignment needs rights, not just preferences

Preference-learning approaches are powerful where the task is to infer what users find useful or congenial. They are much weaker where the purpose of governance is to limit what majorities, managers or agencies may do. Rights exist precisely because some interests should not be traded away by efficient aggregation. Privacy, non-discrimination, freedom of expression, access to remedy and the presumption against arbitrary interference are constraints on optimisation, not merely variables within it.

That is why rights-based frameworks have become increasingly central in Europe and in multilateral settings. They provide a language for hard limits where welfare calculations are indeterminate or politically manipulable. They also make clear that some forms of alignment are suspect. A perfectly compliant system that silently suppresses lawful dissent or denies meaningful recourse is not ethically aligned simply because users click through it without complaint.

The research agenda is widening beyond computer science

None of this diminishes the importance of technical safety research. Robustness, interpretability, uncertainty estimation, red-teaming and misuse prevention remain indispensable. But the centre of gravity is shifting. The most consequential questions increasingly require jurisprudence, public administration, political theory, sociology of organisations and sector-specific expertise. The problem is not merely to build systems that do what we mean. It is to establish institutions that can say what they mean in a manner that is lawful, revisable and publicly defensible.

That may sound slower and less elegant than the search for a general alignment solution. It is. Yet history suggests that complex societies do not solve moral conflict by discovering a master objective. They build procedures for living with disagreement while protecting the vulnerable from concentrated power. AI is now forcing that old lesson into a new register.

A more sober definition of success

By mid-2026, the most credible measure of progress in alignment may not be whether a model appears more obedient or more personable. It may be whether institutions using AI can specify their aims clearly, document trade-offs honestly, expose systems to adversarial review, and provide practical avenues of redress. In other words, success may look less like perfectly behaved machines and more like more accountable human governance around imperfect machines.

The philosophical turn in AI ethics, then, is not an abstract distraction from engineering. It is a recognition that high-capability systems are being inserted into political orders, not laboratories. Once that happens, alignment ceases to be only a question of control. It becomes a question of legitimacy. The ethics gap is moving from models to institutions, and the societies that understand this first will govern AI with fewer illusions and firmer constitutional instincts.

Sources & Further Reading

  1. 1.
  2. 2.
  3. 3.
  4. 4.
  5. 5.
  6. 6.
  7. 7.
  8. 8.
  9. 9.
  10. 10.
ethicsalignmentgovernanceinstitutionspublic-policyai-riskaccountability
The engine behind the Signal

Where this connects to Society OS

The Sovereign Intelligence Hub is the free, open front door of Society OS — the sovereign operating system that turns the ideas you just read into working governance. Where this piece names a problem, Society OS is building the machinery to solve it: AI agents that act with your authority, trust you can verify, and compliance that runs as code.

The 42-Protocol Stack

The governance engine beneath every article — led by the Sovereign Trinity: Human-Twin-Agent identity, HEARTrank trust, and WISE Contracts that execute law, not just code.

F-ACT — the open agent standard

The vendor-neutral framework for governing AI agents before they act: Authority, Scope, Data, Audit, Revocation — free to read, cite and implement.

The Sovereign Platform

Put it to work: govern a fleet of AI agents with verifiable authority, tamper-evident evidence, and compliance-as-code across your whole operation.

Explore membershipRead the F-ACT standard

Related Reading

The Accumulative Threshold: A Sovereign Paper on Civilizational Risk in the Age of Autonomous Intelligence
Civilisational Risk & Safety

The Accumulative Threshold: A Sovereign Paper on Civilizational Risk in the Age of Autonomous Intelligence

18 min read

Reputation Will Be the Hidden Infrastructure of the Agent Economy
Reputation Systems

Reputation Will Be the Hidden Infrastructure of the Agent Economy

18 min read

From Seed Vaults to Sequence Files: How Genetic Ownership Moved Upstream
Genetic Rights & Ownership

From Seed Vaults to Sequence Files: How Genetic Ownership Moved Upstream

11 min read

The Sovereign Intelligence Hub — Society OS

© 1989–2026 Society OS Pty Ltd. All rights reserved.