Few technologies in criminal justice have travelled so quickly from managerial promise to political controversy as predictive policing and risk assessment. The sales pitch was straightforward: where officials had long relied on hunches, habit and uneven discretion, statistical systems would offer consistency. Judges could use recidivism scores to inform bail, sentencing or parole. Police commanders could use forecasts to deploy patrols where crime was deemed most likely. Public agencies, under pressure to do more with less, found the language of prediction attractive.
Yet after more than a decade of scrutiny, the central question is no longer whether these systems can process large quantities of data. It is whether they can do so in a way compatible with equal treatment, due process and meaningful accountability. In many jurisdictions, the answer has become increasingly uncertain. Researchers have documented disparate impact; civil liberties groups have challenged opacity and surveillance; some cities have imposed bans or moratoria; and courts have struggled to decide what evidentiary weight an algorithmic output should carry.
The long shadow of COMPAS
No instrument has become more emblematic than COMPAS, a proprietary risk-assessment tool used in parts of the United States to estimate the likelihood of reoffending. Public attention sharpened in 2016 when ProPublica reported that the tool appeared to produce racially skewed errors: Black defendants who did not reoffend were more likely than white defendants to be labelled high risk, while white defendants who did reoffend were more likely to be labelled low risk. The report provoked immediate methodological debate, but it achieved something more important than a settled verdict on one model. It forced wider recognition that algorithmic fairness is not singular.
As critics and defenders soon showed, a system cannot usually satisfy every fairness criterion at once when base rates differ across groups. That point, developed in technical work by researchers including Alexandra Chouldechova and others, did not exonerate criminal justice scoring. Rather, it exposed a deeper problem: every model embeds normative choices. Deciding whether to minimise false positives, preserve calibration, or equalise error rates is not a merely technical exercise. It is a distributive judgement about who bears risk, error and state coercion.
The central problem is not that an algorithm makes mistakes when humans would be flawless; it is that a statistical tool can give old patterns of unequal enforcement the appearance of scientific neutrality.
Prediction built on unequal histories
Risk scores are only as good as the data and assumptions on which they are built. In criminal justice, that creates immediate difficulty. Arrest records, stops, calls for service and conviction data are not neutral inventories of wrongdoing. They are traces of policing practice, prosecutorial priorities, plea bargaining and neighbourhood surveillance. If certain communities have historically been watched more closely, their residents will be overrepresented in the data, not necessarily because they offend more, but because more of what they do is observed, recorded and pursued.
This matters especially when tools rely on proxies that correlate with race or poverty even when race is excluded. Residential instability, employment history, educational background, family criminal history and prior contact with the justice system can all function as conduits for structural inequality. The language of “race-neutral” modelling can therefore mislead. A model may omit protected characteristics and still reproduce their effects through variables that reflect segregated housing, unequal schooling, selective enforcement or cumulative disadvantage.
The practical consequence is that a score can appear to describe a person while actually summarising features of the social environment and institutional history around them. In this sense, a risk score is not simply discovered; it is constructed.
The central problem is not that an algorithm makes mistakes when humans would be flawless; it is that a statistical tool can give old patterns of unequal enforcement the appearance of scientific neutrality.
The feedback loop problem
Place-based predictive policing reveals the problem even more starkly. Such systems typically forecast where certain offences are likely to occur using historical incident data. But when those forecasts direct officers towards particular blocks or districts, police presence increases there. Increased presence tends to generate more stops, more observations and more recorded incidents. Those new records then feed the model, reinforcing the apparent wisdom of the original forecast.
This is the classic feedback loop described by legal scholars and policy analysts. The model does not merely measure reality; it helps make the reality it later cites as evidence. Areas with less police attention generate less data and may appear quieter than they are. Areas already under scrutiny become ever more legible to the system. The result can be a self-confirming cycle in which statistical confidence grows as epistemic quality declines.
For street crime and low-level offences, the loop is especially potent because recorded crime depends heavily on detection effort. A burglary report may enter the system regardless of patrol density; a stop for loitering, drug possession or disorderly conduct often does not. Where the state looks hardest, it tends to find most.
Disparate impact without explicit discrimination
The legal and political significance of these dynamics lies in disparate impact. A tool need not classify by race on its face to burden racial minorities disproportionately. In cities marked by residential segregation and unequal policing, location itself can serve as a powerful proxy. A risk model that sends more patrols to historically marginalised neighbourhoods may produce formally equal rules while imposing materially unequal exposure to police contact.
European institutions have become increasingly alert to this issue. The European Union Agency for Fundamental Rights has warned that artificial intelligence in law enforcement may affect fundamental rights including non-discrimination, privacy and effective remedy. Similar concerns animate the OECD’s recommendations on trustworthy artificial intelligence, which emphasise transparency, accountability and respect for the rule of law. The point is not that all statistical assistance is unlawful. It is that equality harms can emerge from ordinary administrative practice when data-intensive systems are introduced into already unequal institutions.
What courts have and have not accepted
The most cited judicial encounter with risk scoring remains State v. Loomis, decided by the Wisconsin Supreme Court in 2016. The court permitted use of a COMPAS assessment at sentencing but attached warnings: the score could not determine the sentence; the proprietary nature of the tool limited what a court and defendant could know about it; and concerns about group-level prediction, including gendered and other embedded factors, required caution. The decision was often read as a compromise. In fact it illustrated a profound judicial unease. The court allowed the tool into the process while acknowledging that its operation was partially shielded from adversarial scrutiny.
That tension persists. In evidentiary terms, an algorithmic score occupies an awkward category. It is not quite expert testimony, not quite documentary fact, not quite policy guidance. Yet it can exert gravitational force in a courtroom, particularly when presented in numerical form. Judges and parole boards may treat a score as a disciplined second opinion. Defendants may struggle to challenge it when the underlying model is proprietary, difficult to explain, or based on inaccessible data. This asymmetry cuts against familiar due process principles, especially where liberty turns on claims that cannot be meaningfully contested.
A risk score is not a fact about a person in the way a fingerprint or a blood test might be; it is a probabilistic inference built from choices about data, proxies and objectives.
A risk score is not a fact about a person in the way a fingerprint or a blood test might be; it is a probabilistic inference built from choices about data, proxies and objectives.
The authority of numbers
One reason these systems are so influential is cultural rather than technical. Quantification carries prestige. A number seems cleaner than a narrative, less vulnerable to whim and prejudice. In public administration, metrics often function as a shorthand for rationality itself. But criminal justice decisions are not only predictive; they are moral and legal judgements. They concern blame, proportionality, rehabilitation and the permissible use of state power. A model can estimate correlation with a future event; it cannot decide what degree of uncertainty justifies detention, surveillance or punishment.
This is where the language of “decision support” can obscure as much as it clarifies. Officials may insist that algorithms merely inform human judgement. Yet decades of research on automation bias suggest that people often defer to machine outputs, especially under time pressure or institutional constraint. A nominally advisory score can shape outcomes even when no rule compels deference. The burden then shifts subtly: instead of authorities justifying coercive action from first principles, individuals may find themselves expected to disprove a statistical suspicion.
Bans, moratoria and retrenchment
That scepticism has translated into policy reversal in several cities. In the United States, some police departments have ended or suspended predictive policing programmes after criticism from auditors, community groups and local lawmakers. Santa Cruz prohibited predictive policing and facial recognition in 2020. New Orleans’ city council placed limits on facial recognition use and heightened oversight of surveillance technologies. Elsewhere, programmes in Los Angeles and Chicago were curtailed or abandoned after sustained controversy over efficacy, civil liberties and racial bias.
These measures vary in scope, and they do not amount to a comprehensive prohibition on all analytics in policing. Still, they mark an important shift. The early framing presented predictive systems as inevitable modernisation. The newer approach treats them as exercises of public power requiring explicit democratic authorisation, auditable documentation and, in some cases, prohibition where safeguards are judged insufficient.
Why auditing matters, and why it is not enough
In response to criticism, many scholars and watchdog groups have called for regular algorithmic audits, impact assessments and public reporting. These can be valuable. Independent evaluation can reveal whether a model performs differently across groups, whether predictions drift over time, and whether users are applying outputs in ways that exceed their intended purpose. Documentation can also clarify training data, performance metrics, procurement decisions and redress mechanisms.
But audits are not a cure-all. Much depends on what is being measured, who conducts the review, and whether agencies disclose enough information to make testing meaningful. A model can pass a narrow technical audit while remaining objectionable in law or policy. It may be well calibrated and still rely on tainted proxies. It may show acceptable average accuracy and still produce intolerable harms for a minority of individuals. It may perform as designed in a system whose basic incentives are themselves unequal.
There is also a danger that audit language becomes a legitimating ritual: once measured, a system appears governed. Yet some problems are upstream of model performance. If the target variable reflects discriminatory enforcement, then improving prediction may simply improve reproduction of that discrimination.
Where police presence is already concentrated, predictive systems can turn attention into evidence and evidence into renewed attention, creating a loop that looks empirical while deepening distortion.
Trade secrets and public power
Proprietary secrecy has sharpened objections in both sentencing and policing contexts. When private intellectual property shields core details of a model used by the state, the ordinary logic of accountability becomes difficult to sustain. Defence counsel may not know how variables are weighted. Independent researchers may lack access to training data. Communities subject to algorithmic deployment may be unable to assess claims about efficacy or fairness. This has prompted repeated criticism from legal scholars, journalists and civil rights advocates.
The issue is not merely transparency for curiosity’s sake. Public power generally requires reasons that can be inspected and challenged. If a person’s liberty is constrained partly because a score deems them high risk, then the basis of that judgement should be open to adversarial testing. Secrecy may be commercially ordinary in other sectors; in criminal justice it collides with constitutional and democratic principles.
The evidentiary limits of algorithmic outputs
A sober legal approach would treat algorithmic outputs as contestable administrative artefacts, not objective truths. Their relevance depends on context. A score generated for resource allocation should not automatically migrate into sentencing. A model trained on one population may not be valid for another. Correlations observed at group level may say little about a specific individual standing before a court. These are familiar evidentiary cautions, yet the technological framing often weakens them by suggesting a degree of precision the underlying methods cannot support.
Moreover, recidivism itself is a slippery endpoint. Different tools predict rearrest, reconviction or return to custody, each shaped by institutional behaviour as much as by personal conduct. Rearrest may reflect policing intensity; reconviction depends on prosecution and plea dynamics; breach of supervision conditions often tracks the stringency of monitoring. A model that predicts such outcomes is therefore predicting interaction with the justice system, not pure criminal propensity.
What a more realistic understanding requires
The strongest critique of predictive policing is not that statistical methods have no place in public administration. It is that criminal justice is a domain where measurement is inseparable from power. Data are generated through coercive institutions; outcomes are influenced by unequal exposure to those institutions; and decisions affect liberty, stigma and citizenship. Under those conditions, claims of objectivity require exceptional scrutiny.
A more realistic understanding begins by abandoning the fiction that these tools stand outside politics. They encode choices about safety, acceptable error, and whose uncertainty counts. They can sometimes standardise practice, but they can also standardise injustice. Used incautiously, they displace responsibility: the official points to the score, the vendor to the data, the data to history, and history to no one in particular.
That is why the debate has moved beyond accuracy. The crucial issues are institutional. Who authorises deployment? What evidence establishes necessity and proportionality? Can affected people inspect, challenge and appeal outcomes? Are the systems producing public benefits that outweigh foreseeable equality and due process harms? In many places, those questions were asked only after deployment. The backlash now under way reflects a belated recognition that predictive systems in policing and sentencing are not mere administrative upgrades. They are exercises in governing by inference, and they deserve to be judged as such.
Where police presence is already concentrated, predictive systems can turn attention into evidence and evidence into renewed attention, creating a loop that looks empirical while deepening distortion. That is the limit of the risk-score imaginary. It promises to tame uncertainty with numbers, but in criminal justice the numbers often inherit the uncertainty, prejudice and selective vision of the institutions that produce them. The trial facing predictive policing, then, is not only legal. It is epistemic and democratic: whether a society should permit opaque statistical estimates to shape coercive state action when the data beneath them are marked by unequal power from the start.
