The Post Office Horizon scandal is the British archetype for what happens when institutional, legal, and evidentiary processes defer to machine output. Over 900 sub-postmasters were wrongfully convicted. At least 13 suicides have been linked to the scandal. Fujitsu's European head admitted at the 2024 inquiry that bugs had been known since 1999 and not disclosed in prosecutions. The harm pattern - errors concentrate in populations least able to contest them, institutional defence of the system outlasts internal evidence of its failure - is now repeating in algorithmic systems across UK and EU public services.
Why this matters
The capability is not the automation itself. It is the transfer of decisional authority into systems whose internal logic is opaque, whose errors are hard to detect individually, and whose effects are distributed across populations in ways that evade normal review.
The Horizon scandal was not about AI. But its template - a trusted automated system, institutional defence in the face of individual protest, legal processes that deferred to machine output as evidence, two decades of sustained harm before the pattern broke - is the template every UK citizen can reason from. The AI-specific cases (DWP fraud-detection, Dutch toeslagenaffaire, UK A-level 2020, facial recognition) are the same template with different technology.
The failures share one feature: decision-making authority was transferred to the system faster than the accountability architecture was built around it. Where regulation was mature (FCA-supervised robo-advisors, SRA-regulated AI law firms), automation absorbed into it. Where regulation was absent or weak (benefits administration, exam regulation during an emergency, private sub-postmaster contracts), automation displaced it.
Documented incidents
Evidence timeline
Discussed in Theory
The Black Box Society. Consequential decisions in finance, search, and reputation are increasingly made by systems simultaneously opaque to subjects, protected by trade-secrecy law, and treated as neutral. The combination produces a "one-way mirror" - institutions see citizens in full resolution; citizens cannot see how they are being scored.
Pasquale, Harvard University Press →Weapons of Math Destruction. Three properties make an algorithmic system a "WMD": opacity, scale, damage. Systems with all three reliably produce feedback loops in which existing inequality is encoded as signal and re-projected as prediction.
O'Neil, Crown Publishing →Automating Inequality. Ethnographic study of three US systems (Indiana automated welfare, LA coordinated-entry for homeless, Allegheny County child-welfare risk model). Algorithmic systems in public services disproportionately deployed on the poor - populations with the least political capacity to contest them.
Eubanks, St. Martin's Press →Demonstrated in Lab
Amazon built an experimental recruiting tool trained on ten years of applications predominantly from men. The model learned to penalise CVs containing "women's" (as in "women's chess club captain") and to downgrade graduates of two all-women colleges. Amazon scrapped the system in 2017. A model ostensibly predicting "good candidate" was actually predicting "candidate who resembles past hires."
Amazon hiring algorithm incident →Gender Shades. Audited three commercial facial-analysis systems (IBM, Microsoft, Face++). Error rates for classifying gender on darker-skinned women reached 34.7%, versus 0.8% for lighter-skinned men - a more than 40-fold disparity. IBM exited general-purpose facial recognition in 2020.
Buolamwini & Gebru, PMLR →COMPAS, used in US pre-trial and sentencing decisions, falsely flagged Black defendants as future criminals at roughly twice the rate of white defendants. Northpointe responded that COMPAS satisfied a different fairness criterion (calibration across groups). The debate itself became evidence: plausible fairness definitions are mutually incompatible, and choosing between them is a political decision, not a technical one.
ProPublica, "Machine Bias" (COMPAS analysis) →Demonstrated in Real World
The Horizon accounting system produced apparent shortfalls in sub-postmaster accounts. Over 900 sub-postmasters wrongfully convicted of theft, fraud, or false accounting between 1999 and 2015, with approximately 700 prosecutions pursued by the Post Office itself. Fujitsu's European head admitted at the 2024 inquiry that Fujitsu had known about Horizon bugs from 1999 and failed to disclose this in prosecutions. 77 CCRC referrals to appeal courts, 69 overturned. At least 13 suicides linked.
Post Office Horizon scandal (UK) →DWP internal "fairness analysis" obtained by Big Brother Watch via Computer Weekly 2024 found "statistically significant referral and outcome disparity" across all protected characteristics analysed: age, disability, marital status, nationality. A separate DWP/local-authority housing-benefit fraud-detection algorithm has been reported to have wrongly flagged approximately 200,000 people for investigation.
DWP fraud-detection algorithm →The Dutch tax administration used an algorithm that treated "dual nationality" and "foreign-sounding names" as fraud-risk indicators. Approximately 26,000 families (later estimates up to 35,000) were wrongly accused of childcare-benefit fraud. More than 2,000 children were removed from affected families into state care. Rutte III cabinet resigned January 2021. Dutch government formally acknowledged institutional racism within the tax administration in May 2022.
Dutch toeslagenaffaire →Strongest Counterargument
Human-only decision-making is also systematically biased, opaque in a different way (the inside of a decision-maker's head), and inconsistent. Unlike human decision-makers, algorithmic systems apply the same rules to each case and can in principle be audited across millions of decisions simultaneously. Kleinberg, Lakkaraju, Leskovec, Ludwig and Mullainathan (2018, Quarterly Journal of Economics) argued in the bail context that a well-designed algorithm could reduce both crime rates and incarceration while improving outcomes across racial groups. Regulated AI (FCA-supervised robo-advisors, SRA-regulated AI law firms) already operates inside mature accountability regimes.
Source: Kleinberg et al. (2018, Quarterly Journal of Economics). FCA supervisory framework for robo-advisors. SRA authorisation of Garfield.Law Ltd (2025).
Why this deserves weight: The counterarguments do not dissolve the capability concern; they refine it. The question is not 'algorithm or human?' but 'under what accountability regime?' The failures documented in Horizon, DWP, toeslagenaffaire, and A-level 2020 share a common feature: the decision-making authority was transferred to the system faster than the accountability architecture was built around it. Where regulation was mature and pre-existing (FCA), automation absorbed into it. Where regulation was absent or weak (benefits administration, exam regulation during an emergency, private sub-postmaster contracts), automation displaced it. The policy implication is not 'ban algorithmic decision-making' but 'no deployment without a prior, tested accountability regime.'