Bristol's program started with a genuinely reasonable idea. Social workers, schools, and police each held fragments of information about at-risk children. A child missing school regularly, while also living in a home where domestic abuse had been reported: each piece alone might not trigger action, but together they should. Starting in 2015, a joint team began combining that data into a single database, the Think Family Database, which eventually held records on close to half a million people.
The problems started when the team moved from sharing information to predicting behavior. They built at least 23 separate models, scoring people for the likelihood they would commit burglary, fail to appear in court, go missing, or become victims of exploitation. One model in the Offender Management App was designed to hold data on around 300,000 people in the region. Children were scored for risk of sexual or criminal exploitation using a mix of police intelligence, housing data, welfare records, and school meal eligibility.
Cardiff University researchers flagged the problem early: many of the variables used were simply proxies for poverty. If you score higher because your family receives housing support or your child gets free school meals, the algorithm is not detecting risk. It is tracking deprivation. The children with the highest scores were mostly children already well known to social workers. As the original program head Gary Davies acknowledged, "most of the output told you what you already knew."
Then the models got worse. Police decided to expand the system across five councils rather than just Bristol, but data-sharing agreements with those councils fell through. So the models were retrained using only police data, which meant no welfare records, no housing context, no school data. Staff quickly noticed the difference. Girls who were known victims of criminal exploitation stopped appearing in results. One staffer described spending hours going through the algorithm's outputs and gradually stopping because the effort produced nothing reliable. Another said they would not cite the system in a meeting because they were not confident it was accurate enough to use.
When an independent auditor tried to test the models, the source code could not be found. No one had kept records of how they worked or why they were eventually shut down.
An independent AI audit firm reviewed over 36,000 model performance scores released by the police. The verdict: most models produced what the auditors called "genuinely poor predictive performance." The burglary prediction model appeared to work at below 10 percent precision for more than three years. That means for every ten people it flagged as likely burglars, nine or more were flagged incorrectly. The police say that particular model was never deployed. But they also said the performance records were generated automatically and never deleted after that decision, which makes it difficult to verify what was actually in use.
The bias testing was equally thin. The police showed regulators a screenshot of an app that compared average risk scores for white people and people of color and concluded there was no significant difference. The auditors pointed out that comparing averages says almost nothing about whether the system produces discriminatory outcomes at the individual level.
This matters far beyond Bristol because the UK government just launched PoliceAI, backed by £75 million to identify, test, and spread AI tools across all 43 police forces. The center is hosted by the College of Policing, led by Andy Marsh: the same former chief constable who oversaw Avon and Somerset's program. Trials begin in 2026, with a national rollout planned for 2027. The government says AI models will be independently tested for accuracy and bias, and a public registry of tools in use will be published by autumn this year.
Those are the right commitments. The question is whether they hold. Bristol's ethics committee discussed predictive analytics once, in 2016, and apparently never met on the subject again. Staff described models built and run by single individuals, with no documentation, no public disclosure, and no mechanism for affected people to find out they had been scored.
John Pegram, a community accountability activist in Bristol, only found out he was in the Offender Management App in 2024, years after it was created, and only after hiring solicitors. He still does not know what data is held about him or how it might affect his next encounter with police.
The EU's AI Act, in force since February 2025, now prohibits using AI to predict the probability that a named individual will commit a crime. The UK, outside the EU, has no equivalent rule. PoliceAI's public registry and independent testing are voluntary commitments, not legal requirements. That gap is where the next Bristol gets built.