What generative AI means in a policing context
Most of the artificial intelligence covered elsewhere on this site does one narrow thing very quickly: match a face against a watchlist, read a number plate, score a risk. Generative AI is different in kind. It produces new text, and it produces it fluently, which turns out to matter a great deal when the output is going into a police record.
The systems now reaching policing are, for the most part, not policing products at all. Microsoft 365 Copilot is a general-purpose assistant built into standard office software, sold to organisations of every type. Its arrival in English policing came through the Police Digital Service's National Police Capabilities Environment, an Azure-based shared platform launched in 2025 to give forces common access to cloud tools. Copilot came with it, in the way that a spellchecker comes with a word processor, rather than as a considered policing procurement in its own right.
That distinction matters more than it might sound. A facial recognition system bought for policing arrives with accuracy testing, a threshold setting, an oversight framework and a body of case law. A general office assistant arrives with none of those, because nothing about its design anticipated that its output might end up in an intelligence assessment or a case file.
Report writing, the first mainstream use
The administrative burden in policing is real and well documented. Officers spend a substantial proportion of their time writing: incident reports, statements, handover notes, case file summaries. Any technology that credibly reduces that burden has an obvious appeal, and it is the reason generative AI has spread through forces faster than more visible technologies with far smaller footprints.
A Sky News survey found that at least 21 English police forces continued using Copilot after West Midlands Police withdrew it. Typical described uses are summarising long documents, drafting routine correspondence, and condensing case material. Separately, some US vendors sell tools built specifically for policing that draft an incident narrative directly from body-worn camera audio, which is a materially more consequential application than summarising an email thread.
The efficiency argument for all of this is straightforward and largely uncontested. The dispute is about what happens at the edges: what the system does when it does not know something, and whether anyone downstream can tell.
The West Midlands incident
In November 2025, West Midlands Police were preparing an intelligence assessment ahead of a Europa League fixture between Aston Villa and Maccabi Tel Aviv. Copilot was used during that preparation and produced a reference to a match involving disorder that had never taken place. The fabricated detail fed into the deliberations of a local Safety Advisory Group, which decided to bar away fans from attending.
The subsequent handling is as instructive as the error. Chief Constable Craig Guildford initially told the House of Commons Home Affairs Committee that no artificial intelligence had been used in producing the assessment, and later corrected that account in a written admission to the committee. He retired in January 2026. The force withdrew Copilot from operational use in March 2026 pending a review.
What makes this case genuinely significant is not that a language model produced something false, which is a well-understood property of the technology. It is that the false detail travelled: from a drafting tool, into an intelligence assessment, into a multi-agency decision-making forum, and out into a decision that affected several thousand people, without anyone in that chain identifying it as unverified. Full sourcing for this account is on the deployment tracker, and the system itself is covered on the Microsoft 365 Copilot technology page.
Why hallucination is a different kind of risk here
Large language models generate text by predicting plausible continuations. They do not have a mechanism for distinguishing between a fact they have encoded and a fluent invention, which is why fabricated output tends to arrive with exactly the same confident register as accurate output. In most settings that produces embarrassment. In policing it can produce a wrongful decision.
The specific difficulty is that policing documents are trusted downstream precisely because of where they come from. An intelligence assessment carries weight in a Safety Advisory Group, a case summary carries weight with a prosecutor, a report carries weight in court. That trust is institutional rather than textual: nobody reads a force intelligence document expecting to have to fact-check its internal references. Generative AI inserts a source of confident, unattributed error into exactly the part of the process that is designed not to be questioned.
This is a distinct problem from the accuracy debates around facial recognition or risk scoring, where the failure mode is a measurable false-positive rate that can at least be tested, published and argued about. There is no equivalent published error rate for a general-purpose assistant used to draft an intelligence report, because the task has no defined ground truth to measure against.
Evidential and disclosure problems
Criminal proceedings run on disclosure. The defence is entitled to understand how material against a defendant was produced, and the prosecution carries continuing obligations to disclose material that might undermine its case or assist the defence. Generative AI complicates that in ways the law has not yet fully worked through.
If an officer's statement was drafted with AI assistance from body-worn footage, several questions follow that currently have no settled answer. Is the fact of AI assistance itself disclosable? Does the original audio become the primary record with the statement as a derived document? If the model summarised or paraphrased, what happened to the material it left out, and would anyone know it had been left out? A model does not flag its own omissions.
There is also the question of what a defendant can meaningfully challenge. Where a risk-assessment score is disputed, there is at least a model, a set of inputs and a methodology to interrogate, however opaque. Where a narrative was drafted by a general-purpose assistant that no longer exists in the same state, that has been updated repeatedly since, and whose output is not reproducible, the thing being challenged is much harder to pin down.
Where the rules currently sit
There is no dedicated statutory framework for police use of generative AI in the UK. The governing rules are the general ones: data protection law where personal data is processed, human rights law where a decision engages a protected right, disclosure obligations in criminal proceedings, and force-level policy.
Force-level policy appears to vary considerably. Following the West Midlands incident, only eight of the forces surveyed confirmed that Copilot could not be used in investigations, which implies that the others either permit some investigative use or have not drawn the line explicitly. That variation is itself notable: a technology deployed through a shared national platform is being governed by forty-three separate local judgements about what it may be used for.
The Home Office's Algorithmic Transparency Recording Standard offers a voluntary route for public bodies to publish structured records of algorithmic tools in use, which would in principle cover this. Uptake among police forces has so far been limited, which means the clearest public picture of how generative AI is being used in policing has come from journalism and parliamentary questioning rather than from routine disclosure.
Follow the coverage
PoliceAI News tracks generative AI in policing as it develops: new force deployments, procurement decisions, parliamentary scrutiny, court rulings touching AI-assisted evidence, and vendor product launches. The feed refreshes every 30 minutes.
View Generative AI Stories