Skip to main content

Governance

AI in Criminal Evidence Review: The Safeguards That Must Come First

How AI used to triage or interpret evidence changes the control set — EU AI Act classification, UK disclosure and expert-evidence duties, and the reproducibility record a court will expect.

AI in Criminal Evidence Review: The Safeguards That Must Come First

A compliance team is asked to sign off a large language model that will triage several hundred thousand messages from a whistleblowing investigation. The tool ranks material by relevance so that a small legal team can review the top slice. Nobody in the room calls this evidence. But if the investigation is referred to a prosecutor, or a regulator passes the file on, the material that the model pushed down the ranking becomes as legally interesting as the material it pushed up — and the organisation will be asked how the ranking was produced, whether it can be reproduced, and what was in the part nobody read.

That is the practical shape of AI entering criminal evidence review. It rarely arrives as a decision to automate justice. It arrives as a document-triage tool, a transcription pipeline, a face or voice matching feature bundled into a security platform, or an anomaly detector in a fraud team. The governance question is not whether an algorithm can understand human intent. It is narrower and answerable: can you demonstrate, months later and to a hostile reader, what the system did, on what input, under what version, and what a human independently checked.

Two routes by which this lands on your desk

The first route is as a deployer. Your own investigations, security operations, or legal team uses AI on material that could later become evidence in criminal or regulatory proceedings. Most mid-size organisations are here without having decided to be.

The second route is as a provider. Your product is sold to, or used by, police forces, prosecutors, courts or their contractors. Once that is true, the obligations attach to the system you built, regardless of how small a share of revenue the public-sector contract represents.

What the EU AI Act actually says about this area

The Act treats law enforcement and the administration of justice as areas of concentrated risk rather than as ordinary software markets. Annex III lists as high-risk, among others, systems used to evaluate the reliability of evidence during investigation or prosecution, systems assessing the risk of a person offending or re-offending, profiling during detection and investigation, and systems assisting a judicial authority in researching and interpreting facts and law. Article 5 goes further and prohibits outright the use of AI to assess or predict the risk that a person will commit a criminal offence where that assessment rests solely on profiling or on assessing personality traits and characteristics — with a carve-out only where the system supports a human assessment already grounded in objective, verifiable facts directly linked to criminal activity.

ScenarioLikely positionPractical consequence
Internal investigation triage, no law-enforcement involvementNot high-risk on Annex III grounds alone; data protection law still applies in fullDPIA, lawful basis for criminal offence data, retention and reproducibility records
You supply a tool that weighs evidential reliability to a prosecutor or police forceHigh-risk provider under Annex IIIRisk management, data governance, technical documentation, logging, human-oversight instructions, conformity assessment
Your public-sector customer uses that toolHigh-risk deployerUse per instructions, assigned competent oversight, input-data relevance, log retention, fundamental rights impact assessment for public bodies
Predicting offending from profiling or personality traits aloneProhibited practiceMust not be placed on the market or used; the Act's highest exposure, up to 7% of global turnover

Other high-risk breaches sit at up to 3% of global turnover. Where personal data is mishandled alongside this, UK and EU GDPR exposure runs to 4%. The Act's Article 4 duty to ensure a sufficient level of AI literacy among staff dealing with these systems applies broadly, not only to high-risk deployments.

The UK position: no single statute, several sharp edges

The UK has no equivalent cross-cutting AI statute, which is often misread as an absence of obligation. The sharp edges are elsewhere. Part 3 of the Data Protection Act 2018 governs processing by competent authorities for law enforcement purposes and constrains decisions based solely on automated processing. UK GDPR restricts solely automated decisions producing legal or similarly significant effects, and criminal offence data carries its own conditions for processing. In England and Wales, the Forensic Science Regulator's statutory Code of Practice sets validation and method-quality expectations for forensic activities including digital forensics. Expert evidence is governed by the Criminal Procedure Rules, and the reliability factors courts apply — whether the method has been properly tested, whether its error rate is known, whether the reasoning can be explained — are precisely the questions most AI vendors are least prepared to answer.

The disclosure regime is the one most often overlooked. Prosecutors must disclose unused material capable of undermining the prosecution case or assisting the defence. A triage model that suppressed material is not a neutral piece of infrastructure in that analysis: how it was configured, what it excluded, and how reliably it performs can themselves become matters the defence is entitled to probe.

Safeguards to have in place before an AI system touches evidential material

  • Original material is never mutated. AI outputs are derived work-product held separately from the source, with hashes taken on ingest and verified on export.
  • A reproducibility record per run: model identifier and version, hosting endpoint, sampling parameters, prompt or query template version, the snapshot of any retrieval index, operator identity, and timestamp. Without the version, the run cannot be repeated.
  • A change log for model updates. A silently upgraded hosted model can change outputs on identical input; treat provider version changes as a controlled change, not a maintenance event.
  • Documented recall behaviour, not just precision. What the tool surfaces is easy to inspect; what it buries is the risk. Sample the discard pile at a defined rate and record the result.
  • Human oversight that is real. Named reviewers, defined competence, explicit authority to disagree with and discard the output, and a record when they do. The Act requires oversight designed against automation bias — a reviewer who cannot see the basis for a ranking cannot exercise it.
  • Subgroup error testing where the system touches people — language, dialect, accent, image quality and demographic variation for biometric or voice tools. Where a vendor cannot supply this, record that gap as an accepted risk with a named owner rather than assuming it away.
  • Retention and legal hold covering logs, prompts and intermediate outputs, not only final reports.
  • A prohibited-use statement naming what the tool must never be used for — predictive risk scoring of individuals, credibility assessment, or anything presented to a decision-maker as a conclusion rather than a pointer.

Limits of this guidance

This is a governance and control framing, not legal advice, and it does not tell you whether a specific output is admissible in a specific case. Admissibility, disclosure obligations and expert-evidence duties are jurisdiction-specific and fact-specific; Scotland and Northern Ireland differ from England and Wales, and EU member states differ from one another in criminal procedure even where the AI Act is uniform. Nothing here addresses the position of competent authorities themselves, whose statutory powers and oversight arrangements sit outside a corporate compliance programme.

Take specialist criminal and data protection advice before AI output is relied on in proceedings, before deploying any biometric identification or emotion-inference capability in an investigative context, and before supplying a system into a justice or law-enforcement setting for the first time. Where a use case sits near the prohibited category, treat it as a stop rather than a risk to be managed, and get that assessed independently.

Watch this as a video

  • EU AI Act
  • Criminal Evidence
  • High-Risk AI
  • Human Oversight
  • Digital Forensics
  • Disclosure

More guides

Start Free AI Compliance Review