Every investigation, audit, and legal case begins the same way: with a heap of records nobody has read yet. Emails, scanned contracts, call logs, transaction ledgers, camera footage  all of it real, some of it accurate, none of it doing anything on its own. The job of analysis has always been the slow conversion of that heap into something a person can stand behind in a courtroom, a boardroom, or a published story.

What has changed is that machines can now perform a large share of that conversion, at a scale and speed no team of humans could match. That shift is worth examining carefully rather than celebrating loudly, because it helps in some places and creates fresh problems in others.

Hold on to one distinction from the start. A record is raw and inert. Evidence is a claim that has survived scrutiny and can be defended when challenged. The path from one to the other runs through a middle stage of organised, searchable information  and it is at each of these transitions that AI now does its work. The rest of this piece follows that path, from the pile of documents to the point where someone is willing to sign their name to a conclusion.

What “raw records” really are

The instinct is to describe the problem as one of size, and size is part of it. But volume alone is a solvable problem; you can always add storage. The genuinely hard property of raw records is that they are wildly inconsistent with one another.

A single case might contain typed reports, handwritten notes, PDFs that are really photographs of paper, audio recordings, database exports, and message threads that reference events without ever describing them. Dates appear in five formats. The same person shows up as “J. Alvarez,” “Jose A.,” and an email handle with no name attached at all. Half the useful context lives in what the records leave unsaid.

This is why analysts have long said that the intelligent part of the job is the smallest part. Historically, the overwhelming majority of effort went into reading, cleaning, and reconciling material before anyone could ask a meaningful question of it. Heterogeneity, not the sheer count of files, is the wall that analysis keeps running into.

The old way, and where it stalls

Before AI tooling matured, three methods carried most of the load: keyword search over whatever had been digitised, line-by-line manual review, and sampling  looking closely at a representative slice and inferring the rest.

Each works, and each has a hard ceiling. Keyword search only finds what you already knew to look for; it is blind to the document that matters but uses none of your search terms. Manual review is thorough but does not scale, and a reviewer on their four-hundredth file is not the reviewer they were on their fourth. Sampling is statistically respectable but structurally unable to catch the rare item  and in investigations, the rare item is frequently the whole point.

The real cost of the traditional approach was never just its slowness. It was the evidence that went undiscovered because no one ever had the hours to reach it.

The AI toolkit, sorted by what it does

useful-evidence-1jpg_1790160582.webp

It is more useful to group these tools by function than by product name, because the products change every year while the functions stay stable. Across a working pipeline, AI tends to operate at six distinct points, each handing its output to the next.

FunctionWhat it does to the dataTypical use in practice
Ingestion & extractionTurns unreadable formats into machine-readable textOCR on scanned files; speech-to-text on interviews and calls
StructuringTags entities and resolves duplicates into consistent recordsRecognising people, places and dates; linking “J. Alvarez” to “Jose A.”
RetrievalFinds material by meaning rather than exact wordingSemantic search that surfaces relevant files sharing no keywords
Pattern discoveryFlags what is unusual or connected across the whole setAnomaly detection in ledgers; mapping who contacted whom
ReconstructionAssembles scattered fragments into an ordered accountBuilding a timeline from out-of-sequence sources
SynthesisDrafts a readable narrative from the structured findingsSummarising a cluster of documents into a first-pass memo

Two of these deserve a second look, because they are where the recent gains are largest. Semantic retrieval is the quiet breakthrough: a system that understands that “the payment was reversed” and “funds returned to the account” describe the same event will find records that keyword search walks straight past. And structuring  the dull work of making a million inconsistent records to agree on who and when  is precisely the labour that used to consume analysts’ weeks.

The same pipeline, across four fields

The abstract stages become concrete when you watch them run in different settings. The underlying moves are identical; only the stakes and the vocabulary shift.

1. Legal discovery : review platforms use predictive coding: a lawyer tags a few hundred documents as relevant or not, and the system extends that judgment across millions, ranking the rest by likely relevance.  

2. Investigative journalism : the Panama Papers leak of 11.5 million documents was navigable only because the files were first extracted, indexed, and linked into a searchable graph of names and companies  a task no newsroom could have finished by hand in a lifetime.

3. In fraud and audit, anomaly detection reads entire transaction histories rather than a sample, surfacing the handful of entries that break the pattern of the rest.

4. In forensics and intelligence, the work is increasingly multimodal  reconciling an audio clip, an image’s embedded location data, and a call record into a single account of what happened and when.

When does analysis become evidence?

This is the hinge of the whole subject. A model can hand you a striking correlation, a cluster, or a confident summary. None of that is yet evidence, because evidence is not defined by how it was produced but by whether it can be justified. Three questions decide the matter.

• Can it be traced back? A finding you cannot follow from conclusion to source record is an assertion, not evidence. Provenance of an unbroken line from the claim to the original material  is what lets anyone else check your work, and it is the first thing a serious challenge will attack.

• Would it happen again? Reproducibility matters. If running the same material through the same process yields a different answer on Tuesday than it did on Monday, the output cannot carry weight. Many AI systems are non-deterministic by default, which makes this harder to guarantee than it sounds.

• Is it relevant and reliable? An accurate finding about the wrong question proves nothing, and a finding from an unreliable method proves less than it appears to. Both tests are old; AI has not changed them, only changed how much material now reaches them.

The honest summary is that AI is excellent at generating candidate evidence and incapable, on its own, of conferring the legitimacy that turns a candidate into something usable. That step still belongs to people and institutions.

useful-evidence-2jpg_1790160592.webp

The failure modes worth naming

Vague warnings about AI being “risky” are not much use to anyone doing the work. The specific failure modes are more instructive, because each calls for a different guard. The danger is rarely that these systems are wrong; it is that they are wrong persuasively, fluent and confident regardless of accuracy.

• Fabrication. Generative systems can invent citations, quotations, and facts that exist nowhere in the source material, and present them in the same tone as everything they got right. In an evidentiary setting this is not a minor quirk; it is disqualifying.

• Inherited bias. If the raw records encode a historical skew  who was policed, who was flagged, who was audited  a model tuned on them reproduces that skew and applies it at machine scale, lending old prejudice the appearance of neutral computation.

• Misplaced confidence. A clean, well-structured output signals reliability to a reader whether or not it has earned it. The polish is a property of the format, not the truth, and it steadily erodes the healthy scepticism a reviewer should bring.

• The opaque conclusion. A finding no one can explain is a finding no one can defend. If a system cannot show why it reached an answer, that answer struggles to survive cross-examination, audit, or peer review of the exact places evidence is meant to hold up.

The analyst’s job moves, it does not vanish

A common assumption is that better tools mean fewer people. What actually happens is that the human role relocates. When a machine reads everything, the analyst stops being the person who finds and becomes the person who decides what the findings mean.

That is not a demotion. Judgment about relevance, weight, and defensibility is harder and more consequential than retrieval ever was. And there is a quiet irony in it: as AI takes on more of the labour, the human’s accountability for the result grows rather than shrinks, because someone still has to be answerable for a conclusion the machine merely proposed.

Governance and admissibility

Whether an AI-assisted finding counts as usable depends heavily on where you are standing. A courtroom, a newsroom, and an audit committee apply different bars to the same output, and each is developing its own expectations for how such findings must be documented.

The common thread is that defensibility is built in advance, not argued after the fact. It rests on practical machinery: audit trails that record what was run and when, documentation of the models and settings used, validation testing that shows the method works on known material, and enough transparency that an independent party can retrace the path. Where these exist, an AI-derived conclusion can be defended. Where they are missing, even a correct conclusion may be unusable  and there is no way to add them retroactively once a challenge has begun.

What actually changes, and what is merely faster

Not every improvement is a transformation, and it is worth being precise about which is which. Some of what AI brings is the old work done more cheaply. A smaller part is genuinely new capability  analysis that simply was not possible before at any price.

Mostly just faster or cheaperGenuinely different in kind
Reviewing more documents per hourFinding evidence that manual review would never have reached at all
Transcribing and searching text at lower costRetrieving material by meaning when it shares no keywords with the query
Producing summaries and first drafts quicklyCross-linking thousands of sources no single person could hold in mind at once
Digitising archives that were previously on paperDetecting patterns visible only across the entire set, never in a sample

The left column is valuable and unglamorous. The right column is where the real shift lives  not in doing familiar work faster, but in making certain evidence discoverable for the first time.

Where this is heading

The near-term direction is reasonably clear, and each likely development is best understood as an answer to a current limitation. Agentic systems that plan and run multi-step investigations on their own address the bottleneck of an analyst manually chaining each tool to the next. Fuller multimodal fusion  treating text, audio, image, and location as one integrated body of evidence  responds to how badly today’s pipelines still handle anything that is not text.

Real-time analysis of streaming records points at the last gap: the fact that most analysis today happens well after the events it describes. None of these removes the core tension. Each of them widens the distance between what can be surfaced and what can be trusted, which means the verification burden only grows.

From volume to verdict

Return to the heap of records at the start. AI has, for practical purposes, collapsed the distance between that pile and organised, searchable information, the part of the work that used to swallow whole teams for months. That is a real and lasting change, and it is worth recognising plainly.

But the final step, the one that turns information into evidence someone will vouch for, has not been automated and shows little sign of being. If anything, it has become more demanding, because there is now far more material clearing the machine’s filters and arriving at a human’s desk for judgment.

The tools have moved the boundary of what is knowable. Deciding what is trustworthy  and being willing to answer for that decision  remains exactly where it always was. Which leaves the question the technology cannot settle: as the volume of what we can surface keeps climbing, are we building the discipline to keep pace with the judgment it demands?