Every investigation, audit, and legal case begins the same way: with a heap of records nobody has read yet. Emails, scanned contracts, call logs, transaction ledgers, camera footage all of it real, some of it accurate, none of it doing anything on its own. The job of analysis has always been the slow conversion of that heap into something a person can stand behind in a courtroom, a boardroom, or a published story.
What has changed is that machines can now perform a large share of that conversion, at a scale and speed no team of humans could match. That shift is worth examining carefully rather than celebrating loudly, because it helps in some places and creates fresh problems in others.
Hold on to one distinction from the start. A record is raw and inert. Evidence is a claim that has survived scrutiny and can be defended when challenged. The path from one to the other runs through a middle stage of organised, searchable information and it is at each of these transitions that AI now does its work. The rest of this piece follows that path, from the pile of documents to the point where someone is willing to sign their name to a conclusion.
What “raw records” really are
The instinct is to describe the problem as one of size, and size is part of it. But volume alone is a solvable problem; you can always add storage. The genuinely hard property of raw records is that they are wildly inconsistent with one another.
A single case might contain typed reports, handwritten notes, PDFs that are really photographs of paper, audio recordings, database exports, and message threads that reference events without ever describing them. Dates appear in five formats. The same person shows up as “J. Alvarez,” “Jose A.,” and an email handle with no name attached at all. Half the useful context lives in what the records leave unsaid.
This is why analysts have long said that the intelligent part of the job is the smallest part. Historically, the overwhelming majority of effort went into reading, cleaning, and reconciling material before anyone could ask a meaningful question of it. Heterogeneity, not the sheer count of files, is the wall that analysis keeps running into.
The old way, and where it stalls
Before AI tooling matured, three methods carried most of the load: keyword search over whatever had been digitised, line-by-line manual review, and sampling looking closely at a representative slice and inferring the rest.
Each works, and each has a hard ceiling. Keyword search only finds what you already knew to look for; it is blind to the document that matters but uses none of your search terms. Manual review is thorough but does not scale, and a reviewer on their four-hundredth file is not the reviewer they were on their fourth. Sampling is statistically respectable but structurally unable to catch the rare item and in investigations, the rare item is frequently the whole point.
The real cost of the traditional approach was never just its slowness. It was the evidence that went undiscovered because no one ever had the hours to reach it.
The AI toolkit, sorted by what it does

It is more useful to group these tools by function than by product name, because the products change every year while the functions stay stable. Across a working pipeline, AI tends to operate at six distinct points, each handing its output to the next.
| Function | What it does to the data | Typical use in practice |
| Ingestion & extraction | Turns unreadable formats into machine-readable text | OCR on scanned files; speech-to-text on interviews and calls |
| Structuring | Tags entities and resolves duplicates into consistent records | Recognising people, places and dates; linking “J. Alvarez” to “Jose A.” |
| Retrieval | Finds material by meaning rather than exact wording | Semantic search that surfaces relevant files sharing no keywords |
| Pattern discovery | Flags what is unusual or connected across the whole set | Anomaly detection in ledgers; mapping who contacted whom |
| Reconstruction | Assembles scattered fragments into an ordered account | Building a timeline from out-of-sequence sources |
| Synthesis | Drafts a readable narrative from the structured findings | Summarising a cluster of documents into a first-pass memo |
Two of these deserve a second look, because they are where the recent gains are largest. Semantic retrieval is the quiet breakthrough: a system that understands that “the payment was reversed” and “funds returned to the account” describe the same event will find records that keyword search walks straight past. And structuring the dull work of making a million inconsistent records to agree on who and when is precisely the labour that used to consume analysts’ weeks.
The same pipeline, across four fields
The abstract stages become concrete when you watch them run in different settings. The underlying moves are identical; only the stakes and the vocabulary shift.
1. Legal discovery : review platforms use predictive coding: a lawyer tags a few hundred documents as relevant or not, and the system extends that judgment across millions, ranking the rest by likely relevance.
2. Investigative journalism : the Panama Papers leak of 11.5 million documents was navigable only because the files were first extracted, indexed, and linked into a searchable graph of names and companies a task no newsroom could have finished by hand in a lifetime.
3. In fraud and audit, anomaly detection reads entire transaction histories rather than a sample, surfacing the handful of entries that break the pattern of the rest.
4. In forensics and intelligence, the work is increasingly multimodal reconciling an audio clip, an image’s embedded location data, and a call record into a single account of what happened and when.
When does analysis become evidence?
This is the hinge of the whole subject. A model can hand you a striking correlation, a cluster, or a confident summary. None of that is yet evidence, because evidence is not defined by how it was produced but by whether it can be justified. Three questions decide the matter.
• Can it be traced back? A finding you cannot follow from conclusion to source record is an assertion, not evidence. Provenance of an unbroken line from the claim to the original material is what lets anyone else check your work, and it is the first thing a serious challenge will attack.
• Would it happen again? Reproducibility matters. If running the same material through the same process yields a different answer on Tuesday than it did on Monday, the output cannot carry weight. Many AI systems are non-deterministic by default, which makes this harder to guarantee than it sounds.
• Is it relevant and reliable? An accurate finding about the wrong question proves nothing, and a finding from an unreliable method proves less than it appears to. Both tests are old; AI has not changed them, only changed how much material now reaches them.
The honest summary is that AI is excellent at generating candidate evidence and incapable, on its own, of conferring the legitimacy that turns a candidate into something usable. That step still belongs to people and institutions.

The failure modes worth naming
Vague warnings about AI being “risky” are not much use to anyone doing the work. The specific failure modes are more instructive, because each calls for a different guard. The danger is rarely that these systems are wrong; it is that they are wrong persuasively, fluent and confident regardless of accuracy.
• Fabrication. Generative systems can invent citations, quotations, and facts that exist nowhere in the source material, and present them in the same tone as everything they got right. In an evidentiary setting this is not a minor quirk; it is disqualifying.
• Inherited bias. If the raw records encode a historical skew who was policed, who was flagged, who was audited a model tuned on them reproduces that skew and applies it at machine scale, lending old prejudice the appearance of neutral computation.
• Misplaced confidence. A clean, well-structured output signals reliability to a reader whether or not it has earned it. The polish is a property of the format, not the truth, and it steadily erodes the healthy scepticism a reviewer should bring.
• The opaque conclusion. A finding no one can explain is a finding no one can defend. If a system cannot show why it reached an answer, that answer struggles to survive cross-examination, audit, or peer review of the exact places evidence is meant to hold up.
The analyst’s job moves, it does not vanish
A common assumption is that better tools mean fewer people. What actually happens is that the human role relocates. When a machine reads everything, the analyst stops being the person who finds and becomes the person who decides what the findings mean.
That is not a demotion. Judgment about relevance, weight, and defensibility is harder and more consequential than retrieval ever was. And there is a quiet irony in it: as AI takes on more of the labour, the human’s accountability for the result grows rather than shrinks, because someone still has to be answerable for a conclusion the machine merely proposed.
Governance and admissibility
Whether an AI-assisted finding counts as usable depends heavily on where you are standing. A courtroom, a newsroom, and an audit committee apply different bars to the same output, and each is developing its own expectations for how such findings must be documented.
The common thread is that defensibility is built in advance, not argued after the fact. It rests on practical machinery: audit trails that record what was run and when, documentation of the models and settings used, validation testing that shows the method works on known material, and enough transparency that an independent party can retrace the path. Where these exist, an AI-derived conclusion can be defended. Where they are missing, even a correct conclusion may be unusable and there is no way to add them retroactively once a challenge has begun.
What actually changes, and what is merely faster
Not every improvement is a transformation, and it is worth being precise about which is which. Some of what AI brings is the old work done more cheaply. A smaller part is genuinely new capability analysis that simply was not possible before at any price.
| Mostly just faster or cheaper | Genuinely different in kind |
| Reviewing more documents per hour | Finding evidence that manual review would never have reached at all |
| Transcribing and searching text at lower cost | Retrieving material by meaning when it shares no keywords with the query |
| Producing summaries and first drafts quickly | Cross-linking thousands of sources no single person could hold in mind at once |
| Digitising archives that were previously on paper | Detecting patterns visible only across the entire set, never in a sample |
The left column is valuable and unglamorous. The right column is where the real shift lives not in doing familiar work faster, but in making certain evidence discoverable for the first time.
Where this is heading
The near-term direction is reasonably clear, and each likely development is best understood as an answer to a current limitation. Agentic systems that plan and run multi-step investigations on their own address the bottleneck of an analyst manually chaining each tool to the next. Fuller multimodal fusion treating text, audio, image, and location as one integrated body of evidence responds to how badly today’s pipelines still handle anything that is not text.
Real-time analysis of streaming records points at the last gap: the fact that most analysis today happens well after the events it describes. None of these removes the core tension. Each of them widens the distance between what can be surfaced and what can be trusted, which means the verification burden only grows.
From volume to verdict
Return to the heap of records at the start. AI has, for practical purposes, collapsed the distance between that pile and organised, searchable information, the part of the work that used to swallow whole teams for months. That is a real and lasting change, and it is worth recognising plainly.
But the final step, the one that turns information into evidence someone will vouch for, has not been automated and shows little sign of being. If anything, it has become more demanding, because there is now far more material clearing the machine’s filters and arriving at a human’s desk for judgment.
The tools have moved the boundary of what is knowable. Deciding what is trustworthy and being willing to answer for that decision remains exactly where it always was. Which leaves the question the technology cannot settle: as the volume of what we can surface keeps climbing, are we building the discipline to keep pace with the judgment it demands?