LLMs in Case Narrative Generation
LLMs in Case Narrative Generation: Where They Help and Where They Hallucinate
Large language models can summarise a complex AML or KYC case in seconds. They can also confabulate transactions that never happened. Both things are true at once.
Case narrative writing is one of the least loved parts of financial crime compliance. An analyst pulls together account history, transaction patterns, KYC records, and prior alerts, then turns all of it into a coherent written narrative — often under a filing deadline. It's exactly the kind of task large language models are good at on the surface: dense source material in, fluent prose out.
It's also exactly the kind of task where LLMs are prone to a specific, well-documented failure mode: generating text that reads as coherent and confident, while containing details that are not actually supported by the source material. Researchers refer to this as an extrinsic hallucination — plausible-sounding content with no grounding in the input — and note it appears specifically in narrative-generation and open-ended tasks, not just in factual question-answering.
Where LLMs genuinely help
Used as a drafting layer on top of structured case data, LLMs are strong at a few specific tasks:
- Compressing volume. Turning dozens of transaction records and account notes into a first-pass narrative draft is a language task LLMs handle well, provided the underlying facts are supplied to the model rather than recalled from it.
- Consistency of structure. A model can reliably follow a fixed narrative template — background, activity summary, red flags, conclusion — across hundreds of cases, which is harder to enforce with manual drafting at volume.
- First-draft speed. A 2025 paper on agentic SAR-generation frameworks describes narrative drafting as a genuine bottleneck in AML workflows precisely because it's high-cost and low-scalability when done manually — which is the gap LLM-assisted drafting is aimed at.
Where they hallucinate
The same research makes an equally direct point: general-purpose LLMs used on their own for SAR narratives carry factual hallucination, limited alignment to known typologies, and poor explainability — described as posing unacceptable risk in compliance-critical domains. A separate empirical study testing LLMs on finance-specific tasks — explaining financial terminology and retrieving historical stock data — found off-the-shelf models produced factually incorrect output at a rate the authors called a serious deficiency, not an edge case.
In a case-narrative context, that failure mode tends to show up in specific, identifiable ways:
- Inventing a transaction, date, or dollar amount that resembles the case but doesn't appear in the source file.
- Asserting a typology conclusion — "consistent with structuring" — without a traceable link back to the data that supports it.
- Smoothing over gaps in the record with plausible-sounding filler rather than flagging the gap.
What the regulatory signal actually says
It's worth being precise here, since this is easy to overstate in either direction. FinCEN's proposed rule, published in April 2026, includes a general endorsement of institutions that responsibly experiment with innovative technologies in their AML/CFT programs, stating this alone won't trigger additional supervisory or enforcement risk. That is an openness to AI adoption — it is not a relaxation of the underlying requirement that BSA/AML programs maintain independent testing and documented internal controls, which remains a core pillar of the FFIEC examination framework. In practice, that means an LLM-drafted narrative still needs to move through the same review, documentation, and sign-off process any analyst-written narrative would.
The practical distinction
| Task | Where it sits |
|---|---|
| Drafting narrative prose from a supplied, structured case file | Helps — a legitimate first-draft accelerator |
| Standardizing structure and tone across many narratives | Helps — consistency at volume |
| Generating a narrative from a short prompt with no source data attached | Hallucination risk — nothing to ground the output |
| Letting the model assert a final typology conclusion unreviewed | Hallucination risk — needs analyst sign-off |
The rule of thumb compliance teams keep coming back to: an LLM should draft from your data, not instead of it. If a claim in the narrative can't be traced to a specific record in the case file, that's not a drafting shortcut — it's an unreviewed fact waiting to become a filing error.
Want to go deeper on where AI genuinely fits into a compliance program — and where the guardrails need to hold?
Track 1: AI Foundations in Financial Crime Compliance walks through exactly this: grounding, review controls, and the regulatory context, built for AML and KYC teams evaluating these tools now.
Explore Track 1 →
Registered at District Court Munich HRB 302338
VAT ID DE454846466 | nanoacademy@ai-thea.com