Aug 4 / Aneta Klosek

LLMs in Case Narrative Generation

AITHEA Insights

LLMs in Case Narrative Generation: Where They Help and Where They Hallucinate

Large language models can summarise a complex AML or KYC case in seconds. They can also confabulate transactions that never happened. Both things are true at once.

General-purpose LLMs show serious hallucination behavior on finance-specific tasks (Kang & Liu, 2023)
SAR drafting is described as a high-cost, low-scalability bottleneck in AML workflows (Co-Investigator AI, 2025)
2026 FinCEN's proposed rule offers a general endorsement of responsible AI experimentation in AML/CFT programs

Case narrative writing is one of the least loved parts of financial crime compliance. An analyst pulls together account history, transaction patterns, KYC records, and prior alerts, then turns all of it into a coherent written narrative — often under a filing deadline. It's exactly the kind of task large language models are good at on the surface: dense source material in, fluent prose out.

It's also exactly the kind of task where LLMs are prone to a specific, well-documented failure mode: generating text that reads as coherent and confident, while containing details that are not actually supported by the source material. Researchers refer to this as an extrinsic hallucination — plausible-sounding content with no grounding in the input — and note it appears specifically in narrative-generation and open-ended tasks, not just in factual question-answering.

Where LLMs genuinely help

Used as a drafting layer on top of structured case data, LLMs are strong at a few specific tasks:

  • Compressing volume. Turning dozens of transaction records and account notes into a first-pass narrative draft is a language task LLMs handle well, provided the underlying facts are supplied to the model rather than recalled from it.
  • Consistency of structure. A model can reliably follow a fixed narrative template — background, activity summary, red flags, conclusion — across hundreds of cases, which is harder to enforce with manual drafting at volume.
  • First-draft speed. A 2025 paper on agentic SAR-generation frameworks describes narrative drafting as a genuine bottleneck in AML workflows precisely because it's high-cost and low-scalability when done manually — which is the gap LLM-assisted drafting is aimed at.

Where they hallucinate

The same research makes an equally direct point: general-purpose LLMs used on their own for SAR narratives carry factual hallucination, limited alignment to known typologies, and poor explainability — described as posing unacceptable risk in compliance-critical domains. A separate empirical study testing LLMs on finance-specific tasks — explaining financial terminology and retrieving historical stock data — found off-the-shelf models produced factually incorrect output at a rate the authors called a serious deficiency, not an edge case.

In a case-narrative context, that failure mode tends to show up in specific, identifiable ways:

  • Inventing a transaction, date, or dollar amount that resembles the case but doesn't appear in the source file.
  • Asserting a typology conclusion — "consistent with structuring" — without a traceable link back to the data that supports it.
  • Smoothing over gaps in the record with plausible-sounding filler rather than flagging the gap.
Same task, two very different outputs GROUNDED DRAFTING Input: structured case file → transactions, dates, accounts → supplied directly to the model Output: narrative draft where every claim traces back to a specific record in the file UNGROUNDED GENERATION Input: brief description → "write a SAR narrative for a structuring case" Output: fluent narrative with invented dates, amounts, or an unsupported typology claim
The difference isn't the model — it's whether the output is tied to verifiable source records or generated from a prompt alone.

What the regulatory signal actually says

It's worth being precise here, since this is easy to overstate in either direction. FinCEN's proposed rule, published in April 2026, includes a general endorsement of institutions that responsibly experiment with innovative technologies in their AML/CFT programs, stating this alone won't trigger additional supervisory or enforcement risk. That is an openness to AI adoption — it is not a relaxation of the underlying requirement that BSA/AML programs maintain independent testing and documented internal controls, which remains a core pillar of the FFIEC examination framework. In practice, that means an LLM-drafted narrative still needs to move through the same review, documentation, and sign-off process any analyst-written narrative would.

The practical distinction

TaskWhere it sits
Drafting narrative prose from a supplied, structured case fileHelps — a legitimate first-draft accelerator
Standardizing structure and tone across many narrativesHelps — consistency at volume
Generating a narrative from a short prompt with no source data attachedHallucination risk — nothing to ground the output
Letting the model assert a final typology conclusion unreviewedHallucination risk — needs analyst sign-off

The rule of thumb compliance teams keep coming back to: an LLM should draft from your data, not instead of it. If a claim in the narrative can't be traced to a specific record in the case file, that's not a drafting shortcut — it's an unreviewed fact waiting to become a filing error.

Want to go deeper on where AI genuinely fits into a compliance program — and where the guardrails need to hold?

Track 1: AI Foundations in Financial Crime Compliance walks through exactly this: grounding, review controls, and the regulatory context, built for AML and KYC teams evaluating these tools now.

Explore Track 1 →
Sources: Naik et al., "Co-Investigator AI: The Rise of Agentic AI for Smarter, Trustworthy AML Compliance Narratives," arXiv:2509.08380, 2025; Kang & Liu, "Deficiency of Large Language Models in Finance: An Empirical Examination of Hallucination," arXiv:2311.15548, 2023; Frontiers in Artificial Intelligence, "Survey and analysis of hallucinations in large language models," 2025; Foley Hoag, "Modernizing BSA/AML Compliance: FinCEN's Proposed Program Rule," June 2026; FFIEC BSA/AML Examination Manual (bsaaml.ffiec.gov). This article is for informational purposes and is not legal or compliance advice.
Disclaimer & Copyright

This article is published by AITHEA GmbH for informational and educational purposes only. It is intended to provide general insights into compliance, technology, and related topics, and does not constitute legal, regulatory, or professional advice. Readers should seek appropriate professional guidance before making decisions based on the content provided.

The views and opinions expressed in this article are those of the author(s) and do not necessarily reflect the official position of AITHEA GmbH, its partners, or affiliated organisations.

Artificial intelligence (AI) tools are actively used in the creation of this content. AI may support the drafting of text and the improvement of structure, clarity, and readability. However, all content is based on human-defined topics, reviewed against relevant sources, and critically validated by the author(s). AI is not used as a source of truth, but as a supporting tool to enhance communication.

All content is the intellectual property of AITHEA GmbH unless otherwise stated. Reproduction, distribution, or translation for non-commercial purposes is permitted, provided that appropriate credit is given and the source is clearly acknowledged. For any commercial use, prior written permission is required. © Aithea, 2026.
Created with