The AI Mistake That Actually Worries Me Isn't the Dramatic One
Every AI-in-finance governance conversation eventually gets to the dramatic hypothetical: what if an agent puts a hallucinated number in a board deck, or drafts external guidance language that’s wrong, and nobody catches it before it goes out. It’s a fair thing to worry about, and it’s the scenario that gets all the attention because it’s easy to picture and genuinely bad if it happens.
I actually think it’s not the failure mode most likely to hurt you. The dramatic scenario has enough obvious stakes that most organizations build a real review step around it—board materials get eyes on them regardless of how they were drafted. The failure mode I actually worry about is quieter, more frequent, and specifically designed by circumstance to slip past review.
The quiet version
A small, plausible, internally-consistent error, in a document nobody’s scrutinizing that closely because it’s “just” an internal working file, that gets copied forward into the next version, and the next, until three months later it’s load-bearing in a way nobody intended.
I’ve seen a version of this that didn’t even involve a dramatic hallucination—it was a rounding convention. A tool summarizing a set of regional figures rounded consistently in a way that was individually immaterial in any single output, but the rounded, slightly-off number got pasted into a working model, which fed a forecast, which got compared against actuals a quarter later by someone who assumed the model’s baseline was the exact original figure. Nothing about any single step looked wrong. The error was small enough at every stage to not trigger anyone’s “that doesn’t look right” instinct, which is exactly the instinct that used to catch mistakes before.
Why this specific failure mode is more dangerous, not less
A dramatic, obviously-wrong output tends to get caught, because it’s obviously wrong—someone’s instinct fires immediately. A small, plausible, boring error is precisely the kind that friction-based human review has always been worst at catching, AI or not. The difference now is volume and speed: a human making a small transcription error does it occasionally, in one document, that one person might remember making. A tool making a small, systematic error does it consistently, silently, across every document it touches, until someone happens to compare two versions closely enough to notice a pattern instead of a one-off.
What actually helps here, and it’s not “review everything more carefully”
Telling reviewers to “just be more careful” doesn’t work against an error type specifically built to not look wrong. What’s actually helped in practice:
Periodic reconciliation against an independent source, on a schedule, not just when something feels off. If a tool is producing recurring outputs, spot-check a sample against the original source data on a fixed cadence, specifically looking for small, consistent drift, not just obvious errors.
Treating “this has been running fine for months” as a reason to check harder, not a reason to relax. The failure mode described here gets more dangerous the longer it goes unnoticed, because more downstream work builds on the quietly-wrong number. Comfort with a tool’s track record should trigger a deeper check on the assumptions underneath it, not a lighter one.
Naming an owner for drift, specifically, separate from the owner for obvious errors. Most control frameworks are built to catch things that look wrong. This failure mode requires someone whose job is explicitly to ask “has anything subtly changed” on things that still look fine.
The dramatic AI failure stories make for a better conference talk. The quiet, plausible, small-and-consistent error is the one I’d actually spend governance budget defending against, because it’s the one built specifically to survive the kind of review most organizations already have in place.
Get monthly Finance × AI notes
One concise monthly email with practical finance AI strategy notes, field-tested patterns, and new project updates.
By subscribing, you agree to receive email updates. Unsubscribe at any time.
~Pedro Alizo