Your Model Risk Policy Was Written for Spreadsheets, Not Language Models
Somewhere in most finance organizations there’s a model risk inventory. It has entries for the DCF template, the bad-debt forecasting model, maybe an interest rate hedge model. Each entry has an owner, a validation date, a documented set of assumptions, and a tolerance for how wrong it’s allowed to be before someone gets a call.
Then last year someone added “ChatGPT Enterprise” to the same spreadsheet, in the same format, and called it done.
It isn’t done. The framework asks the wrong questions for this class of tool, and pretending otherwise is how you end up with a governance program that looks complete on paper and catches nothing in practice.
The old framework assumes a fixed function
A traditional financial model has a knowable structure: these are the inputs, this is the transformation, this is the output, and if you run it twice with the same inputs you get the same answer. Validation means checking the math and the assumptions.
A language model doesn’t have that property. Same prompt, same day, can produce a differently worded—sometimes differently reasoned—answer. It can be confidently, fluently wrong in a way a spreadsheet formula error usually isn’t (a formula error tends to look like an obviously wrong number; a hallucinated commentary paragraph reads exactly as polished as a correct one). Your validation process needs to test for that specific failure mode, and most legacy model risk teams have never had to.
Three questions your inventory should actually ask
1. What’s the blast radius if this is wrong and nobody catches it? Not “is this used in external reporting”—more granular than that. A tool drafting internal FAQ answers for a shared drive has a different blast radius than one drafting language that ends up, lightly edited, in an earnings call script. Match your review intensity to that, not to the tool name.
2. Who is the actual human checkpoint, and do they have the expertise to catch a wrong answer? “A human reviews it before it goes out” is not a control if the human reviewing it doesn’t know the underlying number well enough to notice it’s off. I’ve seen review steps that were real controls on paper and rubber stamps in practice because the reviewer was assigned by availability, not by subject expertise.
3. What data is this tool allowed to see, and did anyone check the vendor contract, not just the marketing page? Enterprise tiers vary in what’s retained, what’s used for further training, and what jurisdiction data sits in. This is boring and it’s also the part that actually protects you. If your legal and security teams haven’t reviewed the current contract terms for every AI tool with production or MNPI access, that’s the first gap to close, and it usually takes one meeting to find out nobody has.
What to add, not replace
You don’t need a brand-new governance framework built from scratch—that’s usually an excuse to delay doing anything. Take your existing model risk categories (documentation, validation, ongoing monitoring, incident response) and add an AI-specific supplement that addresses non-determinism, hallucination risk, and data handling. Keep the ownership and escalation structure you already have; people already know how to use it.
The regulatory environment here is still moving—the EU AI Act’s phased obligations and various national guidance are landing at different times through 2026, and U.S. state-level rules are adding another layer on top. Whatever you build should assume the compliance bar moves up, not stay flat, over the next 18 months.
The organizations I’ve seen handle this well didn’t wait for a perfect framework. They took a real inventory of every AI tool actually in use (usually larger than IT’s official list, because shadow usage is real), triaged by blast radius, and started documentation on the three or four highest-risk use cases first. Perfect coverage on day one isn’t the goal. A defensible, improving process is.
Get monthly Finance × AI notes
One concise monthly email with practical finance AI strategy notes, field-tested patterns, and new project updates.
By subscribing, you agree to receive email updates. Unsubscribe at any time.
~Pedro Alizo