Skip to main content

CFOs Are Starting to Ask the Question Nobody Wanted to Ask in 2024

6 min

Two years ago, the question in almost every finance leadership meeting was “what are we doing with AI.” Nobody wanted to be the department with nothing to show. Budget got approved on vibes and a good demo.

That era is ending. The question I’m hearing now, from peers at other companies and in my own conversations, is different: “we’ve been running this pilot for five quarters—what did it actually save us.” And a lot of teams don’t have a clean answer, because nobody set up the measurement at the start, back when the goal was just to look like you were doing something.

Why the honest answer is uncomfortable

Most AI pilots in finance were sized against soft goals: “save analyst time,” “improve accuracy,” “modernize the stack.” Soft goals don’t survive a renewal conversation with a CFO who’s now under pressure themselves to justify software spend line by line. If you can’t say “this saved 6 hours a week across 4 analysts, worth roughly $X in redeployed capacity, and reduced variance-commentary turnaround from two days to same-day,” you’re going to lose the tool in the next budget cycle regardless of whether it’s actually useful—because “useful but unmeasured” and “not useful” look identical on a spend review.

I think a meaningful number of genuinely good tools are about to get cut for this reason. Not because they don’t work. Because nobody built the measurement discipline alongside the pilot, and now there’s nothing to point to except anecdotes.

What actually measurable looks like

The teams I’ve seen survive this scrutiny did three unglamorous things from day one of the pilot, not at renewal time:

  • Picked a baseline before turning the tool on. How long did this task take last quarter, with the old process, measured the same way you’ll measure it after? If you don’t have a “before” number, you can’t prove an “after” number means anything.
  • Tracked adoption, not just access. A tool with 40 licenses and 6 active users isn’t a finance transformation story, it’s a procurement problem. Track weekly active use per workflow, not seat count.
  • Separated time saved from quality improved, and measured both, because they get argued differently in a budget meeting. Time saved is easy to quantify and easy to challenge (“did they really redeploy that time, or did they just work fewer hours”). Quality improved is harder to quantify and, honestly, often the bigger real win—fewer restatements, fewer late nights before close, commentary that a business partner can actually act on instead of ignoring.

The uncomfortable honesty part

Some of what got funded in 2024 and 2025 genuinely wasn’t worth it, and the right answer for those pilots is to say so and shut them down cleanly rather than keep them alive on momentum. I’d rather be the person in the room who says “we tried this, it didn’t move the number we said it would, here’s what we learned, here’s the one thing we’re keeping”—that’s a stronger position with a skeptical CFO than defending a tool nobody can prove is working.

The pilots that survive the next 12 months of budget scrutiny won’t be the ones with the best demo. They’ll be the ones where someone did the unglamorous work of measuring from the start.

Get monthly Finance × AI notes

One concise monthly email with practical finance AI strategy notes, field-tested patterns, and new project updates.

By subscribing, you agree to receive email updates. Unsubscribe at any time.

~Pedro Alizo