Your Finance AI ROI Number is Unfalsifiable, and Your CFO Instincts Should be on Alert

Most finance AI ROI claims are unfalsifiable: they are constructed so that no possible outcome could disprove them. The fix is not better dashboards, it is borrowing the discipline you already apply to every other number you touch. Define the counterfactual before you deploy. Pick one hard metric that moves the P&L. Hold out a control. Pre-register what failure looks like. If you cannot state in advance what result would make you kill the tool, you are not measuring ROI.

The Tell: An ROI Number that Survives Every Outcome

The close came down two days. ROI from AI. The close did not come down, but the team felt less stressed. ROI from AI. Errors fell. ROI. Errors did not fall, but you caught them earlier. ROI. Headcount stayed flat while volume grew. ROI. Headcount grew, but productivity per head rose. Also ROI.

Notice what just happened. Every branch of reality routed to the same conclusion. A claim that is confirmed by all outcomes is confirmed by none. Karl Popper called this the demarcation problem. Your audit committee calls it a finding. If no result could have disproven your ROI number, you did not measure anything. You decorated a decision you had already made.

The Four Ways the Number gets Fabricated, All in Good Faith

This rarely involves intentional misleading, it involves smart people skipping the step they would never skip anywhere else.

1. The hours-saved phantom. You multiply 12 analysts by 5 hours saved per week by a loaded rate and book it as savings. Walk into the room nine months later. Same 12 analysts. Same payroll. The hours did not convert to cash, headcount, or net new output. They converted to slack, and slack is real, but it is not the 1.4 million you put in the deck. Time saved is only ROI if you can name where the time went.

2. The baseline that moves with you. You measure forecast accuracy after deployment and call the improvement AI. But you also changed the cadence, added a reviewer, and cleaned three data feeds in the same quarter. You cannot attribute the lift to any one of them. The AI gets full credit because it is the thing you bought.

3. The denominator nobody books. The license is 200k. The number in the deck is the 200k. Missing: the 0.5 FTE of analyst time spent prompting, checking, and re-checking. The data engineering to wire the source systems. The two weeks the controller spent building trust in the output before she would use it. Real all-in cost is frequently 2 to 3x the license. The honest ROI is a different number, and sometimes it is negative.

4. The survivorship edit. You ran four pilots. One worked. The deck features the one. The three that died quietly are not in the denominator, so your hit rate looks like 100 percent when it was 25.

How to Build an ROI Number that Can Actually Fail

Treat the deployment like a controlled test, because it is one.

Pre-register the kill criterion. Before go-live, write one sentence: if metric X is not at level Y by date Z, we stop. Put it in the project charter. The act of writing it forces honesty, because now a real outcome can end the project. If you cannot write that sentence, you do not understand the use case well enough to fund it.

Pick one P&L-linked metric, not a basket. Days sales outstanding. Close cycle days. Forecast error in basis points against actuals. Dollars of leakage recovered. One number that a non-finance executive would recognize as money. Soft metrics like confidence and morale are real and worth tracking, but they go in a separate column labeled what they are, never in the ROI line.

Hold out a control. Run AI-assisted collections on half the portfolio and your existing process on the other half, matched by aging and segment. The DSO gap between the two is your effect. This is the single highest-leverage discipline on this list and almost nobody does it, because it requires admitting you do not already know the answer.

Measure the counterfactual, not the before-and-after. Before-and-after captures every other change in the quarter. The right question is never did it get better. It is did it get better than it would have without the tool. Those are different numbers and only the second one is ROI.

Book the full denominator. License plus internal labor plus data plumbing plus the trust-building tax. All-in or do not bother.

Top Use Cases, With the Test that Makes Each One Real

  • Collections and DSO. AI prioritizes the accounts, drafts the outreach, predicts who pays late.

The test: hold-out portfolio, matched by aging. Measure DSO delta and dollars recovered against the control, not against last year. Last year had a different economy.

  • Close acceleration. AI auto-reconciles, drafts journal narratives, flags anomalies.

The test: close days is necessary but not sufficient. Pair it with restatement rate and post-close adjustments. A faster close that produces more corrections is not a win. It is a deferral of the work into next month, dressed as speed.

  • Forecasting. AI generates the rolling forecast and the driver commentary.

The test: forecast error in basis points versus actuals, against a frozen baseline model you keep running in parallel. If the AI cannot beat the simple model you already had, you bought a more expensive way to be equally wrong.

  • Spend and AP anomaly detection. AI surfaces duplicate invoices, threshold gaming, off-contract spend.

The test: dollars of leakage recovered that you can trace to a specific flag, minus the cost of chasing false positives. Count the false-positive hours. They are the part the vendor demo never shows.

Pro Tips the Vendor Will Not Volunteer

Run the model against last year before you trust it on this year. Back-test the AI on a closed period where you already know the answer. If it cannot reproduce a quarter you have actuals for, it has no business forecasting one you do not.

Demand the confusion matrix, not the accuracy number. Ask any vendor: of the items you flagged, how many were real, and of the real items, how many did you miss. Precision and recall, separately. A model that flags everything is 100 percent on recall and useless. The single number on the slide is hiding the tradeoff.

Price the cost of being wrong before you trust the output. A misclassified marketing accrual costs you a cleanup. A misclassified revenue recognition entry costs you a restatement. Match the level of human review to the cost of the error, not to a blanket human-in-the-loop policy that everyone clicks through at 4:55 on close day.

Put a sunset clause on every pilot. A pilot with no end date is not a pilot. It is an unmonitored production system with a smaller budget line and no owner.

What Nobody Tells You: The Uncomfortable Parts

The ROI is often real and your number is still fake. Both can be true. The tool genuinely helped and the 340 percent is still fiction, because you never isolated the effect. Sloppy measurement of a good outcome is still sloppy measurement, and it will burn you the day a board member asks how you know.

The honest number is sometimes negative in year one, and that is fine. Most worthwhile finance AI deployments lose money in the first year once you book the full denominator, then compound after. An unfalsifiable positive number robs you of the credibility to ask for the patience the real curve requires.

You will be punished for rigor. The peer who reports 340 percent with no control gets the applause and the bigger budget. You, reporting 18 percent against a hold-out with a stated kill criterion, look conservative in the room. This is a real career tension and pretending it away helps no one. The defense is durability: your number survives the audit, the board follow-up, and next year. Theirs evaporates the moment someone asks for the experiment.

The Constraint that Prohibits Honest Measurement

The reason teams default to unfalsifiable ROI is not laziness. It is that the data required to measure it honestly lives in systems that do not talk to each other. Your hold-out test needs collections data from the ERP joined to payment timing from the bank feed joined to the AI tool's own activity logs. Your close metric needs the sub-ledger detail tied to the consolidation layer tied to the adjustment journal. The all-in denominator needs labor hours from one system and license cost from another. Nobody can assemble that by hand every month, so the honest measurement gets replaced by the vibe-based number that requires no plumbing.

The unfalsifiable ROI claim is, underneath, a data integration failure wearing a finance costume. This is exactly the gap Una is built to close. By unifying revenue, operations, and finance data into a single live source of truth, the full denominator and the counterfactual become computable instead of reconstructed from memory in a slide. Scenario modeling gives you the side-by-side control-versus-treatment comparison the honest number requires, not just before-and-after. And the Action Tracker is the part that matters most for this argument: it turns the plan into owned initiatives with tracked outcomes and measured results, so the saved hours and claimed impact are followed to where they actually landed in the P&L, rather than assumed into existence.

The discipline above is only sustainable if the underlying data is actually joined and the outcomes are actually tracked. Without that, even the most rigorous CFO defaults to the vibe number by month three, not from weakness, but from arithmetic that no human can do by hand.

The Thought Experiment to Run on your Own Deck This Week

Open your last AI ROI slide. Ask your team, if I deleted this tool tomorrow and ran the next quarter without it, what exactly would get worse, by how much, and how would I see it in the P&L? If no one can answer, that is the most important thing you will learn about your AI tool.

Interested in your ROI with an AI-native FP&A tool like Una? Try out the ROI Calculator, and if you have questions about the results, ask Agent Una.

Now Your Turn

What is the most unfalsifiable ROI number you have seen presented with a straight face, and what was the real number once you booked the full denominator? Who here has actually run a hold-out control on a finance AI deployment, and what did the gap turn out to be versus what the before-and-after suggested? Has anyone killed a tool because of a pre-registered criterion, and what did it cost you in the room to do it?

Start Using Una Today

Ready to take control of your budgeting, planning, forecasting and reporting? Schedule a demo and see how Una’s financial planning & analysis software can work for you.

Schedule Demo