Build an LLM Trust Map for Faster FP&A Reforecasting

AI is genuinely useful for about half of what FP&A teams do. The other half is where people are quietly making mistakes they haven't been caught on yet. This post draws the line.

Let's say you're two hours from the board pack deadline. You paste the actuals into ChatGPT, ask it to write the variance commentary, and the output looks pretty good. Clean sentences, right accounting language, confident tone. You do a quick read, change two words, and drop it in the deck.

Three weeks later, your controller asks how you attributed the margin miss to something like unfavourable mix shift. You don't have a great answer. Neither does the AI, since it inferred the driver from the numbers and kinda just invented a story that sounded right.

Now this may look like a horror story about AI being an awful tool, but I'd say it's more of a story about using the right tool in the wrong place.

Here's a simple framework that helps: The 2×2 trust map

Two questions determine whether a task is appropriate for AI right now:

Is the output deterministic or probabilistic?

  • Deterministic = there's one correct answer, and being wrong has consequences (a reconciliation, a balance, a formula result).
  • Probabilistic = the output is a draft, an interpretation, a summary. "Good enough" is the real differentiator here.

Does the output require an audit trail?

  • High audit = SOX scope, board materials, and anything a controller or external auditor will review.
  • Low audit = internal communication, working drafts, and exploratory analysis.

Plot those two axes, and you get four quadrants. Thankfully, the community has collectively, mostly through trial and error, figured out where their work lives on this map.

Where FP&A tasks actually fall

Bottom-right quadrant: Probabilistic output, Low audit requirement

This is where AI earns its keep. Tasks here have real-time savings and low risk:

  • Board commentary drafts, variance narratives, community reports, all things that can save 2-3 hours per close cycle
  • Meeting summaries from call transcripts
  • Exception narrative drafts like "here's why this cost center is over"
  • Email drafts to department heads requesting budget inputs
  • Job descriptions, internal SOPs, process documentation

The keyword is draft. These outputs go through human review before they go anywhere, but the draft is still useful, and the review is fast.

Caution/Verify zones: These two zones are described below and include deterministic output with Low audit requirements (caution) and probabilistic output with more audit requirements (important to verify).

Bottom-left quadrant: Deterministic output, low audit requirement

This is where code generation lives, and it's more trusted than people expect, given the right workflow:

  • SQL queries for ad-hoc analysis
  • Python scripts for data transformation
  • VBA macros for Excel automation
  • DAX measures for Power BI models
  • Power Query transformations

AI-generated code gets tested before it touches anything real. It doesn't need an audit trail, but it does need a human who understands what the code is supposed to do and that combination works well in practice.

Top-right quadrant: Probabilistic output, High audit requirement

These are outputs where there's no single "correct" answer, so the task itself is language/judgment-shaped, but the output still ends up in a formal, defensible context:

  • MD&A commentary in SEC filings
  • Board-approved forward guidance language
  • Non-GAAP reconciliation explanations
  • Audit committee memos and presentations

The output feels like a safe AI use case (i.e., it's natural language, it's a draft), but it requires the human reviewer to have deep enough expertise to catch subtle errors in framing, consistency with disclosed data, or regulatory nuance. AI can draft these, and probably saves real time doing so, but the review bar is substantially higher than for an internal variance narrative.

Top-left quadrant: Deterministic output, High audit requirement

This is where it's probably best to stop.

  • Actuals close and period-end reconciliations
  • Variance bridge math (volume/mix/price decomposition)
  • SOX workpapers and internal control documentation
  • Financial model logic and formula architecture

Here you can find practically anything going into an SEC disclosure or board-approved financial statements.

The issue here isn't just accuracy, but that even a correct AI output without a traceable logic chain is unusable. SOX requires controls you can document and defend. If an auditor asks, "How did you get this number?", "I asked Claude" is not a great answer. As described in Grant Thornton's 2025 guidance on AI in SOX compliance (also a good read link), auditability and explainability aren't optional features; they're the baseline requirement.

The specific failure mode to watch for: AI is particularly dangerous for variance attribution. LLMs are trained on financial language, including volume/mix/price frameworks, so they'll produce confident-sounding variance narratives even when that logic isn't encoded in the underlying data. The analysis looks real, but it's not, and this is a documented pattern, not a theoretical risk.

Why the line exists where it does

It comes down to how LLMs work. They're trained to produce confident, fluent, well-structured output regardless of whether the underlying reasoning is correct. That's a feature for drafting work (fluency is the goal) and a bug for financial analysis (correctness is the goal, and confident wrongness is harder to catch than obvious wrongness).

Current hallucination rates for frontier models "vary widely, from <1% in constrained tasks to over 90% in complex benchmarks", according to this SQ article (link). For a draft email, a 7% error rate is manageable with review. For a variance bridge calculation, going to your CFO, not so much.

The AFP's 2025 FP&A Benchmarking Survey found that only 23% of FP&A professionals use AI daily or weekly, but 40% are in the testing phase (another good read for FP&A people: link). The teams moving from testing to production are exactly the ones who need this map before they deploy.

The map will shift, but hasn’t yet

As tools improve, the line will move. Source-connected AI that can trace its reasoning back to actual data, with audit logs and explainable output, will eventually be deployable in higher-stakes contexts. Some platforms are building toward this now (Una.ai... cough- cough).

But right now, the trust boundary is real, and crossing it without knowing where it is is exactly how a variance memo with fabricated driver logic ends up in front of the board.

The practical move for most teams:

Audit your current AI use against the four quadrants.

The sweet spot zone: expand and systematize; the time savings are real

Caution/Verify zones: use freely with a test-before-deploy rule

No-go zone: hold the line until the tooling has the audit trail you need

The framework in one sentence: Use AI when the output is a draft, and the review is fast; don't use it when the output is a calculation, and the defence has to hold up to audit.

For more information on whether you or your company are using AI to its limits, you may find that Una's AI Readiness Assessment (https://benchmark.una.ai/) can help in testing your current use cases and where you can improve. Give it a try if you're up for it!

Start Using Una Today

Ready to take control of your budgeting, planning, forecasting and reporting? Schedule a demo and see how Una’s financial planning & analysis software can work for you.

Schedule Demo