The Hallucination Tax You Can't Afford as a CFO
The Hallucination Tax You Can't Afford as a CFO
26 Jun 2026 · 8 min read
Key Takeaways
- Generic AI assistants fill in metric definitions from training data. Not from your organization's methodology.
- The hallucination tax is the hidden cost of verifying AI-generated financial insights before you can act on them.
- A semantically grounded AI fetches your organization's definition before calculating — same answer every time, whoever asks.
- AI amplifies whatever data quality you already have, good or bad. The foundation comes first.
Most finance leaders have now tried an AI assistant on real data. Most of them had a similar experience. Here is what is actually happening, and what trustworthy finance AI looks like when it works.
You Already Know the Problem
You ask your AI assistant what gross margin looks like across your European entities. You get a number. Your FP&A lead runs the same question and gets a different one, because they excluded a subsidiary that only went live in Q3, or because their active customer count filters out trials and yours does not. You pull the BI dashboard. A third number. None of them are wrong.
This is not new, and you know it. Multi-entity finance leaders have managed definition drift since the first spreadsheet got emailed across a country border. The question that has changed is not whether your organization has this problem. Almost every organization does, at some level. The question is what happens when you put AI on top of it.
“You can't simplistically look anymore at finance versus business versus operations... you have a business meeting and you spend two thirds discussing who has the right numbers, and the last third on the actual business.
The CFO, in my view, is a good place to have one person responsible for the accuracy of data, with a data dictionary, clearly defined, and then everybody leverages it. That's when you create the real benefit.”
Thilo Kusch
Group CFO, P3 Logistic Parks
The Generalist Problem
When Microsoft Copilot, or a general-purpose AI assistant, reads your financial data and answers a question about gross margin, it does something no one thinks to warn you about: it decides what gross margin means.
Not maliciously. Not randomly. It uses its training data, its understanding of standard accounting terms, and whatever context exists in your data, to produce a reasonable answer. The problem is that "reasonable" is not the same as "according to your company's methodology, which excludes capitalized software costs from the calculation because of a 2022 board decision specific to your industry reporting requirements."
The AI was never told about that decision. It fills in the gap.
A generalist AI decides what your metric means. A grounded AI fetches what your organization decided it means.
This is why finance leaders who have tested general AI assistants against their data report the same pattern: inconsistent results. Ask the same question twice and get different numbers. Ask it in a slightly different way and get a different answer still. The tool sounds authoritative every time, which makes the inconsistency harder to catch and more expensive to correct.
Jiří Maňas, COO at Keboola, frames it precisely: "If you leave AI running on top of your data without the semantic layer, it will run on the assumptions it brings in from the outside. Those assumptions are not your company."
What the Confident Hallucinator Looks Like in Practice
At a CFO breakfast in Prague this April, twenty-six finance leaders compared notes. Around 70% of the room had already tried AI on real financial data. Most of them had watched it not work in a specific and recognizable way.
One CFO showed a year-end forecast that was 30% off from what their BI system reported. The OLAP cube behind the BI system had been built for human analysts over a decade. Nobody had rebuilt the definitions layer for the way a large language model interprets queries. The AI was answering correctly from its own perspective. It was answering incorrectly from the CFO's.
The phrase that recurred across the room: "...but the data isn't trustworthy enough yet."
Michal Hruška, Head of Solution Engineering at Keboola, has been implementing financial data platforms for seven years. His description of what happens in those moments is direct:
“Without the definitions, AI assistants will be confident hallucinators. You need the semantic layer for it to give you correct answers, to give you the same answers every time and to always work in the context that you need it to work in.”
Michal Hruška
Head of Solution Engineering, Keboola
Confident is the operative word. The problem is not that the AI says "I'm not sure." The problem is that it says "Here is your gross margin" and it is wrong, and it is wrong in a different way the next time you ask.
The Moment Kai Actually Works
In Keboola's Financial Intelligence platform, the AI assistant is called Kai. During a demo, Michal Hruška walked through what happens when a finance leader asks about group EBITDA — and how to trace it back to source data.
Before answering, Kai queries the semantic layer to understand what the organization means by EBITDA. It retrieves the exact definition — not an approximation — and confirms it is using it. Then it shows the result alongside the formula, the accounts used, and the full data lineage.
Then it shows what that definition contains: the formula, the adjustments, the scope. Then it gives the result.
Semantically grounded AI (e.g. Kai)
- +Fetches your organization's definition before answering
- +Same answer every time, for every user, however they ask
- +Cites the formula and definition it applied
- +Designed for multi-entity, multi-methodology environments
- +Flags when a metric has no official definition rather than guessing
Generic AI assistant
- −Decides what gross margin means from training data
- −May give a different answer if you rephrase the question
- −Cannot explain which methodology it used
- −Works well on general questions, struggles on entity-specific ones
The consequences are significant for multi-entity organizations. Jakub Žalio, Group CTO at Creditinfo Group, which operates across 30 markets, describes the boardroom experience after building this kind of foundation:
“Imagine sitting in the boardroom and someone asks about a metric in two different countries, and you have the answer in a minute or two.
That is the benefit. That is the ROI. That is when you can make informed decisions without asking anyone else.”
Jakub Žalio
Group CTO, Creditinfo Group
The reason that works is not just clean data. It is that the AI behind the answer is grounded in definitions that were agreed on across all 30 markets. Same question, same answer, whoever asks, wherever they are.
When You Can Trust It, and When You Probably Can't
The question finance leaders should be asking about any AI assistant is not "is this AI good?" It is "does this AI know what my company means by gross margin?"
For generic tools, the answer is almost always no. They are designed to be useful in many contexts, which means they adapt to the data in front of them rather than inheriting your organizational definitions. That adaptability is the source of the inconsistency.
You can trust it when
- It explicitly retrieves a definition before calculating, and can show you which definition it used
- It gives the same answer to the same question regardless of who asks or how they phrase it
- It tells you when a metric has no official definition rather than approximating one
Watch for inconsistent results when
- The AI has no access to your organization's metric definitions and is interpreting terminology from context
- Results vary depending on how you phrase the question
- The tool cannot trace an answer back to a specific formula or methodology
For a CFO presenting to a board, this distinction is everything. A hallucinating AI produces insights that require verification before you act on them. The time your team spends on that verification is the hallucination tax. A grounded AI produces insights you can present directly, because the methodology behind the answer is auditable.
What Changes When You Get This Right
Thilo Kusch is the Group CFO at P3 Logistic Parks. He runs finance across eleven countries. Before AI became a boardroom topic, he built a single data foundation across fourteen systems. The goal had nothing to do with AI. He wanted to stop spending management meetings arguing about whose numbers were correct. That problem got solved first. Then AI arrived. And the foundation he had built turned out to be exactly what AI needed.
Jiří Maňas worked alongside Thilo on that project. He describes what happens when you get the sequence right:
“Without this foundation, you cannot use AI. When you have this, then you can put AI on top and it will do magic.
Without that, you will just amplify the errors you already have in the data.”
AI does not create bad data. It finds it faster. Whatever uncertainty you already have, AI multiplies it — at a speed and scale no analyst can match. The same is true in reverse. Clean, governed data plus AI means answers in the room. Fewer follow-ups. Fewer spreadsheets sent around after the meeting. Decisions made while everyone is still at the table.
Organizations that run AI on a solid data foundation describe a different experience. The numbers can be trusted. Decisions happen in the meeting. Not in a follow-up. Not after someone checks the figures. In the room.
The Investment That Actually Matters
Right now, most finance leaders are asking the same question: which AI tool should we buy? That is the wrong conversation to start with.
The better question is simpler. Does this AI know what your company means by gross margin? Some tools approximate. They make an educated guess based on training data. Others use your definitions — exactly as your team agreed on them, every time, for every user.
The second kind exists today. Kai is one example. But the principle applies beyond any single product. Any AI that works from an explicit, company-owned definition layer will give you consistent, auditable answers. Any AI that does not will give you what most finance leaders have already experienced: impressive in the demo, inconsistent on Monday morning.
The CFOs at that Prague breakfast had seen every wave of this. BI tools. Data warehouses. Dashboards. Analytics platforms. Each one promised to fix the numbers problem. Each one fell short in some way. At this breakfast, they said something they had not said before: this time feels real.
It does. The only question is whether your AI is built to use what your organization already knows.