
Before agentic AI acts in treasury, CFOs need to know if they can trust its advice

By Bob Stark
Global Head of Market StrategyShare
Mastercard’s Agent Pay announcement is a signal that AI agents are moving closer to the payments layer. For consumers, that may mean AI agents that can shop, subscribe, or transact within defined permissions. For treasury leaders, it raises a more interesting question: what happens when an AI agent is not just advising but actually executing financial decisions?
That next level matters. A flawed AI-generated insight can be challenged. A flawed recommendation can be reviewed. But a confirmed payment, investment, or FX hedge is harder to unwind.
That’s why treasury leaders should not wait until autonomous execution arrives to define how they will trust AI and where humans are in-the-loop. The first stage is more basic, and more immediate: can treasury teams trust AI-generated insights and recommendations?
Many AI readiness conversations in finance start with data quality. And I understand why. Treasury cannot trust incomplete data, whether that means missing bank reporting, unreconciled cash positions, or sketchy payables and receivables data. AI can help identify and enrich missing data, but without a proper harness, meaning treasury-specific workflows, controls, best practices, and domain expertise built around the LLM, it can just as easily create more problems than it solves.
Without a harness, the same question asked of the same data can produce different answers on different days. Similar behavior may apply when AI is asked to find missing bank transactions, categorize cash flows, or interpret liquidity trends. In treasury, that lack of precise repeatability is not a technical nuisance. It is a trust problem.
Trust requires confidence in the data, but it also depends on understanding how the agent was built. CFOs need to know which parts of the analysis are driven by the LLM, which parts are governed by treasury-specific controls, and where repeatable workflows are enforcing consistency.
As the ancient Greeks might have said, “know thy agent.”
The first trust problem: AI as advisor
For many treasury teams, the first practical use of AI will not be jumping straight into automated decision making and transaction execution. It will be offering insights and recommendations.
That may include questions like:
Why did cash move differently than expected?
Which entities are showing unusual liquidity patterns?
Which exposures are most sensitive to a rate or FX shock?
Which payments are suspicious?
What happens to our cash if the Strait opens or closes again?
What should we do next?
These are high-value use cases. They can help finance teams deliver meaningful input and faster interpretations when something material changes. Yet even in this advisory role that AI supports, trust matters.
One example not too long ago was where a company had data duplication in their data lake. One of their banks sent intra-day reporting in both MT and CAMT formats for a period of time. Prior day wasn’t the issue; but intra-day duplication occurred. With no harness to help the LLM know that the same transactions were consistently being sent twice, their AI model created forecasts built upon phantom patterns.
While that’s clearly not the same as an AI agent initiating, approving, and releasing a payment, it still created a fear factor around relying on AI to navigate the nuances of daily treasury life.
If an AI tool recommends a cash movement, identifies a potential shortfall, or flags a payment exception, treasury teams need to feel confident that the calculations are right. They need to know that AI completed the job like they would have done it themselves.
This exemplifies the need for treasury expertise to be built into treasury agents, with those same skilled agents displaying their reasoning to “show their work”, so to speak. Otherwise, AI becomes another black box layered on top of already fragmented treasury data.
Data quality matters
A lot of treasury leaders talk about clean data as a precursor for successful AI. I’m typically in that crowd too; I often say that there is no AI strategy without a data strategy. But does that mean that data needs to be perfect? I would argue no. Because if that was the price of admission, then no one would ever get started with AI.
In Kyriba's 2026 CFO Survey, only 40.7% of CFOs reported having fully complete, real-time cash visibility. This is meaningful because nearly six in ten finance leaders are reporting imperfections in their cash reporting. It also means that their AI analyses may operate with delayed, partial, or fragmented cash data. The survey didn’t ask how many have 90% or 95% visibility, but we all know that the number of finance teams that have come to terms with 90% or more (but not 100%) is significantly higher.
What does this mean for AI? It means that you can have incomplete information and still make complex liquidity decisions, including hedging cash exposures. Treasury teams operate quite effectively without all the data. But what they can’t sacrifice is having wrong data.
AI can help fill in the gaps of your cash visibility, or even make assumptions about what we don’t see. But AI needs the data you do have to be correct. Reliability is important, otherwise AI is built on top of a shaky foundation, and when asked to model liquidity scenarios or flag payment anomalies, the risk is not just that the output is wrong. The real issue is that the output may sound confident while being wrong.
Now, in that same CFO survey, 31.2% of CFOs identified data reliability as a critical operational priority for 2026. That is a meaningful number, but it also means that nearly seven in ten CFOs aren’t treating data reliability as a priority. That is concerning, especially as AI becomes more involved in treasury analysis and recommendations. It also shows where AI readiness needs to begin: with reliable data and clear controls around how agents use it.
Five questions before treasury trusts AI advice
Before deploying AI into treasury workflows, CFOs and treasury leaders should be able to answer five core questions.
1. Can we reproduce AI’s answer?
Treasury teams cannot rely on insights that sometimes change every time the same question is asked of the same data. If AI flags a liquidity risk, recommends a cash movement, or identifies a potential fraud event, users should be able to understand whether the same conditions would produce the same recommendation again.
Repeatability matters. In an un-harnessed agent, the same prompt against the same data can still produce different answers. That may be acceptable when AI is summarizing a document. It is not acceptable when the output may influence liquidity, payments, hedging, or investment decisions.
“The AI said so” is not a treasury control. But neither is “the AI said something different this time.”
2. Do we know what data, tools, and skills the AI used?
AI is only as reliable as the data and capabilities it is allowed to use. If a bank feed is delayed, an ERP file is incomplete, reconciliation has not happened, or an agent uses the wrong tool for the task, the output may still sound confident.
Treasury teams need visibility into source data, data freshness, exclusions, known gaps, and the skills or tools used to generate the recommendation. A recommendation based on stale or partial data should not carry the same weight as one based on complete and reconciled information. A recommendation produced outside an approved harness should not carry the same weight as one generated through controlled data access, defined skills, and repeatable workflows.
3. Can users challenge the recommendation?
Trust does not mean accepting every output. It means knowing how to interrogate it.
In practice, I have seen treasury teams ask whether an AI-generated recommendation is accurate. I have seen fewer ask what alternatives the system considered, what assumption would change the answer, or whether the recommendation still holds if one source is stale.
Treasury users should be able to ask why a recommendation was made, what alternatives were considered, and what conditions would change the answer. If a liquidity shortfall is flagged, users should be able to test whether the conclusion changes under different assumptions or data corrections.
4. Where does human review happen?
Even when AI is only advising, human oversight needs to be explicit. A cash forecast recommendation may require review by treasury operations. A hedge recommendation may require input from risk management. A payment exception may require escalation before any action is taken.
This is where many governance conversations become too abstract. “Human-in-the-loop” sounds reassuring, but it only matters if the organization can name the human, get the decision right, and show where that review happens before the recommendation becomes action.
Human-in-the-loop governance does not need to slow everything down, but it does need to define where judgment and accountability enter the process.
5. Can we test AI against real treasury scenarios?
Readiness is not theoretical. Treasury teams should test AI against the conditions they actually face: a delayed bank file, a liquidity shortfall, an FX shock, a sanctioned counterparty, a payment anomaly, or a failed approval.
The purpose is not simply to see whether AI produces a plausible answer. It’s to see whether the organization can explain, challenge, escalate, or reject that answer before it influences a treasury decision.
In every treasury AI deployment we have guided at Kyriba, the teams that moved fastest were not the ones that experimented most aggressively. They were the ones that defined their escalation paths, approval boundaries, and data transparency requirements before the agent entered the workflow.
Crawl before you run
As AI agents move closer to financial execution, the trust question becomes more urgent, but treasury should not start with the moment an agent moves money. It should start one step earlier, with the recommendation that leads someone to approve a payment, investment, hedge, or cash sweep.
That is the crawl stage of AI-driven treasury: proving that AI-generated insights and recommendations are repeatable, explainable, and governed before AI-initiated action ever arrives. The next stages will be harder because auditing AI advice is not the same as auditing AI action after execution.
73.5% of CFOs in our survey said AI and emerging technology fluency is the most important skill for future finance leaders. The first test of that fluency is not whether a CFO can recognize the promise of AI. It is whether the organization can govern the advice AI gives before that advice becomes action.
Before AI acts, treasury teams need to answer one question: can we prove why the agent gave the advice before anyone acts on it?
Written By

Bob Stark
Global Head of Market Strategy
Bob Stark is the Global Head of Market Strategy at Kyriba and has been a product and go-to-market financial technology leader for 25 years and works directly with clients, partners, and industry influencers to ensure Kyriba is at the forefront of financial technology. He has empowered finance leaders at some of the world’s largest companies, and is a frequent speaker and author on treasury, risk management, and payments.
Related resources


