Blog

Before you trust what the agent says, you need to know the agent

With most treasury teams just beginning the adoption of AI for their operations, we find that most treasury AI agents initially run in read-only mode. They offer insight, but don’t take action…yet.

Whether your organization is days or years away from allowing AI to approve or execute transactions, relying on AI-generated insights and recommendations requires trust. Trust has to be earned, and earning it starts with knowing the agent.

After my last blog, I received feedback from someone who argued that trust doesn't matter very much when a human remains in the loop. If the AI hallucinates, the argument went, the human can simply ignore the recommendation.

Of course, a treasury professional can ignore an AI recommendation. But if the only control is ignoring the answer, what value is AI actually providing? A system that produces advice no one can rely on is not improving much of anything. Rather, it is actually adding another review burden.

If you’re using AI for finance, cash, payments, hedging, etc., at some point you’ll want it to perform tasks and work that you can’t easily validate within seconds. While I very much like the idea of testing AI with questions that you already know the answer to, this is the ‘crawl’ stage. As you move towards walking and running, you’ll want AI to make decisions and act on its recommendations eventually. You need AI to be accountable to you, which changes how we think about and validate accountability for Agentic AI.

Why deployment isn't the same as governance

Gartner's data from this year's Finance Symposium makes the gap concrete: 84% of finance organizations have implemented AI or are planning to. Only 7% report high or very high impact. The blockers Gartner names are data quality gaps and unclear ROI. I'd add a third: most organizations have deployed AI into workflows without defining who is accountable for auditing the agent behind the recommendation.

Kyriba's updated Risk Radar survey puts a number on the cost of that gap. Across nine countries, nearly 4 in 5 finance leaders (79%) reported a material financial impact from inadequate risk visibility in the past 12 months. The organizations most actively deploying AI are also the ones most likely to have already absorbed the cost of operating without adequate visibility controls.

The first piece of this series laid out five questions for trusting AI advice. But those questions assume something more foundational: the organization needs to know which agent produced the insight, what it was permitted to access, and what governed its reasoning. AI's opportunity in treasury is to improve reliability across cash, liquidity, and hedging, so finance teams can make better, more valuable decisions. We can call that the cost of getting it right: the value created when an agent's recommendation is reliable enough to act on. Getting there requires trust in our agents. In treasury, where agentic AI is moving from advice toward execution, trust starts with knowing what the agent is doing.

The ‘know your agent’ test

Finance professionals will recognize the structure immediately: it parallels the Know Your Customer (KYC) frameworks that have governed counterparty risk for decades. KYC applies to external parties. Know Your Agent applies to the AI operating inside your own workflow.

The test has four questions. They should be answerable without searching through multiple systems or calling the vendor.

  • Which agent produced the recommendation, and who is accountable for its operation?

  • What data, tools, and permissions did it use?

  • What did it conclude, and what was the basis for that conclusion?

  • Which authorized human reviewed, approved, rejected, or changed the recommendation before action was taken?

If those questions can't be answered cleanly, the organization has a deployment, not a governance model.

Those answers need to sit in the same treasury record as the decision itself, not in a separate AI log. For a recommendation about cash, liquidity, payments, or hedging, management and auditors should be able to see the relevant balances, forecasts, bank statements, exposures, and policies; when those inputs were refreshed; what the agent calculated or excluded; which controls and approvals applied; and who reviewed the next step. Agent activity, payment history, bank data, workflow approvals, and policy checks should form one connected audit trail, so no one has to reconstruct a transaction or payment workflow across multiple systems.

The audit trail tells you what happened. A treasury-specific harness helps ensure that the recommendation was produced within the organization’s operating rules. In practical terms, the agent should show the treasury work, not just the model’s conclusion. Which data did it use? Which policy did it apply? What calculation did it run? What did it exclude? Which exception did it identify? Where was human approval required?

Kyriba’s TAI is designed to operate within that controlled treasury workflow, using the permissions, repeatable processes, and audit controls that govern the underlying decision. The point is not to create a separate AI control framework. It is to make the agent part of the treasury control framework already in use.

That gives finance leaders a practical way to apply the Know Your Agent test: understand what the agent was permitted to see, follow how it reached its conclusion, and identify where human authority entered the workflow. The technology should make those answers easier to produce, not harder to find.

Know which agent you're dealing with

Every AI agent operating in a treasury workflow should be identifiable: which agent produced the output, what was changed, and who authorized it to operate in the context where it was used.

Treasury teams often run multiple AI tools simultaneously: a forecasting model, a payment anomaly detector, a cash categorization assistant, a scenario modeler. When a recommendation surfaces, the appropriate response depends in part on which agent produced it, under what conditions, and whether that agent was operating within its defined scope.

An agent with a narrow mandate to flag intraday exceptions is a fundamentally different thing from an agent with broad access to forecast, categorize, and recommend. The recommendation from each should carry different weight and different governance requirements.

What the agent saw shapes what it said

An agentic sequence should confirm that the required data is present before it produces an output. A treasury team would reasonably expect it to check that bank files arrived, reconciliation completed, and required entity positions are available. Those controls can catch obvious gaps. They do not guarantee that the data is correct.

A bank can send duplicate transactions. A feed can be current but incomplete. A reconciliation can match the wrong items. Data presence is not data reliability. The agent needs to identify those conditions and stop or escalate when it cannot establish a trustworthy data state.

Risk Radar measures the scale of the data-currency problem directly. Fewer than 1 in 4 finance leaders said their organization can quantify the financial implications of an emerging risk in real time. Most take one to six days. Even in the USA, real-time quantification sits at only 23.9%.

Data alone does not make an agent an effective treasury assistant. Consider an agent reviewing an unusual $4 million payment from a subsidiary. Based on the amount and beneficiary, it flags the payment as anomalous. A treasury professional knows that the payment is a recurring intercompany funding transfer, scheduled at month-end, within approved limits, and tied to a known liquidity shortfall.

Without the entity relationship, payment purpose, approval history, and policy context, the agent has identified an unusual pattern, but not necessarily an unusual payment. It has performed the pattern recognition. It has not yet done the treasury work.

A treasury harness supplies that context: entity hierarchies, committed cash obligations, policy thresholds, cut-off times, and approved workflows. Without those inputs, AI can be technically accurate and operationally wrong. That is why an effective treasury assistant needs more than access to data. It needs the treasury context required to interpret that data.

The reasoning needs to be followable

Explainability in treasury is the mechanism that allows a human reviewer to decide whether to act on a recommendation, escalate it, challenge it, or set it aside.

Risk Radar adds a layer to that argument. Fewer than half of finance leaders in any country or industry rate their ability to analyze risk exposure as “high confidence.” In the Technology sector, which leads all industries in AI adoption, the figure is just 36.5%. A reviewer operating with moderate confidence in their own analytical foundation needs to follow the agent's reasoning step by step. Without that, the review step is an approval, not an audit.

Consider FX hedging. A treasury manager knows what to do if their forecast is uncertain versus if they have complete confidence in foreign currency payables and receivables projections. They need this level of transparency from an AI agent so they know what actions to take. The agent cannot operate as a black box. In treasury, black boxes don't belong in the decision chain.

Human oversight needs an address

Human-in-the-loop oversight is the fourth control, and the one that is most vague. The phrase appears everywhere yet is rarely well defined.

Human oversight in treasury works only if the policy, execution and audit trail all completely align. If the sequence of actions - both AI and human - can pinpoint which human, at which step, with what authority to approve, escalate, or reject, and under what conditions that step can be skipped – then the entire workflow can be mapped to policy and later reconstructed by auditors to ensure compliance.

At this early stage of AI adoption, the very important role of approving transactions is often performed by the human-in-the-loop. And even the most basic audit trail will capture that action. But as treasury teams prepare for payment and transaction execution by AI, it is the entire process that must be trackable and enforceable in a single system with a single set of controls. In addition, who performed which action must be distinguishable. We need to know what the AI did.

Before the advice becomes action

The first piece in this series ended with a question: can we prove why the agent gave the advice before anyone acts on it? The Know Your Agent test is the structure for answering that question: four things that should be documented before a recommendation influences a treasury decision.

Organizations that can't pass the test aren't ready to be implementing AI for autonomous execution. The accountability has to come before the autonomy. AI agents in treasury that can't be identified, that don't surface their data state, that can't show their reasoning, and that operate without named human oversight aren't ready to be trusted with the walk stage, let alone the run.

“The AI said so” is not an audit record. Know your agent before you trust its advice.

Written By

Bob Stark

Bob Stark

Global Head of Market Strategy

Bob Stark is the Global Head of Market Strategy at Kyriba and has been a product and go-to-market financial technology leader for 25 years and works directly with clients, partners, and industry influencers to ensure Kyriba is at the forefront of financial technology. He has empowered finance leaders at some of the world’s largest companies, and is a frequent speaker and author on treasury, risk management, and payments.

Related resources

Blog

5 signals from the front lines of finance leadership

Learn more
News

Kyriba brings AI-orchestrated treasury to European enterprises as it hosts KyribaLive Exchange “KLX” London

Learn more
Success Story

Voltava builds scalable, AI-powered forecasting with Kyriba Liquidity Planning

Learn more