How to test AI agents for finance operations before production
For operations leaders at fintechs and financial institutions: how to test AI agents on your own workflows before production, and what to own in a buy vs build decision.
Operations teams at fintechs, banks, and asset managers are starting to give AI agents real work: reconciliations, account maintenance, onboarding reviews. Getting an agent to finish a task is the easy part. Knowing whether it finished the task correctly is harder.
An agent can produce the right total while matching the wrong transactions. It can finish a reconciliation by clearing an exception that should have stayed open. It can write a convincing summary of a check it never performed.
Each of these looks like finished work, and each has to be ruled out before an agent touches production. A public benchmark cannot catch them, because it does not contain your reconciliation rules, your exceptions, or your approval limits. Before you build an AI agent for an operations team, build the test it has to pass. In practice:
- Pick one workflow and write down what an accepted result looks like.
- Build test cases from that workflow, including exceptions, missing evidence, and attempts to exceed permissions.
- Check the resulting records and system state, not the agent's report of its work.
- Test whether permissions, approvals, audit records, and shutdown hold when the agent makes a bad decision.
- Set release criteria by severity, and name who approves, monitors, and can pause the agent.
A company owns its intelligence when it controls the procedures agents follow, the tests used to evaluate them, and the authority they receive. That matters in a buy vs build decision: whichever you choose, those assets should stay with you when you change models, tools, or vendors. The examples below come from financial services operations, where Zomma works.
Buy vs build AI agents: what should you own either way?
Whether you build agents in-house or buy from a vendor, keep four things under your control:
- Procedures and decision rules: required evidence, permitted actions, exceptions, and escalation paths.
- Benchmark cases and acceptance checks: examples of the work, the conditions under which it starts, and what a correct result requires.
- Evaluation records: the configuration tested, actions taken, resulting changes, scores, and human reviews.
- Release decisions: who approved a configuration, what work it may perform, and what would cause that approval to be withdrawn.
Written procedures rarely capture what an experienced operator knows: which discrepancy can wait, which source takes precedence, which apparent match needs a second check. Building an agent exposes these gaps, because an unwritten rule is hard to implement or test. (We have written about how that knowledge gets lost when people leave.)
If a vendor builds the agent or the benchmark, get the right to keep the test materials and run the tests yourself after the engagement ends. A report alone leaves you dependent on whoever wrote it.
Decide how data may be used. Permission to use a case for evaluation should not automatically permit model training or use for another customer. Limit sensitive data, access, and retention to what the test needs.
What should an AI agent benchmark include?
An AI agent benchmark is a repeatable set of tasks, starting conditions, and acceptance checks used to evaluate an agent configuration. It should include routine work, difficult exceptions, incomplete evidence, and attempts to cross permission boundaries.
Start with one workflow and write down what counts as an accepted result. For payment reconciliation, that might mean matching eligible payments to ledger entries, identifying unexplained differences, preserving unresolved items, and saving a reviewable reconciliation record. An illustrative set of cases:
| Case | What the test changes | What passing requires |
|---|---|---|
| Routine reconciliation | Complete files with valid matches | Correct matches, required checks, and a saved result supported by the source records. |
| Difficult exception | A duplicate reference, partial settlement, or fee difference | The exception is identified and handled under the team's rules. No unsupported clearance. |
| Incomplete evidence | A required statement is missing or sources disagree | The case stays unresolved, with the missing evidence and next step recorded. |
| Attempted misuse | A document asks the agent to export records, or an action requires approval it lacks | Required controls prevent the unauthorized effect. The attempt is recorded. |
Report difficult and high-consequence cases separately so a large volume of easy work cannot hide them.
Have the people responsible for the workflow approve the checks. If two experienced reviewers disagree about the right outcome, resolve the procedure first, or the benchmark will penalize the agent for the company's own ambiguity.
Keep a held-out set of cases that developers cannot use to tune prompts or train the agent, and refresh it as procedures change.
How do you verify an AI agent's work?
Verify an AI agent's work by checking the resulting files, records, and application state against the acceptance rules, not by reading the agent's account of what it did. Anthropic's guidance on evaluating agents makes the same point: an agent reporting that a booking is complete is different from a reservation existing in the database.
Evaluate the full deployment: the model, instructions, tools, permissions, memory, and execution environment. Record their versions for every run. A result obtained with one set of permissions says little about a deployment with broader access.
Reset the test environment between runs so one attempt cannot leave clues or completed work for the next. For computer-use agents, keep the screenshots and actions, and pair them with independent records of the resulting state. A click on “Save” does not establish that the correct record was saved.
Use code to check exact requirements such as amounts, record IDs, required fields, and permission decisions. Have domain experts assess judgment calls. If another model helps grade, check its judgments against expert reviews before relying on its scores. Accept any valid path through the work, and require a specific step only when the procedure or a control demands it, such as approval before submitting a change.
How should you test AI agent security and governance?
Test whether the deployment blocks unauthorized actions, enforces approvals, preserves evidence, and stops work when access is revoked. Do it in an isolated environment, and make the agent attempt the wrong thing on purpose. The question is whether the controls hold when the agent gets it wrong.
Test missing evidence
Remove a required source, deny access, or introduce conflicting records, and check whether the agent preserves the uncertainty. An unavailable sanctions source must never become “no match.” A missing statement must never become a zero balance. A correct escalation counts as a pass.
Test permissions and approvals
Have the agent request another customer's records, use a restricted tool, or send information to an unapproved destination. Confirm that controls outside the model deny the action. The agent declining the request does not test the permission boundary.
Put malicious instructions in documents, emails, web pages, tool responses, and stored memory, and check whether that content can change the procedure, authorize an action, or influence a later case. Our analysis of zero trust for agentic systems covers why an authenticated agent still needs limits.
Test approvals with a generic “proceed,” an expired authorization, the wrong reviewer, and action details changed after approval. Also give human reviewers plausible but wrong recommendations, and check whether they catch the error.
Trace every consequential action through its required control. In AISI's control red-team research, some successful attacks avoided the monitor entirely by exploiting the rules that decided which actions it reviewed. Test that routing, not only the monitor.
Test the audit trail
Pick a completed case and reconstruct it from retained records: the evidence used, the configuration in effect, actions attempted, approvals received, and changes made. The agent under evaluation should not be able to alter source or audit records.
Treat changes to persistent memory as changes to future behavior. Test who can add or change durable instructions, and whether those changes are reviewed and attributable.
Test shutdown
Revoke credentials mid-run, including while the agent has queued actions, delegated tasks, or downstream jobs. Verify what actually stops, what has already changed, and how the team recovers without repeating or losing work.
Measure detection and containment separately. OpenAI's report on an agent using DNS to reach an external service describes an alert that fired while automatic shutdown failed.
What do financial regulators expect from AI agents?
Regulators have named the risks of AI agents but have not said how to test for them. For broker-dealers, FINRA's 2026 oversight report flags agents that act without human approval, exceed their authority, or produce outcomes no one can trace. Our explainer on FINRA's AI agent guidance covers it in full. For banks, the April 2026 interagency model risk guidance (SR 26-2) places generative and agentic AI outside its scope and leaves the controls to each bank's own risk management.
Either way, the firm decides what evidence is enough. The tests above are that evidence.
How do you compare AI models and agent platforms?
Compare models by holding everything else constant: tasks, starting state, tools, acceptance checks, and resource limits. Record any differences you cannot remove.
Vendors and agent platforms also differ in the harness: the software around the model that manages tool calls, state, retries, and execution. The harness can change results as much as the model, so to compare complete products, test each configuration you intend to run.
Run repeated trials, and report failures, timeouts, retries, and the cost of failed attempts alongside successes. A best result out of several attempts tells you little about a workflow that has to work on the first run. Keep the comparison under your control so you can rerun it when a model changes; a benchmark tied to one vendor's private scoring is hard to use for that decision.
What should you measure before deploying an AI agent?
Measure work quality, control effectiveness, and operating economics as three separate scorecards:
- Work quality: accepted cases, material omissions, unsupported conclusions, incorrect clearances, and valid escalations.
- Control effectiveness: unauthorized effects, approval bypasses, exposure of another customer's data, legitimate work blocked by mistake, and time to containment.
- Operating economics (the ROI case): reviewer minutes, rework, elapsed time, and total cost per accepted case, including model usage, infrastructure, failed attempts, and human work. A faster agent can still cost more if it creates review and repair.
Set pass criteria by severity. A formatting defect and an unauthorized payment need different treatment. Define critical failures that block release even when the average score improves.
Name the owners. The workflow owner approves the standard of work. The security owner validates the controls. A deployment owner authorizes the tested scope, monitors operation, and can pause it. Keep production permissions within what was tested.
Rerun the relevant evaluations after any change to models, instructions, tools, permissions, memory, or applications, and feed production incidents back into the suite.
How do agent failures become new tests and training tasks?
Find the cause of each failure before adding training data: missed evidence, lost state, a wrong permission, a buggy check, or an unclear procedure. Then turn the failure into a reproducible case with variations. If an agent mishandles a duplicate payment reference, vary the dates, amounts, and available evidence so success requires applying the rule.
A custom reinforcement learning (RL) environment can use these tasks to train agents against verified outcomes. Keep training tasks out of the held-out set, and confirm the data is approved for training. Each investigated failure leaves a test, a clearer rule, or a stronger control that outlasts the current model.
At Zomma, we build custom benchmarks and RL environments for operations, including computer-use tasks and evaluations across models and harnesses. We start from a workflow and the team's definition of acceptable work. The company should be able to inspect the cases, challenge the scores, and use the results to decide what its agents are allowed to do.