← Blog

AI Agents for Transaction Monitoring in Cross-Border Payments: A 2026 Guide

Why transaction monitoring queues at cross-border payment companies grow faster than the teams that work them, what an alert investigation actually involves, why regulators forbid the obvious shortcuts, and how AI agents add review capacity inside the systems you already run.

An AI agent for transaction monitoring is a computer-use agent that works a compliance team's alert queue the way an analyst does: it reads the rule that fired, pulls the customer's KYC file, checks the payment's supporting documents, researches the counterparty in registries, sanctions lists, and public sources, and writes a draft case assessment in which every claim is cited to evidence, with a human analyst making every disposition decision. It adds review capacity to the queue. It does not decide what is suspicious, and it does not file anything.

This guide is for the people who own that queue at cross-border payment companies: what an alert investigation actually involves, why the queue grows faster than the team, why regulators forbid the obvious shortcuts, and where agents fit.

Key takeaways

  • Cross-border payment companies drown fastest: in the FCA's Starling Bank final notice, a payment screening review alerted on roughly one in five payments, while the equivalent customer screening review alerted on about 1.4% of customers. Every new corridor adds volume, new registries, and new document formats.
  • Almost all of the work ends in "this was fine." Across 19 large US banks, 16 million alerts produced 640,000 SARs, a 4% conversion rate (Bank Policy Institute). The other 96% still had to be investigated and documented.
  • The investigation blocks the payment. Operators we speak with describe payments that clear screening and then stall one to two days waiting for manual review, document checks, and client RFIs. In trade corridors, a hold measured in hours can mean cargo ships without the money.
  • The obvious shortcut is illegal. U.S. Bank paid $185M in penalties partly for capping alerts to match staffing levels. When volume grows, adding review capacity is the only compliant option.
  • The documentation burden is measured: banks report 21.4 hours per SAR against FinCEN's official 1.98-hour estimate (Bank Policy Institute survey).
  • An agent changes the arithmetic by doing the evidence assembly, which is where the time goes. The judgment stays with your analysts, which is where regulators require it to stay.

Why cross-border transaction monitoring is its own problem

Every regulated payment company screens transactions against rules and watchlists. Screening is a solved, crowded software market. The unsolved part starts one second later, when a rule fires and a human has to investigate the alert.

Cross-border payments make this structurally worse than domestic ones. The screening rate is an order of magnitude higher, because corridors, counterparties, and currencies multiply the ways a transaction can look unusual. The evidence is scattered across jurisdictions: a company registry in one country, a bill of lading issued in another, adverse media in a third language. And the case has a clock on it that domestic monitoring rarely has, because while the alert is open, the customer's money is frozen mid-journey. The pain is deadline-shaped.

There is also a second queue most people outside the industry never see: RFIs, or Requests for Information, arriving at the payment company from correspondent and partner banks through the SWIFT chain. When the request concerns a payment still in flight, operators tell us the clock is measured in days, sometimes as short as two or three, before the bank gives up and returns funds. At scaled players, working that mailbox is a dedicated full-time role.

What an alert investigation actually involves

A single transaction-monitoring alert, worked properly, is a chain of steps across many systems:

  1. Read the trigger. Which rule fired, on what amounts, dates, and corridor.
  2. Pull the KYC file. What business did this customer claim at onboarding, and what activity was expected? (This is the ongoing half of customer due diligence.)
  3. Review transaction history. Is this a new pattern, or normal behavior for this customer?
  4. Examine the payment's documents. For B2B cross-border, that means the commercial invoice, the bill of lading or airway bill, sometimes contracts and customs declarations, checked for internal consistency: do the goods, amounts, ports, and parties line up with each other and with the customer's stated business?
  5. Research the counterparty in public sources. Registry lookups, sanctions and PEP checks, adverse media in English and the local language. The Wolfsberg Group's guidance requires this public-domain review to be completed before any information request goes to the customer.
  6. Answer the relationship question. Does the connection between sender and receiver make commercial sense? This is literally a question on Wolfsberg's standard RFI battery, and it is the heart of the judgment.
  7. Decide and document. Close as a false positive with a written rationale, escalate to a senior investigator, or draft a SAR narrative for the MLRO to decide on. If the evidence is thin, issue an RFI to the customer and wait.

Notice what this list is made of. One step is judgment. Six steps are gathering, cross-checking, and writing.

Why the queue outgrows the team

The volumes are unforgiving: McKinsey puts false positives at more than 90% of transaction-monitoring alerts at most banks, and every one of them needs a documented rationale anyway. Regulators examine the quality of rationale on closed alerts specifically: the FCA's Monzo notice records that in the first half of 2019 the bank closed 45% of its transaction-monitoring alerts as "undecided," and the FCA's review of challenger banks named "inconsistent and inadequate rationale for discounting alerts" as a recurring failure. The negative finding must be documented as carefully as a hit.

When teams fall behind, the failure modes are all on the public record. Cash App's alert backlog grew from 18,000 to 169,000, with an average of 129 days from alert to SAR, before the NYDFS penalty arrived. And the pressure lands hardest on smaller firms: financial-crime compliance costs smaller firms about 2.3% of revenue versus 0.43% at large ones, per the LexisNexis / Oxford Economics UK study.

The two traditional answers both fail:

  • Suppressing alerts is illegal. U.S. Bank was penalized $185M partly for capping monitoring alerts to the number its staff could work; an internal memo flagged the practice as a risk item. Regulators expect financial-crime resources to grow with the business. So when volume grows, review capacity has to grow. There is no threshold dial you are allowed to turn.
  • Hiring and outsourcing scale slowly and unevenly. In-house analysts take months to hire and train. BPO benches take roughly two weeks of training per batch and around four weeks to stand up a new market, and the quality record is public: in Coinbase's NYDFS settlement, over half of 73,000 contractor-cleared alerts failed quality control, and TD Bank's consent order described outsourced review work in language no vendor wants quoted.

This is the trap for a growing cross-border business: alert volume scales with payment volume and with every new corridor, while review capacity scales with hiring cycles.

Where the time actually goes

One investigation commonly touches 10 to 20 systems that do not talk to each other: the case manager, the screening tool's hit detail, the company's own admin panel, email in both directions, SWIFT message viewers, company registries, sanctions search tools, news search, sometimes a paid terminal. The analyst is the integration layer, copying context between screens and assembling it into a narrative.

That assembly, not the judgment, is where the one-to-two-day payment stall comes from. Once the evidence is in one place, the disposition call itself usually takes minutes. The write-up burden explains the most striking number in the category: banks measured 21.4 hours per SAR against the regulator's own 1.98-hour estimate.

And most of that evidence lives behind screens with no API. The US Treasury's sanctions search tool has no API. UK shareholder detail exists only inside downloadable PDF filings. Dozens of company-registry jurisdictions can only be searched by hand. Sponsor and correspondent banks send their requests in their own formats through their own channels, and no fintech will ever get an API into its partner bank's request queue. This is why a decade of integration-first automation never reached this work: the backlog lives in the systems that don't have APIs.

What an AI agent for transaction monitoring actually does

A computer-use agent operates the same screens your analysts already use, so it can do the assembly steps end to end: read the trigger, follow your SOP for that rule type, pull the KYC profile and history, check the invoice and shipping documents against each other, run the registry and sanctions lookups, search adverse media, and write the draft assessment your analyst would have written. Where evidence is missing, it drafts the RFI from your templates. Nothing is sent outward by the agent, and nothing is closed by it either.

Three design properties matter more than speed, because this is regulated work:

  • Every claim is cited. The draft assessment carries numbered citations, and where a check came back clean, it says so explicitly, because examiners expect to see the negative finding stated, not implied.
  • Uncertainty escalates. When the evidence is ambiguous, the agent's job is to hand your analyst a fully assembled case and say so, not to guess. The four-eyes structure your program already has stays intact.
  • The human holds the disposition. Close, escalate, or recommend a SAR remains an analyst's decision, and the filing decision remains the MLRO's. Regulators expressly contemplate outsourced first-level review performed under the firm's control and supervision (Singapore's guidelines to Notice PSN01 and New York's Part 504 both say so directly); what is prohibited is relying on a third party for the monitoring obligation itself. An agent that drafts while your team adjudicates stays on the permitted side of that line by design.

Framed correctly, this is a capacity story, not an alert-reduction story. The enforcement record punishes firms that shrank detection to fit the team. An agent does the opposite: it lets review capacity grow with the queue, so the incentive to suppress detection disappears, and the artifact regulators actually examine, the documented rationale, gets stronger rather than thinner.

Because the judgment is regulated, the agent has to be auditable by design: scoped access, approval gates on anything consequential, and a complete trail of what it did and why. We describe how we build that in the agentic trust framework.

FAQ

Why do cross-border payments generate so many more alerts? More corridors, currencies, and counterparties mean more ways for a transaction to deviate from expected patterns, and payment screening alerts at a far higher rate than customer screening (roughly one in five payments versus about 1.4% of customers in the reviews described in the FCA's Starling Bank final notice). Each new market adds rules, registries, and document formats on top.

What is an RFI in transaction monitoring? A Request for Information: a formal request for documents or explanation, sent to a customer when the assembled evidence is thin, or received from a correspondent or partner bank about a payment in flight. RFIs tied to an in-flight payment carry a short clock, days not weeks, before the bank rejects the payment and returns funds; RFIs on already-settled activity run longer, with the Wolfsberg guidance describing response windows of 10 to 30 business days.

Can an AI agent decide to file a SAR? No. An agent can assemble the evidence and draft the narrative, but the decision that activity is suspicious, and the decision to file, are regulated judgments that stay with your analysts and your MLRO.

Is it permitted to use a third party for alert review? First-level review performed under the firm's control and supervision is expressly permitted in the major regimes (for example, Singapore's PSN01 and New York's Part 504 both contemplate it). What is prohibited is reliance: outsourcing the monitoring obligation itself, or letting a vendor's output stand without your firm's own review and QA.

Does automating investigations let you tune down your monitoring rules? No, and that is the wrong goal. Firms have been penalized specifically for matching detection to review capacity. The compliant direction is the reverse: hold detection where your risk assessment puts it, and scale review capacity to meet it.

Next step

Raise what your team can do.