Most discussion about AI agents treats autonomous software making consequential decisions as something new. Card fraud systems have been declining transactions without human involvement since the 1990s. They do it at enormous volume, with real money and real customers on the other side of every call.
That industry has already made most of the mistakes the current wave is making. It also worked out the answers. Those answers are unglamorous and directly transferable, which is why fraud detection is such a useful place to start before treating agents as an entirely new discipline.
Fraud systems already behave like agents
Strip away the vocabulary and a fraud system matches the definition people now use for an agent. It reads what is happening around it, decides on its own, and takes an action that has consequences.
The system sees several kinds of information before it decides:
- Transaction attributes.
- Account history.
- Device signals.
- Network context.
It then makes a decision autonomously, within milliseconds, with no human in the loop. The action is consequential because the system can approve, decline, or challenge a transaction. It also operates continuously at a volume no human process could handle.
Natural language is the only capability missing from that picture. Everything else about the governance problem is identical. You still have an autonomous system making decisions, affecting customers, and doing it fast enough that manual review cannot sit in the middle of every case.
The expensive error is blocking the legitimate customer
Naive fraud systems optimise for catching fraud. Mature systems optimise for the total cost of both error types, and the ratio is not intuitive.
A missed fraud costs the transaction value. A wrongly declined legitimate transaction costs the transaction value plus the customer relationship, the support contact, and frequently the customer’s future business. Industry work has repeatedly found the second is several times the first in expected value.
The agent parallel is exact. An agent that fails to act costs an opportunity. An agent that acts wrongly costs the action, the reversal, the trust, and the investigation. Teams building agents almost universally optimise for capability. That is the agent version of optimising for catch rate, and teams discover the asymmetry after deployment.
This is one of the first lessons fraud teams learned the hard way. A system can look strong on the narrow measure it was trained to improve and still create a larger business cost somewhere else. The same pattern appears when an agent is rewarded mainly for getting things done.
Uncertainty needs a third path
Early fraud systems were binary: approve or decline. Modern systems have a third option that carries most of the value, which is challenge.
Challenge can take a few forms:
- Ask for another factor.
- Send a verification.
- Hold and check.
The third path is where uncertainty goes. It converts a forced binary decision under uncertainty into a deferred decision with more information, at modest cost.
Most agent designs have no equivalent. They act, or they leave the action undone. Adding a third state, escalate with context, is the single highest-return design change available to most agent systems. It is copied directly from an industry that worked this out two decades ago.
That design matters because uncertainty does not disappear when a system becomes autonomous. Fraud systems learned to give uncertainty a place to go. Agent systems need the same place, especially when the action has money, customers, or compliance attached to it.
Human review should start where the loss is highest
When a fraud system flags transactions for review, the queue is not ordered by arrival. It is ordered by expected loss, which means the score multiplied by the amount.
The reason is capacity. Analysts can review a fixed number per day. Reviewing transactions in arrival order means spending the same attention on a forty pound transaction as on a forty thousand pound one. Ordering by expected loss means finite human attention is spent where it changes the most money.
We built exactly this for a payments client, and it was the single change that produced most of the two million a year we removed. The model was not dramatically better than their rules engine at detection. The queue ordering was what converted detection into savings.
Almost no agent system does this. Escalations arrive in a list, in order, and a person works down it. The reordering is free and the improvement is large.
This is a plain operational point, not a modelling breakthrough. If people can only inspect a limited number of cases, the order of that work becomes part of the system’s value. Fraud teams learned that a good score is less useful when the review process treats every flagged item as equal.
A decision needs reasons people can inspect
Card networks and regulators require that a decline can be justified. They do not need mathematical detail, but the explanation must be good enough for a dispute to be adjudicated and for a customer to be told something true.
This constraint shaped the industry’s architecture. Pure black-box scoring, however accurate, was insufficient, so systems produce reason codes alongside scores. Reason codes are structured labels that show which factors drove the decision. What first looked like a limitation turned out to be the property that made the systems deployable, auditable, and improvable.
Agents face the same requirement and are mostly ignoring it. An agent that took an action needs to be able to say, in terms a person can evaluate, what it was responding to. That does not mean a chain of reasoning text, which is a plausible narrative rather than a record. It means the specific inputs that drove the decision.
This distinction matters. A generated explanation can read convincingly while failing to show what the system actually used. Fraud systems solved that by recording structured reasons as part of the decision itself, so a later review has something stable to examine.
Attackers adapt to your defences
Fraud is adversarial. Every rule you deploy is a rule the other side learns and works around. That is why rules engines lose over time, and why the industry moved to models that could be retrained.
Agents operating on content supplied by third parties are in the same position, and largely do not know it. Prompt injection is the direct analogue of the techniques fraudsters used against rule-based systems: find the input that produces the behaviour you want.
The fraud industry’s answer was architecture rather than better rules. Separate the components that read untrusted input from the components that take action, so that influencing the first does not directly control the second. That is exactly the right pattern for agents, and it is not the pattern most of them use.
The business consequence is the same as it was in fraud. If the party on the other side can learn how your system behaves, they can shape their input around it. Rules alone become a brittle defence once the other side is adapting.
You need a holdout to know whether the system works
Fraud teams maintain populations they deliberately do not act on, so they can measure what would have happened. It is expensive. It means knowingly allowing some fraud through. It is also the only way to know whether the system is actually working, rather than confirming its own decisions.
The alternative creates a feedback loop. The system declines what it thinks is fraud, those transactions never complete, and their outcome is unknown. Accuracy measured only on the population you acted on gives the wrong answer.
Agent systems have precisely this problem and no equivalent practice. When an agent takes an action, the counterfactual is gone. Building a holdout where the agent proposes but does not act, and comparing against what humans did, is the only honest measurement, and it is almost never done.
This is difficult because it deliberately withholds automation from some eligible cases. Fraud teams accepted that cost because measurement without a holdout is circular. Agent systems face the same measurement problem whenever the action changes the outcome being measured.
The model is only a small part of the product
The most striking thing about a mature fraud platform is how little of it is the model. The platform contains many parts around the model:
- Feature computation.
- Real-time data joins.
- A decision layer combining model output with hard rules.
- The challenge mechanism.
- The review queue.
- Case management.
- Feedback capture.
- Retraining.
- A governance process controlling what may deploy.
The model is a component. The system is the product. Teams building agents today are building the model and calling it the system, which is the same mistake fraud teams made in the 1990s and stopped making for good reasons.
This is the part many agent projects understate. The useful system includes the decision path, the fallback path, the review process, the measurement process, and the controls over deployment. Fraud platforms became reliable because the industry treated all of that as part of the product.
A refund agent shows how the fraud pattern transfers
Abstract parallels only matter if they change what you build. Here is the concrete version applied to a support agent that can issue refunds.
The design has the same controls a fraud platform learned to use:
- Three outcomes: issue, decline, or escalate with a written reason and the specific evidence. The third path is the default when confidence is below threshold, and it is not a failure state, it is the designed behaviour for uncertainty.
- Queue ordered by exposure: escalations are ranked by refund value multiplied by uncertainty, not by arrival. The agent handles the routine volume; the human attention lands where the money is.
- Reason codes, not narrative: every decision records the structured factors that drove it, including order age, prior refund count, item category, stated reason, and account tenure. A paragraph of generated explanation reads convincingly and is not a record of anything.
- Read and act separated: the component that reads the customer’s message has no credentials. It produces a structured assessment. A second component, which never sees the raw message, decides and acts on that structure. Injection through the message can influence the assessment; it cannot reach the refund API.
- A holdout: a percentage of eligible cases are routed to humans regardless, and the agent’s proposal is recorded but not executed. Comparing the two populations is the only measurement that is not circular.
- Hard limits outside the model: maximum value, maximum per customer per period, and categories excluded entirely. Deterministic code, not instructions in a prompt, because a limit that can be argued with is not a limit.
None of that is novel. This is a fraud platform with different nouns.
The example also shows why the lessons transfer cleanly. The decision may be a refund rather than a card authorization, but the same governance questions appear. What can the system do, when does uncertainty move to a person, what evidence is recorded, and how do you know whether the automation helped?
This framing makes risk review easier
There is a practical benefit beyond the engineering, and it is the reason we now open enterprise conversations this way.
Security and risk functions are sceptical of agents and reasonably so. They are not sceptical of fraud systems, which they have governed for years and understand. Presenting an agent as the same class of system, with the same controls, moves the conversation from novel and frightening to familiar and reviewable.
The questions become ones the organisation already knows how to answer:
- What is the decision boundary.
- What are the reason codes.
- Who reviews the queue.
- What is the holdout.
- How is the limit enforced.
A risk committee can process those questions. It can process them much more readily than a demonstration of a model doing something impressive.
We have watched this reframing shorten a review from months to weeks, with no change to the underlying system. The difference was purely that the system was described in a vocabulary the reviewers already trusted.
Agents are harder in several ways
Being fair to the differences matters, because the differences are real.
Fraud decisions are narrow and repeated: the same shape of decision, millions of times, with a clean feedback signal when a chargeback arrives weeks later. Agent decisions are heterogeneous, and often there is no feedback signal at all because nobody records whether the outcome was good.
Fraud has a natural loss function measured in money. Many agent tasks do not, which makes both optimisation and evaluation harder.
Fraud systems also operate in a regulatory framework built over decades. Agents operate in one being written now, which cuts both ways: less constraint today, and less clarity about what will be required.
Those differences make agents harder. They are reasons to take the transferable lessons more seriously rather than less.
Payments already learned the agent lessons
An industry has been running autonomous consequential decisions at scale for thirty years. It concluded that the third option matters more than accuracy, that queues should be ordered by expected loss, that decisions must be explainable to be deployable, that adversaries adapt so architecture beats rules, and that you cannot measure a system by the decisions it made.
Current agent discourse leaves out these lessons, and all of them apply. The fastest way to build a deployable agent is to read what payments already learned.
More on how we apply this on the financial services page, and the fraud work itself is written up as a case study.








