Thought Leaders
Your Fraud Stack Was Built For Human Adversaries

For the last 30 years, CAPTCHA has done one job: make a machine fail a test that a person passes without thinking. That job is over. Researchers, testing automated solvers against Google’s reCAPTCHA v2 image challenges, reported a 100% success rate, up from roughly 68% to 71% in earlier work, and found no meaningful difference between the number of challenges a bot needed to clear and the number a human did.
If we think about CAPTCHA as a stand-in for an entire category of control built on the same premise- that certain tasks are trivial for a human and hard for a machine, and thus passing the task proves you are human- it’s time for a revisit of control parameters. AI agents have quietly erased the premise across image recognition, form-filling and multistep navigation alike, and most risk teams have not yet rebuilt their stack in response.
Calibrated Against the Wrong Enemy
Every rule in a mature fraud stack was tuned against a specific opponent, be it a human impersonating another human or a bot farm impersonating many humans at once. Velocity checks catch someone opening accounts faster than a person plausibly could. Device-reuse rules catch someone recycling hardware across identities. Both were built around human limits, the physical time it takes to fill out a form and the practical cost of buying new devices.
The same limits do not constrain an AI agent as they do us, humans. It can complete a form in the time it takes a human to read the first field, and it can be provisioned, used once, and discarded at a cost close to zero. Automated traffic now accounts for more than 53% of all web traffic, up from 51% a year earlier, while human activity is down to 47% and still falling. More than a quarter of bot attacks (27%) now target APIs directly, bypassing the user interface entirely to operate at machine speed, and financial services alone absorbed 24% of all bot attacks and 46% of account-takeover incidents in 2025. Rules built around human pacing were never going to catch a class of traffic that has no pacing at all.
Where the Cracks Are Showing
The industry’s early instinct has been to fight AI with more AI, treating a blackbox model as a silver bullet for a blackbox adversary. That instinct makes the problem worse on its own. A fraud model that cannot explain its own decision is nearly impossible to defend to a regulator, a customer, or an internal auditor once that decision gets challenged, and an opaque model trained to catch an opaque threat compounds the opacity rather than resolving it. Defensibility here isn’t a philosophical nicety; it also means being able to demonstrate model governance: exactly which model is live, what changed from the previous version, what information the model was trained on, who approved the change and why. A system built to auto-update the AI model itself was never designed to produce that kind of paper trail.
Another defense required to tell the good from the mischievous is layered intelligence: signal depth across enough independent dimensions that a fabricated identity, a hijacked account or a money mule runs out of places to hide. Each additional dimension is a separate opportunity to catch an inconsistency, a piece of shared fraud infrastructure or behavior that doesn’t match the person on file. No single signal has to carry the full weight of the decision. The broader the context, the more an isolated, ambiguous clue becomes something a fraud team can actually stand behind.
That same principle holds whether the decision gets made by a person or a model. An AI risk engine is not smarter than the data it’s given. Strip away half the dimensions and you get a system reasoning with total confidence from half the picture, and that’s just as true of a machine learning model scoring a transaction in milliseconds as it is of an analyst working a case by hand. It gets more true, not less, as more of that decisioning shifts to agentic systems making the call with no person in the loop at all. Depth of context isn’t only how you catch an agent committing fraud. It’s the precondition for trusting an agent’s decision about anything.
Feed that context into a model whose reasoning stays legible, rather than a single blackbox system guessing at intent from one signal alone, and the decision stops being just accurate. It becomes defensible. That distinction matters more against AI agents than it did against human fraud rings, because an agent leaves a different kind of trace. It doesn’t fidget, hesitate or make the small errors that behavioral biometrics were built to notice. It leaves a footprint instead; in the age of the credentials it’s using and the coherence of a device history that either does or doesn’t stretch back further than the transaction in front of you. It runs a remote-access tool quietly in the background while someone else’s credentials get spent. None of that requires reading a face. It only requires the ability to read the full session details against a broad signal set.
The Human Tell That Isn’t There Anymore
Nowhere is the recalibration more urgent than in identity verification, where fraud teams built an entire discipline around catching human tells, such as an unnatural blink rate, a mismatched lip sync, or a textured artifact in a photo. Deepfake tooling has closed most of those gaps, and criminals no longer need to build that tooling themselves.
Interpol’s Global Financial Fraud Threat Assessment, published in March, found financial fraud drained more than $442 billion from the global economy in 2025, with AI-enhanced fraud now roughly 4.5 times more profitable than traditional methods. The report identifies “deepfake-as-a-service” operations on dark web marketplaces selling synthetic-identity kits, complete with AI-generated video avatars, voice clones and biometric data, that let an attacker build a convincing digital clone from just 10 seconds of audio harvested from a social media post.
The reliable signal was never the ID document. It is continuity: the digital past a real identity accumulates over years of ordinary account activity and behavioral patterns, the kind of history a synthetic identity, or an agent standing in for one, has no time to build. A system looking only for a fabricated ID document will keep failing. A system looking for a missing history has a harder target to fake.
Retooling for a Different Adversary
None of this is an argument for panic, nor is it an argument that human fraud has gone away. It hasn’t. It is an argument that a stack tuned entirely to human limits, pacing, and tells has a structural blind spot exactly where AI agents operate, and that blind spot will not close on its own. Any one of these checks, tested alone, can still come back clean. It’s only when those checks are read together, across enough independent dimensions, that the pattern an isolated check was built to miss becomes visible.
Agentic AI systems can now autonomously plan and execute a complete fraud campaign, from reconnaissance to the ransom demand, largely without a human operator in the loop. The fix is not a rule bolted onto the old stack. It is a rebuilt premise. The test is no longer whether something looks like a human fraudster — it’s whether there was ever a genuine human behind it at all.












