Interviews

Dr. Musheer Ahmed, PhD, Founder and CEO of Codoxo – Interview Series

mm
Add Unite.AI to your preferred sources on Google

Dr. Musheer Ahmed, PhD, Founder and CEO of Codoxo is a technologist and entrepreneur focused on applying artificial intelligence to solve systemic inefficiencies in healthcare. He founded Codoxo based on research developed during his PhD at the Georgia Institute of Technology, where he built the foundations of a patented AI approach to detecting fraud, waste, and abuse in medical claims. Under his leadership, the company has grown into a provider of AI-driven payment integrity solutions, helping healthcare organizations identify risks earlier and shift from reactive audits to proactive cost containment. His earlier experience in security intelligence at VeriSign (VRSN ), where he worked on identifying emerging cyber threats and vulnerabilities, shaped his focus on using advanced analytics and machine learning to uncover hidden patterns in complex data environments.

Codoxo is a healthcare AI company focused on reducing inefficiencies and unnecessary costs across the healthcare system through its Forensic AI Platform. The platform uses patented algorithms and machine learning to analyze vast volumes of claims data, identifying suspicious behavior, billing anomalies, and emerging fraud patterns earlier than traditional systems. By enabling healthcare payers, government agencies, and pharmacy benefit managers to intervene before or during the claims process, Codoxo shifts the industry toward proactive payment integrity rather than retrospective recovery. Its broader unified cost containment platform integrates data mining, provider education, audit workflows, and case management, helping organizations improve accuracy, reduce overpayments, and streamline operations while addressing an estimated hundreds of billions of dollars lost annually to fraud, waste, and abuse.

You founded Codoxo after researching healthcare fraud detection during your PhD at Georgia Tech. What first convinced you that AI could fundamentally change how fraud, waste, and abuse are detected in healthcare systems?

I became convinced not by a single moment, but by confronting how badly the existing approach was failing. Healthcare fraud in the U.S. represents somewhere around $330 billion lost annually. That’s more than every other form of insurance fraud in the country combined, and yet the dominant detection methods could only catch what they’d already been trained to look for. Even as AI entered the picture, most approaches were reactive, pattern-matching against known fraud schemes rather than surfacing the unknown ones. The moment that crystallized it for me was realizing that fraud is not a static problem. Bad actors adapt. They learn what triggers a flag and route around it. A system built on fixed rules is, by definition, always behind.

What AI offers is the ability to surface patterns that nobody thought to program in advance. During my dissertation work at Georgia Tech, I was building models that could look across a provider’s full claims history, identify behavioral anomalies, and connect signals that a human analyst or a rules engine would never link together. The JASON advisory group, which advises the U.S. government on science and technology, recognized that work as addressing real structural gaps in how health data was being used for payment integrity. That recognition told me the problem was serious enough to build a company around.

The core conviction that drove me then is the same one driving what we’re building at Codoxo now: healthcare claims data contains the signal you need to catch fraud, but you can only extract it with AI that can look at the full picture, quickly and accurately, rather than checking boxes.

Healthcare fraud has long been a multibillion-dollar problem, but generative AI appears to be accelerating it dramatically. How has the rise of tools capable of producing convincing clinical documentation and diagnostic images changed the threat landscape for health insurers and payment integrity teams?

It has changed it in a fundamental way, and the industry hasn’t fully absorbed how significant that shift is yet. The old model of documentation fraud required manual effort. A bad actor committing fraud had to fabricate notes one at a time, alter images individually, and create records that were plausible enough to survive review. That friction created a natural ceiling on the scale of any scheme.

Generative AI removed that ceiling. Today, anyone can ask a large language model to produce 50 therapy session notes for anxiety treatment and receive them in under five minutes. Those notes will use appropriate clinical terminology, follow a plausible narrative structure, and look internally consistent. Most fraud detection systems were never designed to evaluate whether a document is authentic or synthetic. They were designed to check whether billing codes were applied correctly, flag known patterns, and match against existing fraud signatures. So the synthetic documentation passes right through, even past some systems that claim an AI component.

We’ve also seen this with diagnostic imaging. A single legitimate X-ray can be used as the seed for dozens of AI-generated variations, each submitted under a different fabricated patient. A system with no image comparison capability sees 50 unique-looking cases, but the reality is one real scan and 49 synthetic duplicates. The threat landscape has shifted from isolated bad actors to people who can run scalable, repeatable schemes with almost no technical expertise required.

Many traditional fraud detection systems rely on rules-based models and manual review. Why are these approaches increasingly ineffective when dealing with AI-generated medical records or manipulated diagnostic images?

Most fraud detection approaches, whether rules-based or earlier-generation AI, operate on a fundamentally flawed assumption right now: that all documentation entering the system was created by a human following normal clinical processes. Once that assumption breaks down, the entire detection approach breaks down with it.

A rules engine can flag an impossible code combination, a provider billing for more hours than exist in a day, or a procedure performed on a deceased patient. These are all real and useful catches. But rules-based logic cannot look at a progress note and determine whether it was written by a clinician who actually saw a patient or generated by an AI model that has never practiced medicine. The two outputs can be structurally identical.

Manual review has the same ceiling. Studies show that only about 34% of people can identify a deepfake even when they’ve been told one is present and are actively looking for it. An SIU investigator reviewing a pile of progress notes has no specialized forensic training for detecting synthetic text, no image comparison tools for spotting cloned scans, and not enough hours in the day to conduct that level of scrutiny on every claim. The volume problem alone makes comprehensive manual review impossible, and that was true before generative AI started accelerating the volume and sophistication of fraudulent documentation.

There’s also an emerging threat that I think is underappreciated: what researchers call trojan deepfake engines. These are virus-like agents designed specifically to neutralize detection software, working through tactics like image analysis derailment or malicious prompt re-engineering. So the adversarial dynamic isn’t just fraudsters generating better fakes. In some cases, they’re actively trying to break the tools being built to catch them. That arms-race dynamic is part of why static detection approaches, whether rules-based or a fixed AI model, will always fall behind. The defense has to be as adaptive as the offense.

Codoxo recently launched Deepfake Detection to address this emerging risk. At a high level, how does the technology analyze medical documentation and images to determine whether content may have been generated or manipulated by AI?

The core design principle is that we built Deepfake Detection specifically for healthcare documentation rather than retrofitting a general-purpose AI detection tool into a clinical context. That means the models are trained on healthcare-specific signals, not adapted from tools built for other industries or other use cases. That distinction matters because the signals that indicate synthetic content in a medical record or a diagnostic image are different from the signals relevant in other domains.

At a high level, the system analyzes medical documentation and images alongside the full claim context, and it does so in seconds. When an SIU investigator uploads suspect documentation, the AI runs analysis across multiple dimensions simultaneously. It looks for indicators of synthetic or manipulated content, checks for cloning and duplication patterns across the payer’s broader claims history, and evaluates behavioral consistency between the documentation and the provider’s historical patterns.

One thing worth noting is the breadth of formats the system can work across. It handles text documents in PDF, Word, and XML, spreadsheets, medical images, and even handwritten notes. That matters in practice because fraudulent documentation doesn’t arrive in a single tidy format, and a detection system that only covers some of what SIU teams receive leaves gaps that sophisticated fraudsters will eventually find.

From all of that analysis, the system generates a risk score on a 0 to 100 scale with detailed explanations attached, so the investigator understands exactly which signals drove the score. The goal at every step is to produce actionable outputs rather than just alerts, and to do it faster and with greater accuracy than a more generalized system. Speed matters because the only intervention point that meaningfully changes fraud economics is before payment is made.

Your platform highlights capabilities such as cloning detection, partial AI-generation identification, and behavioral cross-referencing against claims history. Could you walk us through how these signals combine to produce a meaningful risk score for investigators?

Each of those capabilities targets a different fraud pattern, and the risk score reflects how they interact in a specific case.

Cloning and duplication detection address the scenario in which a single authentic record is replicated across multiple fabricated patients. What makes this hard to catch without AI is that the variations can be subtle enough that no single document looks suspicious in isolation. The pattern only becomes visible when you’re comparing across the full claims population. Our system can surface that a set of records that appear unique on the surface are actually derivatives of a common source.

Partial AI-generation detection is important because sophisticated fraudsters aren’t always fabricating entire records from scratch. A more common and harder-to-catch pattern is blending, which means taking a legitimate patient record and using AI to add fabricated additional services or procedures. The authentic sections make the document look credible, but the added sections represent claims for care that was never delivered. Our system is specifically tuned to find these instances.

Behavioral cross-referencing connects the documentation being reviewed to the provider’s full claims history. If documentation presents clinical narratives that are inconsistent with how that provider has historically documented similar cases, or if the volume and pattern of supporting records suddenly deviates from baseline, those inconsistencies are signals. On their own, none of them is conclusive.  Together, weighted and explained in the risk score output, they give investigators a meaningful starting point that would take hours or days to develop manually. What makes this different from a system that simply flags AI-generated content is the combination of signals. Content analysis alone can miss blended documents. Behavioral analysis alone can miss a first-time fraudster with no prior pattern. It’s the intersection of all three layers simultaneously, processed in seconds, that catches what other AI systems can’t.

From your perspective, what are some of the most concerning deepfake-enabled healthcare fraud schemes that insurers and regulators should be watching for in the next few years?

The scenarios that worry me most are those that combine scale with plausibility in ways that are hard to track, even with strong detection.

Behavioral health is a real vulnerability. The documentation for therapy services is largely narrative, including session notes, treatment plans, and progress summaries. There are no lab values to cross-reference, no imaging to examine. A fraudulent provider with access to a generalist language model can generate clinically plausible behavioral health documentation at extraordinary volume, and the only practical way to detect it is to use AI that can evaluate whether the documentation shows the linguistic and structural signatures of synthetic generation.

Diagnostic imaging fraud is the other area I watch closely. Free and accessible AI tools can now generate realistic medical imaging variations from a single seed image. As those tools improve, the synthetic outputs will become harder to distinguish from authentic scans without purpose-built detection. Payers whose workflows have no image forensics capability are, at this point, operating on trust that the images they receive are real.

There’s also an emerging concern around identity and credentialing fraud, where AI-generated documentation supports fraudulent provider enrollment or prior authorization for services that were never medically necessary. These schemes are harder to detect because the fraud is embedded in the intake process rather than the claims themselves, and by the time it surfaces in billing data, the damage is already done.

Healthcare claims often involve large volumes of documentation and supporting evidence. How does an AI system evaluate that information quickly enough to stop fraudulent claims before payment is made?

Speed is actually a core design requirement, not a nice-to-have. The only way Deepfake Detection is useful in practice is if it operates at the speed of the claims pipeline. If analysis takes hours or requires a human to initiate a review queue, you’ve already missed the prepay window, and you’re back to doing recovery work after money has left the system.

Our system is designed to complete the analysis in seconds. When documentation is submitted for review, the AI is running its assessment in parallel rather than sequentially. Synthetic content analysis, duplication checks, and behavioral cross-referencing are happening simultaneously rather than in a chain.  The output is a risk score with detailed explanations, so the investigator doesn’t have to interpret raw signals. They get a prioritized, actionable result. The parallel architecture is part of what allows us to do this faster and with greater accuracy. Running all three signal layers simultaneously means the risk score reflects the full picture of a case, not just the first flag that surfaced.

The broader point here is that the shift we’re pushing for across all of Codoxo’s work, what we call Point Zero, is moving payment integrity intervention as far upstream as possible. Catching a fraudulent claim before payment is dramatically more efficient than recovering an overpayment after the fact. Recovery is expensive, slow, and often incomplete. Prevention at the documentation and evidence validation stage changes the economics of the entire problem.

Fraud detection tools must be explainable to investigators, auditors, and regulators. How do you ensure that AI-generated risk scores can be understood and trusted by Special Investigation Units and payment integrity teams?

Explainability isn’t optional in this domain. If an SIU investigator is going to act on a risk score, whether that means holding a claim, opening a case, or building a referral for prosecution, they need to be able to articulate what the system found and why. A black-box output that says “high risk” is not a useful tool in a workflow that has legal and regulatory accountability attached to it.

Every risk score our system generates includes specific fraud indicators, which can include the signals that drove the score, the patterns identified, and the inconsistencies that were surfaced. The investigator can follow the reasoning from the score back to the evidence. That level of specificity is only possible because the underlying detection is purpose-built for healthcare documentation.

We also built in custom prompting capability, which lets investigators tailor the analysis for specific investigation scenarios and unique fraud patterns. That’s important for explainability in practice because it means the system isn’t running a one-size-fits-all analysis and asking investigators to interpret generic outputs. They can shape the inquiry based on what they’re actually looking for in a given case, which makes the results moreectly useful and easier to explain to auditors or in legal proceedings.

On the regulatory side, OIG, CMS, and state agencies are increasing scrutiny on how organizations use AI in fraud prevention. Being able to demonstrate that your detection methodology is interpretable and auditable is not just good practice, it’s a component of responsible deployment that reduces compliance risk.

As generative AI continues to improve, fraudsters will likely become more sophisticated. How does Codoxo design its models to continuously adapt as new forms of synthetic medical documentation emerge?

The challenge is that fraud is adversarial in nature. As detection improves, the tactics on the other side evolve. Any system trained once and deployed without updating will degrade over time, as fraudsters learn what triggers it and adjust. That’s the same fundamental problem that makes any static detection approach inadequate, whether it’s rules-based or an AI model that isn’t designed to update. The sophistication of the tool doesn’t matter if the underlying architecture can’t keep pace with the threat.

Our approach is to treat detection as a continuously updated capability rather than a fixed product. As new fraud patterns surface in the system and as AI generation techniques evolve, those patterns feed back into model improvement. The system is designed to get better at identifying the emerging threat, not just the threats that existed at deployment. This is important given the pace at which generative AI is advancing and the speed at which fraudsters are experimenting with new approaches, including some adversarial techniques like trojan deepfake engines that are designed to undermine detection.

This will be an ongoing contest. There is no final, solved state. The commitment we’ve made is to keep detection capabilities current with the threat, and our agentic architecture is what makes that possible at scale.

Looking ahead, do you believe deepfake detection will become a standard component of healthcare infrastructure, similar to how anti-money-laundering systems operate in finance, or will the industry need entirely new approaches to trust and verification in medical data?

I do think Deepfake Detection will become standard infrastructure, and the timeline on that is shorter than most people in the industry expect. Before AML frameworks became standard in financial services, the industry also relied heavily on rules-based detection and manual review. The shift happened when the threat reached a scale that made reactive detection clearly inadequate, and when the regulatory environment codified the expectation that financial institutions would have systematic, continuously updated controls in place. Healthcare is approaching a similar inflection point.

What’s already happening is that the payers who are deploying deepfake detection now are doing it because the threat is real and present, not because regulation requires it. As those early deployments generate evidence of the losses being prevented, and as AI-generated fraud schemes become more visible in enforcement actions and public reporting, the expectation will broaden across the industry.

Looking ahead, as generative AI continues to improve, the industry may need to rethink how documentation authenticity is established at the point of creation rather than validated after the fact. That could mean cryptographic attestation of clinical records at the EHR level, provider identity verification integrated into the documentation workflow, or other mechanisms that make the origin of a document traceable in ways it currently is not. Detection at the claims level is a necessary response to the current threat. But the durable solution may require building verification deeper into the infrastructure of how medical records are created and transmitted.

Thank you for the great interview, readers who wish to learn more should visit Codoxo.

Antoine is a visionary leader and founding partner of Unite.AI, driven by an unwavering passion for shaping and promoting the future of AI and robotics. A serial entrepreneur, he believes that AI will be as disruptive to society as electricity, and is often caught raving about the potential of disruptive technologies and AGI.

As a futurist, he is dedicated to exploring how these innovations will shape our world. In addition, he is the founder of Securities.io, a platform focused on investing in cutting-edge technologies that are redefining the future and reshaping entire sectors.