Interviews
Thomas Clozel, M.D., Co-Founder and CEO of Owkin – Interview Series

Thomas Clozel, M.D. is the co-founder and CEO of Owkin, an AI biotech company revolutionizing the healthcare industry. With a background in clinical research, Clozel has been instrumental in leading Owkin since its inception in 2016. Under his leadership, the company has achieved significant milestones, including raising substantial funding and forming strategic partnerships with major pharmaceutical companies. He is a former Assistant Professor of Clinical Onco-Hematology at Hopital Henri Mondor in Paris and former member of Dr. Ari Melnick’s oncology lab at the Weill Cornell Medical College.
Owkin is an AI biotechnology company developing autonomous AI systems designed to accelerate biomedical research and drug discovery. Founded in 2016 by oncologist Thomas Clozel and machine learning researcher Gilles Wainrib, the company combines advanced AI models with multimodal patient data, biological research, and clinical expertise to help pharmaceutical teams better understand disease, identify therapeutic targets, discover biomarkers, and make more informed R&D decisions. Its flagship platform, K Pro, acts as an AI scientist that can orchestrate specialized models and research tools to analyze complex biological questions, with Owkin ultimately working toward what it calls Biological Artificial SuperIntelligence: a self-improving AI system capable of conducting increasingly large portions of the drug discovery process autonomously.
Owkin has established strategic relationships with major biopharmaceutical companies including Sanofi, Bristol Myers Squibb, and MSD, and has built a global network spanning academic researchers, clinicians, patient data partners, and laboratory infrastructure.
Your career has spanned clinical hematology, oncology research and computational biology. When you co-founded Owkin in 2016, what limitations did you see in conventional medical research that convinced you AI could help scientists understand disease differently?
I was an oncologist and I was seeing patients of 35 with cancer, but nothing in their history explained why they had it, or why they relapsed. And we didn’t have good science to explain it — even with all the research of the last hundred years. So it made me think: biology is not a human-scale problem. It needs reasoning across scales, from molecules to entire organisms, and across many different types of data at once — from tissue biopsies to a patient’s treatment history. For that, we would need AI. And I thought that the only data that could get us closest to that understanding in humans was patient data. So in 2016 I co-founded Owkin with Gilles Wainrib, and we started with federated learning — a method that lets AI models learn from hospital data without that data ever leaving the hospital — to unlock that data and reach those answers.
Generative AI can produce large numbers of plausible biological hypotheses. How should researchers distinguish a genuinely novel discovery from a convincing but ultimately false AI-generated correlation?
You should consider all your hypotheses wrong until you prove otherwise. This is by design. I believe that frontier AI agents — advanced, autonomous AI systems — when given access to the right tools and data, will be better at generating accurate hypotheses than we are currently. And at Owkin, our models rank hypotheses by how likely they are to hold up, based on patterns learned from the data.
But we still need to test them, first in the lab and then through clinical trials, before we can put our full faith in them. At Owkin we do that through what we call our “lab in the loop” — testing every AI prediction on real biological samples before it’s trusted — and then in clinical trials. So we still need to test each hypothesis experimentally, as we always have. But AI helps us use resources better by directing us toward the testing routes most likely to succeed.
K Pro combines pathology, genomics, spatial biology and clinical outcomes. How does bringing these different data types together produce stronger biological evidence than relying on a single dataset or modality?
We’re trying to see the full picture of what has caused a person’s disease. Each of these data types gives you a snapshot of part of it — but no single one can give you the whole, causal explanation.
If you can look at the disease from multiple angles at once, you get closer to the full picture. That’s what we want to do across every scale of biology — for example, linking gene activity to how tissue is physically structured, and linking a patient’s underlying biology (their genes and molecules) to what’s actually observed in the clinic (tissue images, medical records, how they responded to treatment).
That’s what’s so important about AI: it can combine these very different data types and find the underlying patterns that get closer to the whole picture — something that’s extremely hard for a human mind to do on its own.
Owkin describes AI-generated discoveries as hypotheses rather than facts. What stages must a K Pro hypothesis pass through before your team considers it biologically meaningful enough to influence a drug-development decision?
Four stages.
First, computational replication: does the signal hold up when checked against independent patient datasets, and across data types that weren’t used to generate it in the first place? This step is cheap to run and it eliminates the majority of candidates early.
Second, mechanistic scrutiny: is there a coherent, causal explanation — a plausible new piece of biology? And does the evidence actually say what the model claims it says? We interrogate the underlying evidence, not just the conclusion. This is where input from our expert biomedical team is hugely important.
Third, we test it in the lab. We check the prediction using patient-derived models — including organoids (miniature, lab-grown versions of a patient’s tumor) — through drug testing and gene-editing experiments. This is the step where a correlation either becomes a real causal claim, or doesn’t.
Fourth, patients. We’re currently testing OKN4395 — a drug candidate whose underlying biology K Pro helped us understand — in our INVOKE clinical trial. We’re also using that trial to improve our AI: it collects detailed patient data throughout, so we’re not only testing a drug, we’re testing whether the model’s reasoning about the disease mechanism and the right patient population was correct.
Owkin tests selected K Pro predictions in its wet laboratory using approaches such as patient-derived organoids and genetic perturbation. How does this laboratory feedback improve the system, and what happens when experimental results contradict the original prediction?
K Pro predicts new biology, we design the experiment, and we test it using patient-derived lab models — including organoids, co-cultures (growing different cell types together to see how they interact), and models where specific genes have been switched on or off. The results then go back into the system to improve it. Today, parts of that loop are still done manually; fully automating it is our next step.
We’ve already measured the effect of this feedback. In a study we released on ArXiv (a repository for scientific papers) this year, AI agents that received this kind of iterative lab feedback were 166–185% better than random selection at identifying genes that meaningfully change how a cell behaves — the kind of genes worth targeting with a drug.
What metrics should pharmaceutical companies use to evaluate an AI scientist? Should success be measured by predictive accuracy, reproducibility, the number of validated targets, improved clinical-trial decisions or another standard?
We think the benchmark that matters most is reliability. An AI that can occasionally produce a very accurate answer is interesting, but AI systems only become genuinely useful to the industry when they can answer the same question the same way, reliably, no matter how many times you ask.
That’s what shifts these systems from a flashy demo to the foundation of an actual workflow. That’s how we’ve built K Pro — so that the AI reasoning inside it follows the same reproducible, transparent process every time it answers a question.
Scientific reproducibility is essential in drug research. How does K Pro show researchers which datasets, analytical workflows and reasoning steps contributed to a conclusion so that its findings can be independently reviewed?
K Pro documents and displays the reasoning process the AI goes through to answer every query.
Every dataset it uses is identified by source and version. You can see the patient group narrow at each filtering step — for example, how four thousand patients became three hundred. Every tool it calls on is shown. Every “skill” it uses is shown. Every reasoning step is shown.
A “skill” is a scientific workflow that’s been encoded once — the sequence of analyses, the thresholds, the reasoning — which K Pro then applies identically every time it’s relevant. Skills are especially important for reproducibility, because they’re one of the ways K Pro reduces the natural variability of AI systems — the same task run twice should give the same answer. We’re continuing to build on this, because the ability to verify an AI’s work is a key advantage in our industry.
Owkin is developing purpose-built K Pro agents for AstraZeneca and Sanofi. What have these collaborations taught you about integrating AI scientists into established pharmaceutical workflows without weakening human oversight, governance or scientific rigor?
You have to meet pharmaceutical companies where they are. Every company has a different setup, a different history, and a different appetite for AI adoption — and within each company, different teams have different needs.
Our experience working with pharma companies over almost the last decade is that they will always prioritize oversight and rigor, and you won’t even get in the door without meeting their governance standards.
So we’ve built K Pro to be interpretable — meaning you can see how it reasons, what tools it selects, and what data it calls on — flexible enough to work within a company’s existing systems without requiring them to change infrastructure, and able to access proprietary data exactly where it’s already stored, without ever moving or duplicating it.
That goes a long way toward building the trust needed for pharma adoption.
Multimodal patient data can contain institutional biases, inconsistent annotations and underrepresented populations. How does Owkin prevent these limitations from becoming misleading biological conclusions or narrowing the patients who may benefit from a discovery?
All data contains some form of bias, whether from the lab it came from or the patients it represents. Our approach is to understand our datasets as fully as possible, so we can identify whether bias exists and how to account for it.
We’re able to do this because we have strong relationships with the researchers who originally generated each dataset, and they help us examine it in depth.
We’re confident that if a relevant dataset exists, we can find it. But if there’s a gap — for example, a patient population isn’t well represented in existing data — we’re able to generate new data to rigorous standards ourselves, working with academic partners, as we did with our MOSAIC dataset.
You have described Owkin’s long-term objective as Biological Artificial Superintelligence. What capabilities would distinguish that from today’s AI research tools, and what evidence would convince you that an AI system genuinely understands biology rather than simply identifying increasingly sophisticated patterns?
Biological Artificial Superintelligence (BASI) — our term for an AI system capable of independently proposing genuinely new, correct biology — arrives when an AI system starts reliably suggesting new disease biology that human research hasn’t uncovered, and that later proves to be correct.
We’re building toward that with a system that can:
- Work autonomously, refining its own hypotheses without needing human intervention at every step, including designing and running experiments in automated labs
- Reason about cause and effect in biological systems, not just spot correlations
As for whether the AI genuinely “understands” biology, that’s an interesting philosophical question. At what point do a model’s internal representations become complex enough to count as real understanding, rather than just very sophisticated pattern-matching? For now, I’ll settle for it being able to reliably suggest new biology that turns out to be true.
Thank you for the great interview, readers who wish to learn more should visit Owkin.












