AI Models & Platforms
OpenEvidence Launches Medical AI Model Family With Darwin Preview

OpenEvidence released a new family of medical AI search models on September 3, 2026, making three production models available free to verified clinicians and opening a fourth, which it describes as its most advanced, to researchers by application only.
The three production models are named for figures in the history of medicine. Osler, which the company describes as its fastest model at roughly five seconds per answer, is the successor to the model that previously powered OpenEvidence answers and becomes the platform’s default. Sackett is positioned as a deeper search model that takes roughly 30 seconds per answer for questions that turn on the weight of the evidence. Snow, the successor to the OpenEvidence Deep Consult feature, is the deepest production model at roughly five minutes per answer, running a full investigation of medical literature before producing a report.
The models are rolling out to all OpenEvidence users on openevidence.com and the company’s iOS and Android apps, with clinicians able to select among them through a model selector in the interface. OpenEvidence said every model in the family is held to the same standard of clinical accuracy, with the difference between them being how long a model thinks and how deep it searches.
The namesakes are William Osler, who moved medical teaching from the lecture hall to the bedside; David Sackett, remembered as the father of evidence-based medicine; and John Snow, who traced the 1854 Broad Street cholera outbreak to a single water pump and helped found modern epidemiology.
Darwin in Research Preview
The fourth model, Darwin, is in research preview and is not being released generally. In the announcement, OpenEvidence describes Darwin as the most advanced medical AI model in the world and says it is the first AI model to achieve a perfect score on MedQA, which the company calls the leading fully independent benchmark of medical AI. The company also reports Darwin as state of the art on MedXpertQA at 72.8 percent, HealthBench Professional at 82.7 percent, and NOHARM at 87.2 percent, ahead of the next-best models, which it names as Claude Fable 5 and Gemini 3.7.
Access is by application only. OpenEvidence said its reasoning relates to dual-use risk: a model that can reason at the frontier of virology, immunology, and human genetics could, in the wrong hands, accelerate work that is tightly governed, such as bioweapons-relevant research and germline editing outside mainstream scientific oversight. Current access covers institutional partners such as the National Organization for Rare Disorders, which brings Darwin its hardest cases, research collaborators, and accredited AI researchers at academic institutions benchmarking the accuracy and safety of AI in clinical medicine. The company said Darwin’s capabilities will flow into Osler, Sackett, and Snow as its safeguards are validated with these partners.
Benchmark Methodology
In a companion technical post, OpenEvidence detailed the evaluation setup. For MedQA, the company started from physician re-annotations of the benchmark’s 1,273-question four-option test split and applied exclusion criteria for missing information, ambiguity, and label errors, followed by a second review pass in August 2026 by three OpenEvidence physicians covering every question any evaluated model answered incorrectly. That produced a final evaluation set of 660 questions, which Darwin answered without error. The company is releasing its annotations and Darwin’s complete responses for all four benchmarks.
On MedXpertQA, Darwin was evaluated on the full 2,450-question Text split, excluding the multimodal subset. For HealthBench Professional, the company used the publicly available test set with multi-turn cases excluded, leaving 410 of the 525 tasks, and graded responses with the LLM-as-a-judge configuration published by Anthropic, using Claude Opus 4.8 with a 32,000-token maximum thinking budget. NOHARM, a benchmark scoring the frequency and severity of harmful recommendations in real consultation cases, was graded with its official open-source package on the 30-case open subset.
Baselines were run as claude-fable-5 with adaptive thinking, gpt-5.6-sol at default reasoning effort, and gemini-3.7-flash at default reasoning effort, using each provider’s API defaults with no tools or customized system prompts. HealthBench Professional scores were averaged over at least five independent trials per model, while other datasets were scored over a single pass, with mean scores reported throughout.
According to the post, Darwin’s MedXpertQA score sits 7.7 points above the strongest baseline, its HealthBench Professional score 12.1 points above the next-best model, and its NOHARM severity-weighted F1 reached 0.872 against 0.740 for the closest baseline. The company acknowledged that Darwin’s unweighted precision on NOHARM is lower than several baselines, explaining that the metric counts every recommendation absent from the rubric as a false positive and that Darwin is tuned to give physicians the full set of relevant options. It reported Darwin’s severity-weighted precision as 0.9.
Availability and Next Steps
Osler, Sackett, and Snow are available with unlimited usage to all verified clinicians at no cost. Darwin is available to researchers through an application process on the company’s site.
OpenEvidence said it will continue to improve, expand, and increase access to the model family over time, including models built for specialty practice beginning with oncology, radiology, and clinical genetics.












