OpenEvidence, the AI platform already widely used by clinicians for point-of-care medical search, released a four-model family on September 3, 2026 — and its flagship, Darwin, scored a perfect 100.0% on MedQA, which OpenEvidence describes as the first time any AI has done so. The launch is notable both for the benchmark number and for what OpenEvidence chose not to do with it: Darwin isn't shipping broadly. It's in application-only research preview while its safeguards get validated with partners, even as the three less powerful models in the family go live immediately.
TL;DR
| Question | Direct answer |
|---|---|
| What launched? | A four-model OpenEvidence Model Family: Osler, Sackett, Snow (live now), and Darwin (research preview) |
| What's Darwin's headline result? | 100.0% on MedQA — the first AI, per OpenEvidence, to score perfectly |
| How does it compare to general frontier models? | Ahead of Claude Fable 5, GPT-5.6 Sol, and Gemini 3.7 Flash on all four benchmarks OpenEvidence published |
| Can I use Darwin today? | Only by application, in research preview — not a self-serve rollout |
| Can I use the other three models today? | Yes — live now on web, iOS, and Android for verified clinicians |
| What differs between Osler, Sackett, and Snow? | Primarily how long each model thinks and how deep it searches, not clinical-accuracy standard |
| Is the benchmark comparison independent? | No — it's OpenEvidence's own published comparison, pending independent replication |
The four models, and what each is actually for
OpenEvidence frames the family by clinical moment rather than raw capability tier — a deliberate choice, since it signals these are meant as fit-for-context tools rather than a single model dialed up or down:
| Model | Framed for | Status |
|---|---|---|
| Osler | "The hallway" — quick, in-the-moment lookups | Live now |
| Sackett | "The consult" — deeper single-patient reasoning | Live now |
| Snow | "The tumor board" — multi-specialist review | Live now |
| Darwin | Flagship, most capable | Research preview, application-only |
OpenEvidence's own description of the split is precise: "Every model in the family is held to the same standard of clinical accuracy. What varies is time: how long a model thinks, and how deep it searches." That's a meaningfully different design philosophy than most consumer AI product tiers, which typically trade off capability, not just latency, between a fast and slow option.
Darwin's numbers, against the field
OpenEvidence published a direct head-to-head comparison of Darwin against three general-purpose frontier models across four medical benchmarks:
| Benchmark | OpenEvidence Darwin | Claude Fable 5 | GPT-5.6 Sol | Gemini 3.7 Flash |
|---|---|---|---|---|
| MedQA (USMLE-style, accuracy) | 100.0% | 99.7% | 99.1% | 99.2% |
| MedXpertQA (specialty boards, accuracy) | 72.8% | 64.3% | 57.6% | 65.1% |
| HealthBench Pro (open-ended rubrics) | 82.7% | 70.6% | 66.9% | 59.0% |
| NOHARM (harm-weighted F1) | 87.2% | 72.5% | 74.0% | 71.5% |
Two things stand out in that table. First, the MedQA gap between Darwin and the general models is small in absolute terms — all four models are already clustered near the top of that benchmark's ceiling, which is a familiar pattern as a specific benchmark saturates across the field. Second, the gap widens substantially on the harder, more open-ended benchmarks — MedXpertQA, HealthBench Pro, and NOHARM — where Darwin's lead over the closest general model runs 8-24 percentage points depending on the test. That pattern is consistent with what you'd expect from a domain-specialized model versus general-purpose frontier models on a domain's hardest, least benchmark-saturated tasks.
Why Darwin isn't shipping broadly yet
OpenEvidence's own framing for restricting Darwin is unusually direct for a launch announcement: "Capability at that level cuts both ways in medicine, so Darwin is in research preview, by application only, while its safeguards are validated with partners like RareDiseases."
That's a notable departure from the more common pattern of shipping the most capable model first and adding restrictions reactively. Restricting the most capable model in a medical-AI family, specifically because of its capability, while shipping the less powerful models immediately, is a legible statement about where OpenEvidence sees the actual risk surface — not in whether the model gets clinical facts right (Darwin's benchmark scores are the strongest in the family), but in what happens when a highly capable model's output gets trusted without the same guardrails as the more constrained models.
What people are asking
Is OpenEvidence's benchmark comparison trustworthy? It's a first-party comparison, published by the company whose model is winning it, against three benchmarks (MedXpertQA, HealthBench Pro, NOHARM) that are less universally standardized than MedQA. That doesn't make the numbers wrong, but it does mean independent replication — by researchers or a third-party evaluator without a commercial stake — hasn't happened yet as of this post, and is worth watching for before treating the comparison as settled.
Are Claude Fable 5, GPT-5.6 Sol, and Gemini 3.7 Flash "bad" at medicine? No — all three score above 99% on MedQA in OpenEvidence's own comparison, which is itself a strong result for general-purpose models never specifically trained on OpenEvidence's clinical dataset or fine-tuning pipeline. The comparison illustrates the value of domain specialization on the harder, more open-ended benchmarks specifically, not a broad medical-AI capability gap.
Does a high MedQA score mean an AI is safe to use for real clinical decisions? No, and this is exactly the distinction OpenEvidence's own Darwin restriction seems to acknowledge. MedQA measures knowledge recall against board-style questions; NOHARM's harm-weighted F1 and HealthBench Pro's physician-written open-ended rubrics are closer proxies for real clinical judgment, and even there, a benchmark score is not equivalent to validated safety in live clinical use — which is precisely why Darwin remains gated behind partner validation rather than shipping on benchmark strength alone.
How is this different from a general chatbot answering medical questions? OpenEvidence is built specifically for point-of-care clinical use by verified clinicians, with models tuned and evaluated against medical-specific benchmarks and rubrics, rather than a general-purpose assistant that happens to also answer medical questions. The four-model family structure — differentiated by clinical context and thinking depth rather than a single one-size-fits-all model — reflects that specialization.
Related reading on explainx.ai
- AI Benchmarks: Complete Guide — background on how benchmarks like MedQA and HealthBench Pro are constructed and their limits
- Claude Protein Design and Analytical Chemistry — Anthropic's own domain-specialized scientific AI work, for comparison
- Claude Team Plan for Scientists: 10,000 Seats — general frontier-model adoption in scientific and clinical research settings
- AI Drug Discovery: Clinical Evidence and Benchmarks — a broader look at how AI clinical-benchmark claims hold up against independent scrutiny
- Moderna and Merck: AI-Designed mRNA Cancer Vaccine Phase 3 — another example of specialized medical AI moving from benchmark to real clinical application
- Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6 Comparison — where the general models cited in OpenEvidence's comparison stand outside medicine
- What Are AI Agents? Complete Guide — background context on how domain-specific AI products differ from general-purpose assistants
Official source: OpenEvidence Model Family launch announcement
Benchmark figures in this post are OpenEvidence's own published comparison as of the September 3, 2026 launch and have not yet been independently replicated. Darwin's availability is restricted to research-preview, application-only access as of this post's publication.
