explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR
  • The four models, and what each is actually for
  • Darwin's numbers, against the field
  • Why Darwin isn't shipping broadly yet
  • What people are asking
  • Related reading on explainx.ai
← Back to blog

explainx / blog

OpenEvidence Darwin Hits 100% on MedQA — First AI to Do It

Medical AI, OpenEvidence, Model Releases, Benchmarks, Healthcare

OpenEvidence launched a four-model medical AI family on September 3, 2026. Darwin, its research-preview flagship, scored a perfect 100% on MedQA — the first AI in history to do so — and leads general models on every clinical benchmark tested.

Sep 4, 2026·7 min read·Yash Thakker
add explainx.ai
go deep
OpenEvidence Darwin Hits 100% on MedQA — First AI to Do It

OpenEvidence, the AI platform already widely used by clinicians for point-of-care medical search, released a four-model family on September 3, 2026 — and its flagship, Darwin, scored a perfect 100.0% on MedQA, which OpenEvidence describes as the first time any AI has done so. The launch is notable both for the benchmark number and for what OpenEvidence chose not to do with it: Darwin isn't shipping broadly. It's in application-only research preview while its safeguards get validated with partners, even as the three less powerful models in the family go live immediately.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.


TL;DR

table · 2 cols
QuestionDirect answer
What launched?A four-model OpenEvidence Model Family: Osler, Sackett, Snow (live now), and Darwin (research preview)
What's Darwin's headline result?100.0% on MedQA — the first AI, per OpenEvidence, to score perfectly
How does it compare to general frontier models?Ahead of Claude Fable 5, GPT-5.6 Sol, and Gemini 3.7 Flash on all four benchmarks OpenEvidence published
Can I use Darwin today?Only by application, in research preview — not a self-serve rollout
Can I use the other three models today?Yes — live now on web, iOS, and Android for verified clinicians
What differs between Osler, Sackett, and Snow?Primarily how long each model thinks and how deep it searches, not clinical-accuracy standard
Is the benchmark comparison independent?No — it's OpenEvidence's own published comparison, pending independent replication

The four models, and what each is actually for

OpenEvidence frames the family by clinical moment rather than raw capability tier — a deliberate choice, since it signals these are meant as fit-for-context tools rather than a single model dialed up or down:

table · 3 cols
ModelFramed forStatus
Osler"The hallway" — quick, in-the-moment lookupsLive now
Sackett"The consult" — deeper single-patient reasoningLive now
Snow"The tumor board" — multi-specialist reviewLive now
DarwinFlagship, most capableResearch preview, application-only

OpenEvidence's own description of the split is precise: "Every model in the family is held to the same standard of clinical accuracy. What varies is time: how long a model thinks, and how deep it searches." That's a meaningfully different design philosophy than most consumer AI product tiers, which typically trade off capability, not just latency, between a fast and slow option.

Darwin's numbers, against the field

OpenEvidence published a direct head-to-head comparison of Darwin against three general-purpose frontier models across four medical benchmarks:

table · 5 cols
BenchmarkOpenEvidence DarwinClaude Fable 5GPT-5.6 SolGemini 3.7 Flash
MedQA (USMLE-style, accuracy)100.0%99.7%99.1%99.2%
MedXpertQA (specialty boards, accuracy)72.8%64.3%57.6%65.1%
HealthBench Pro (open-ended rubrics)82.7%70.6%66.9%59.0%
NOHARM (harm-weighted F1)87.2%72.5%74.0%71.5%

Two things stand out in that table. First, the MedQA gap between Darwin and the general models is small in absolute terms — all four models are already clustered near the top of that benchmark's ceiling, which is a familiar pattern as a specific benchmark saturates across the field. Second, the gap widens substantially on the harder, more open-ended benchmarks — MedXpertQA, HealthBench Pro, and NOHARM — where Darwin's lead over the closest general model runs 8-24 percentage points depending on the test. That pattern is consistent with what you'd expect from a domain-specialized model versus general-purpose frontier models on a domain's hardest, least benchmark-saturated tasks.

Why Darwin isn't shipping broadly yet

OpenEvidence's own framing for restricting Darwin is unusually direct for a launch announcement: "Capability at that level cuts both ways in medicine, so Darwin is in research preview, by application only, while its safeguards are validated with partners like RareDiseases."

That's a notable departure from the more common pattern of shipping the most capable model first and adding restrictions reactively. Restricting the most capable model in a medical-AI family, specifically because of its capability, while shipping the less powerful models immediately, is a legible statement about where OpenEvidence sees the actual risk surface — not in whether the model gets clinical facts right (Darwin's benchmark scores are the strongest in the family), but in what happens when a highly capable model's output gets trusted without the same guardrails as the more constrained models.

What people are asking

Is OpenEvidence's benchmark comparison trustworthy? It's a first-party comparison, published by the company whose model is winning it, against three benchmarks (MedXpertQA, HealthBench Pro, NOHARM) that are less universally standardized than MedQA. That doesn't make the numbers wrong, but it does mean independent replication — by researchers or a third-party evaluator without a commercial stake — hasn't happened yet as of this post, and is worth watching for before treating the comparison as settled.

Are Claude Fable 5, GPT-5.6 Sol, and Gemini 3.7 Flash "bad" at medicine? No — all three score above 99% on MedQA in OpenEvidence's own comparison, which is itself a strong result for general-purpose models never specifically trained on OpenEvidence's clinical dataset or fine-tuning pipeline. The comparison illustrates the value of domain specialization on the harder, more open-ended benchmarks specifically, not a broad medical-AI capability gap.

Does a high MedQA score mean an AI is safe to use for real clinical decisions? No, and this is exactly the distinction OpenEvidence's own Darwin restriction seems to acknowledge. MedQA measures knowledge recall against board-style questions; NOHARM's harm-weighted F1 and HealthBench Pro's physician-written open-ended rubrics are closer proxies for real clinical judgment, and even there, a benchmark score is not equivalent to validated safety in live clinical use — which is precisely why Darwin remains gated behind partner validation rather than shipping on benchmark strength alone.

How is this different from a general chatbot answering medical questions? OpenEvidence is built specifically for point-of-care clinical use by verified clinicians, with models tuned and evaluated against medical-specific benchmarks and rubrics, rather than a general-purpose assistant that happens to also answer medical questions. The four-model family structure — differentiated by clinical context and thinking depth rather than a single one-size-fits-all model — reflects that specialization.


Related reading on explainx.ai

  • AI Benchmarks: Complete Guide — background on how benchmarks like MedQA and HealthBench Pro are constructed and their limits
  • Claude Protein Design and Analytical Chemistry — Anthropic's own domain-specialized scientific AI work, for comparison
  • Claude Team Plan for Scientists: 10,000 Seats — general frontier-model adoption in scientific and clinical research settings
  • AI Drug Discovery: Clinical Evidence and Benchmarks — a broader look at how AI clinical-benchmark claims hold up against independent scrutiny
  • Moderna and Merck: AI-Designed mRNA Cancer Vaccine Phase 3 — another example of specialized medical AI moving from benchmark to real clinical application
  • Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6 Comparison — where the general models cited in OpenEvidence's comparison stand outside medicine
  • What Are AI Agents? Complete Guide — background context on how domain-specific AI products differ from general-purpose assistants

Official source: OpenEvidence Model Family launch announcement

Benchmark figures in this post are OpenEvidence's own published comparison as of the September 3, 2026 launch and have not yet been independently replicated. Darwin's availability is restricted to research-preview, application-only access as of this post's publication.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Sep 4, 2026

GPT-6 Astra Is Live. Here's Every Number That Actually Matters.

After a confused false start — press coverage went live before OpenAI's own page did — GPT-6 Astra shipped on September 3, 2026 to ChatGPT Plus, Pro, Business, and Enterprise, plus the API. It matches Fable 5.1's pricing, leads on security and long-context benchmarks, and trails Fable 5.1 on general intelligence. Here is every number, not just the highlight reel.

Sep 5, 2026

Artificial Analysis Intelligence Index v4.2: What Actually Changed

Artificial Analysis published Intelligence Index v4.2 on September 4, 2026, an interim update ahead of v5: two new evaluations added (AA-Briefcase, GDP.pdf), GPQA Diamond retired as saturated, and private held-out test sets now carry 40% of the total weight. Claude Fable 5.1 leads the index, GPT-6 Astra wins on cost-per-task and token efficiency — and Hacker News raised fair questions about the timing.

Sep 5, 2026

EEBench: The Benchmark That Grades Whether AI Can Design Circuits

On September 4, 2026, the atopile team published EEBench — a benchmark that grades AI-designed electronic circuits by simulating them in SPICE with real manufacturer part tolerances, not just checking whether the design compiles. It hit #1 on Hacker News, and the leaderboard has some surprises.