Where clinicians tested Meditron
LiGHT is hosted by
Talk to Meditron
Ask a medical question. Meditron runs on EPFL's servers. It is a research model, not a doctor: it can be wrong, and it knows nothing about you beyond what you type.
Where clinicians tested Meditron
LiGHT is hosted by
Ask a medical question. Meditron runs on EPFL's servers. It is a research model, not a doctor: it can be wrong, and it knows nothing about you beyond what you type.
Apertus is the large language model built in Switzerland by EPFL, ETH Zurich and CSCS. Its weights, training data and training code are all public. That is why we build Meditron on it: if the base model is open all the way down, so is Meditron.
In blind side-by-side comparisons, the AI judges prefer Meditron's answer over the starting model's in 90% of 516 clinician-written cases.
| Model | Open base model | Open medical data | Open data pipeline | Open training code |
|---|---|---|---|---|
| MedGemma-27B (Google) | No | No | No | No |
| Meditron-70B (2023) | No | No | No | No |
| Meditron on Apertus | Yes | Yes | Yes | Yes |
Apertus on Hugging FaceMeditron models
Commercial medical assistants don't show what they learned from. Meditron's training set is on Hugging Face, row by row. Here is one row.
A 12-year-old boy, weight 38 kg, is being considered for fingolimod therapy. What is the recommended starting dose?
Fingolimod (Gilenya) is approved for pediatric patients ≥ 10 years of age, but the dose is adjusted for body weight.
The boy is 12 years old and weighs 38 kg, which places him in the < 40 kg category. Therefore the appropriate starting dose is the lower, weight-adjusted dose of 0.25 mg once daily. The answer is A.
Not perfect: AI reviewers still flag about 1 in 10 sampled rows. You can check them yourself.
Benchmarks made of exam questions miss what clinicians actually ask. On the MOOVE platform, clinicians write the questions they face in their own practice, then read two anonymous answers side by side and vote for the better one.
Clinician votes are the reference we trust most, and the scarcest. AutoMOOVE turns them into an evaluation we can rerun on every new model.
Join the MOOVETry the clinicians' job in AutoMOOVE
HealthBench: 5,000 health conversations, each graded against rubrics written by physicians.
Graded by Gemma-4-31B, so the scores are not comparable to OpenAI's published numbers.
At MOOVE events, clinicians read two anonymous answers to the same question and vote for the better one. Thousands of votes add up to a ranking: the share of matchups each model would win against an average model at that event.
Three events at CHUV, Lausanne, Aug 2024 to Nov 2025: 684 questions from hospital specialists, 1,788 clinician-ranked pairs. Not every pair of models met equally often.
Kenya PHC, Tanzania PHC and Malawi PHC, May to Aug 2026: 605 primary-care questions, 4,235 clinician-ranked pairs, every system against every other. During these events Apertus-70B-MeditronFO was served at temperature 1.0, which cut off many of its answers.
The axis runs from 20% to 80%; bars show the 95% range over 1,000 bootstrap resamples, and the dashed line is 50%. Ranking: Bradley–Terry on clinician votes, ties counting half. An AI judge replaying both events reproduces these rankings (rank correlation 1.00 at CHUV, 0.96 in primary care), which is what lets AutoMOOVE stand in between events.