HealthBench raw score leaderboard
1 ranked models · higher is better · labels show rank and score
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Opus 5 | #1 | Anthropic | 67.1% |
Knowledge · Benchmark profile
Raw score on realistic multi-turn healthcare conversations graded against expert-written rubrics.
Data verified 27 Jul 2026 · Methodology 1.7.0
Visual analysis
Switch between model placement, score distribution and descriptive provider averages. Every view uses the same sourced leaderboard.
1 ranked models · higher is better · labels show rank and score
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Opus 5 | #1 | Anthropic | 67.1% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | Claude Opus 5 claude-opus-5 | Anthropic | closed | Estimated reference | 67.1% |
About HealthBench raw score
Raw score on realistic multi-turn healthcare conversations graded against expert-written rubrics. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
Raw score on realistic multi-turn healthcare conversations graded against expert-written rubrics.
Claude Opus 5 by Anthropic currently leads with 67.1%.
1 model in the LuminaBench cohort have a qualifying score on this benchmark.
No. This benchmark is display-only and does not enter the overall Lumina composite.
Related
2026 · 21 results · Reference
Artificial Analysis MMLU-Pro2026 · 8 results · Reference
Massive Multitask Language Understanding2020 · 16 results · Reference
Graduate-Level Google-Proof Q&A2023 · 140 results · Reference
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines2025 · 38 results · Reference
Massive Multitask Language Understanding Professional2024 · 85 results · Reference