INCLUDE leaderboard
3 ranked models · higher is better
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Opus 4.8 | #1 | Anthropic | 87.6% |
| Qwen3.7-Max | #2 | Alibaba Cloud | 86.2% |
| Qwen3.7-Plus | #3 | Alibaba Cloud | 83% |
Knowledge · Benchmark profile
A multilingual benchmark used in provider tables to measure inclusive language coverage and cross-lingual understanding beyond common high-resource languages.
Data verified 21 Jul 2026 · Methodology 1.6.0
Benchmark score on INCLUDE
Claude Opus 4.8 leads at 87.6%, followed by Qwen3.7-Max (86.2%) and Qwen3.7-Plus (83%).
Visual analysis
Switch between model placement, score distribution and descriptive provider averages. Every view uses the same sourced leaderboard.
3 ranked models · higher is better
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Opus 4.8 | #1 | Anthropic | 87.6% |
| Qwen3.7-Max | #2 | Alibaba Cloud | 86.2% |
| Qwen3.7-Plus | #3 | Alibaba Cloud | 83% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.8 claude-opus-4-8 | Anthropic | closed | Reference only | 87.6% |
| #2 | Qwen3.7-Max qwen3.7-max | Alibaba Cloud | — | Reference only | 86.2% |
| #3 | Qwen3.7-Plus qwen3.7-plus | Alibaba Cloud | — | Reference only | 83% |
The top of this snapshot is led by Claude Opus 4.8 at 87.6%; third place is 4.6 points behind. The top-3 spread is 4.6 points.
About INCLUDE
A multilingual benchmark used in provider tables to measure inclusive language coverage and cross-lingual understanding beyond common high-resource languages. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
A multilingual benchmark used in provider tables to measure inclusive language coverage and cross-lingual understanding beyond common high-resource languages.
Claude Opus 4.8 by Anthropic currently leads with 87.6%.
3 models in the LuminaBench cohort have a qualifying score on this benchmark.
No. This benchmark is display-only and does not enter the overall Lumina composite.
Related
2026 · 10 results · Reference
Artificial Analysis MMLU-Pro2026 · 4 results · Reference
Massive Multitask Language Understanding2020 · 8 results · Reference
Graduate-Level Google-Proof Q&A2023 · 70 results · Reference
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines2025 · 19 results · Reference
Massive Multitask Language Understanding Professional2024 · 42 results · Reference