AGIEval leaderboard
3 ranked models · higher is better
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| DeepSeek V4 Pro Base | #1 | DeepSeek | 83.1% |
| DeepSeek V4 Flash Base | #2 | DeepSeek | 82.6% |
| Soofi S 30B-A3B | #3 | Soofi Project | 66.9% |
Knowledge · Benchmark profile
A human-centric exam benchmark for general knowledge and reasoning reported in DeepSeek-V4 base-model evaluations.
Data verified 21 Jul 2026 · Methodology 1.6.0
Benchmark score on AGIEval
DeepSeek V4 Pro Base leads at 83.1%, followed by DeepSeek V4 Flash Base (82.6%) and Soofi S 30B-A3B (66.9%).
Visual analysis
Switch between model placement, score distribution and descriptive provider averages. Every view uses the same sourced leaderboard.
3 ranked models · higher is better
| Model | Rank | Provider | Score |
|---|---|---|---|
| DeepSeek V4 Pro Base | #1 | DeepSeek | 83.1% |
| DeepSeek V4 Flash Base | #2 | DeepSeek | 82.6% |
| Soofi S 30B-A3B | #3 | Soofi Project | 66.9% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | DeepSeek V4 Pro Base deepseek-v4-pro-base | DeepSeek | open | Estimated reference | 83.1% |
| #2 | DeepSeek V4 Flash Base deepseek-v4-flash-base | DeepSeek | open | Estimated reference | 82.6% |
| #3 | Soofi S 30B-A3B soofi-s-30b-a3b | Soofi Project | open | Estimated reference | 66.9% |
The top of this snapshot is led by DeepSeek V4 Pro Base at 83.1%; third place is 16.2 points behind. The top-3 spread is 16.2 points.
About AGIEval
A human-centric exam benchmark for general knowledge and reasoning reported in DeepSeek-V4 base-model evaluations. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
A human-centric exam benchmark for general knowledge and reasoning reported in DeepSeek-V4 base-model evaluations.
DeepSeek V4 Pro Base by DeepSeek currently leads with 83.1%.
3 models in the LuminaBench cohort have a qualifying score on this benchmark.
Related
2026 · 10 results · Reference
Artificial Analysis MMLU-Pro2026 · 4 results · Reference
Massive Multitask Language Understanding2020 · 8 results · Reference
Graduate-Level Google-Proof Q&A2023 · 70 results · Reference
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines2025 · 19 results · Reference
Massive Multitask Language Understanding Professional2024 · 42 results · Reference