Graphwalks BFS 0K-128K leaderboard
1 ranked models · higher is better
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| MAI-Thinking-1 | #1 | Microsoft | 90% |
Reasoning · Benchmark profile
Long-context graph traversal benchmark using breadth-first search tasks.
Data verified 21 Jul 2026 · Methodology 1.6.0
Visual analysis
Switch between model placement, score distribution and descriptive provider averages. Every view uses the same sourced leaderboard.
1 ranked models · higher is better
| Model | Rank | Provider | Score |
|---|---|---|---|
| MAI-Thinking-1 | #1 | Microsoft | 90% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | MAI-Thinking-1 mai-thinking-1 | Microsoft | closed | Estimated reference | 90% |
About Graphwalks BFS 0K-128K
Long-context graph traversal benchmark using breadth-first search tasks. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
Long-context graph traversal benchmark using breadth-first search tasks.
MAI-Thinking-1 by Microsoft currently leads with 90%.
1 model in the LuminaBench cohort have a qualifying score on this benchmark.
Related