AI-Needle leaderboard
4 ranked models · higher is better
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Opus 4.5 | #1 | Anthropic | 74% |
| Qwen3.5 397B | #2 | Alibaba Cloud | 68.7% |
| Qwen3.6 Plus | #3 | Alibaba Cloud | 68.3% |
| GLM-5 | #4 | Z.AI | 63.3% |
Reasoning · Benchmark profile
A long-context retrieval benchmark that measures whether a model can recover relevant information embedded deep inside very long contexts.
Data verified 21 Jul 2026 · Methodology 1.6.0
Benchmark score on AI-Needle
Claude Opus 4.5 leads at 74%, followed by Qwen3.5 397B (68.7%) and Qwen3.6 Plus (68.3%).
Visual analysis
Switch between model placement, score distribution and descriptive provider averages. Every view uses the same sourced leaderboard.
4 ranked models · higher is better
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Opus 4.5 | #1 | Anthropic | 74% |
| Qwen3.5 397B | #2 | Alibaba Cloud | 68.7% |
| Qwen3.6 Plus | #3 | Alibaba Cloud | 68.3% |
| GLM-5 | #4 | Z.AI | 63.3% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.5 claude-opus-4-5 | Anthropic | closed | Reference only | 74% |
| #2 | Qwen3.5 397B qwen3.5-397b-a17b | Alibaba Cloud | open | Estimated reference | 68.7% |
| #3 | Qwen3.6 Plus qwen3-6-plus | Alibaba Cloud | closed | Reference only | 68.3% |
| #4 | GLM-5 glm-5 | Z.AI | open | Reference only | 63.3% |
The top of this snapshot is led by Claude Opus 4.5 at 74%; third place is 5.7 points behind. The top-4 spread is 10.7 points.
About AI-Needle
A long-context retrieval benchmark that measures whether a model can recover relevant information embedded deep inside very long contexts. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
A long-context retrieval benchmark that measures whether a model can recover relevant information embedded deep inside very long contexts.
Claude Opus 4.5 by Anthropic currently leads with 74%.
4 models in the LuminaBench cohort have a qualifying score on this benchmark.
No. This benchmark is display-only and does not enter the overall Lumina composite.
Related