OpenAI MRCR v2 8-needle 128K-256K leaderboard
2 ranked models · higher is better
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| GPT-5.5 | #1 | OpenAI | 87.5% |
| Claude Opus 4.7 (Adaptive) | #2 | Anthropic | 59.2% |
Reasoning · Benchmark profile
MRCR v2 slice focused on very long contexts at 128K-256K lengths.
Data verified 21 Jul 2026 · Methodology 1.6.0
Benchmark score on OpenAI MRCR v2 8-needle 128K-256K
GPT-5.5 leads at 87.5%, followed by Claude Opus 4.7 (Adaptive) (59.2%).
Visual analysis
Switch between model placement, score distribution and descriptive provider averages. Every view uses the same sourced leaderboard.
2 ranked models · higher is better
| Model | Rank | Provider | Score |
|---|---|---|---|
| GPT-5.5 | #1 | OpenAI | 87.5% |
| Claude Opus 4.7 (Adaptive) | #2 | Anthropic | 59.2% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | GPT-5.5 gpt-5.5 | OpenAI | closed | Estimated reference | 87.5% |
| #2 | Claude Opus 4.7 (Adaptive) claude-opus-4-7-max | Anthropic | closed | Estimated reference | 59.2% |
About OpenAI MRCR v2 8-needle 128K-256K
MRCR v2 slice focused on very long contexts at 128K-256K lengths. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
MRCR v2 slice focused on very long contexts at 128K-256K lengths.
GPT-5.5 by OpenAI currently leads with 87.5%.
2 models in the LuminaBench cohort have a qualifying score on this benchmark.
No. This benchmark is display-only and does not enter the overall Lumina composite.
Related