Overall score
63.0
Rank #15 · A confidence
OpenAI · stable
OpenAI's documented reasoning model for coding and complex professional work.
Overall score
63.0
Rank #15 · A confidence
Capability
64.2
27 eligible benchmark families
Coverage
86%
Confidence A
Benchmarks scored
33
37 raw result rows
Model specification
Identity, configuration and price are preserved as dated source records.
Category profile
A category remains Not evaluated until an eligible source-qualified result is attached.
Benchmark scores
Best published score per benchmark. Open any row for the full per-benchmark leaderboard.
| Benchmark | Category | Score | Date |
|---|---|---|---|
| BenchLM Knowledge priorRanking weighted | research | 96.8% | 2026-07-15 |
| GPQA DiamondRanking weighted | reasoning | 92.8% | 2026-07-15 |
| BenchLM Reasoning priorRanking weighted | reasoning | 87.5% | 2026-07-15 |
| τ²-Bench TelecomRanking weighted | agents | 87.1345% | 2026-07-15 |
| τ-benchRanking weighted | agents | 87.1% | 2026-07-15 |
| CharXivRanking weighted | multimodal | 82.8% | 2026-07-15 |
| BrowseCompRanking weighted | agents | 82.7% | 2026-07-15 |
| MMMU-ProRanking weighted | multimodal | 81.2% | 2026-07-15 |
| LiveBenchRanking weighted | reasoning | 80.28% | 2026-07-15 |
| Terminal-BenchRanking weighted | agents | 78.28% | 2026-07-15 |
| OSWorld-VerifiedRanking weighted | agents | 75% | 2026-07-15 |
| AA Long Context ReasoningRanking weighted | research | 74% | 2026-07-15 |
| IFBench | instruction following | 73.9% | 2026-07-15 |
| IFEval | instruction following | 73.9% | 2026-07-15 |
| MCP AtlasRanking weighted | agents | 70.6% | 2026-07-15 |
| BenchLM Math priorRanking weighted | mathematics | 70.2% | 2026-07-15 |
| BenchLM Coding priorRanking weighted | coding | 64.17% | 2026-07-15 |
| BenchLM Multimodal priorRanking weighted | multimodal | 60% | 2026-07-15 |
| BenchLM Agentic priorRanking weighted | agents | 59.82% | 2026-07-15 |
| SWE-Bench ProRanking weighted | coding | 57.7% | 2026-07-15 |
| Long-Horizon-Terminal-BenchRanking weighted | agents | 57.6% | 2026-07-15 |
| Terminal-Bench HardRanking weighted | agents | 57.5758% | 2026-07-15 |
| SciCodeRanking weighted | coding | 56.6% | 2026-07-15 |
| ToolathlonRanking weighted | agents | 54.6% | 2026-07-15 |
| Humanity's Last ExamRanking weighted | reasoning | 52.1% | 2026-07-15 |
| GAIARanking weighted | agents | 48.2% | 2026-07-15 |
| FrontierMathRanking weighted | mathematics | 47.6% | 2026-07-15 |
| GDPval-AA v2Ranking weighted | agents | 44.734% | 2026-07-15 |
| APEX-Agents-AARanking weighted | agents | 33.2596% | 2026-07-15 |
| τ³-BankingRanking weighted | agents | 30.3093% | 2026-07-15 |
| CritPtRanking weighted | reasoning | 23.4286% | 2026-07-15 |
| ExploitGym | agents | 6% | 2026-07-15 |
| AA-OmniscienceRanking weighted | research | 5.65 index | 2026-07-15 |
Showing the best of 37 raw result rows across 33 benchmark families.
Price and efficiency
Operational values are shown only when the source publishes a comparable measurement or a valid derivation.
Verification summary
News and X posts are structurally excluded from this model's evidence graph.
Missing publication dimensions: value. The last source check for this exact model was 2026-07-15.
Related models
Exact variants remain separate even when they share a provider or family.