About OpenBookQA
Definition and scoring
- Organisation
- Todor Mihaylov, Peter Clark, Tushar Khot, Ashish Sabharwal
- Category
- Knowledge
- Version
- 2018
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
A science question-answering benchmark that tests whether models can apply a small open-book set of elementary science facts to multi-step reasoning questions. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗
