About General AI Assistants
Definition and scoring
- Organisation
- BenchLM registry
- Category
- Agents
- Version
- 2024
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗
