About InferenceBench
Definition and scoring
- Organisation
- Jehyeok Yeon, Ben Rank, Maksym Andriushchenko
- Category
- Agents
- Version
- 2026
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
A benchmark for open-ended LLM inference optimization by AI agents. Agents receive a base model, one H100, and a fixed time budget to build a valid OpenAI-compatible inference server that improves serving speed. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗
