About BFCL v4
Definition and scoring
- Organisation
- Berkeley
- Category
- Agents
- Version
- v4
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Low
Berkeley Function-Calling Leaderboard v4 measuring structured tool and function invocation correctness across multi-turn scenarios. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗
