τ²-Bench Airline Domain leaderboard
1 ranked models · higher is better
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| ZAYA1-74B-Preview | #1 | Zyphra | 56.1% |
Agents · Benchmark profile
τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.
Data verified 21 Jul 2026 · Methodology 1.6.0
Visual analysis
Switch between model placement, score distribution and descriptive provider averages. Every view uses the same sourced leaderboard.
1 ranked models · higher is better
| Model | Rank | Provider | Score |
|---|---|---|---|
| ZAYA1-74B-Preview | #1 | Zyphra | 56.1% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | ZAYA1-74B-Preview zaya1-74b-preview | Zyphra | open | Estimated reference | 56.1% |
About τ²-Bench Airline Domain
τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.
ZAYA1-74B-Preview by Zyphra currently leads with 56.1%.
1 model in the LuminaBench cohort have a qualifying score on this benchmark.
Related