🚨 Grok 4.5 tops the new Long-Horizon-Terminal-Bench
LuminaXspace

Researchers tested frontier AI models on 46 extremely long, real world terminal tasks:
• Grok 4.5 ranked #1 overall with a 19.6% pass rate • Grok fully solved 13 of the 46 tasks • Each task consumed approximately 9.9 million tokens on average • Agents completed roughly 231 episodes per task • Average execution time reached 85.3 minutes • Tasks covered software engineering, scientific computing, multimodal analysis etc
SpaceXAI cooked with Grok 4.5.

