The benchmarks below show Max's performance across the Arena leaderboards most relevant to its capabilities. Because Max is a router, the models we compare Max to reflect a point-in-time snapshot of which models were publicly available and routable when Max was last trained and evaluated. Max is updated periodically to incorporate the latest frontier models.
Max holds Pareto frontier performance when compared to its routing set across every modality it covers. It outranks all other models in this set for every supported arena except Single-Image Edit and Multi-Image Edit, where it places second. In these two arenas, Max offers a large latency benefit over the top model.

Introducing AutoEval to the Arena leaderboards
At Arena, our evaluations are dynamic and grounded in real-world use. But real-world signals take time to collect. Today, we’re introducing AutoEval scores to provide immediate, calibrated model ratings on real tasks when waiting for human votes to accumulate.









