Every comparison here is built from the same data as the AI Coding Leaderboard: SWE-bench Verified as the ranking metric, tiebroken by the harder SWE-bench Pro, and no score printed until it is confirmed against an independent evaluation or a primary source. Figures last changed 2026-09-03.

head-to-head

Two models side by side on score, price, context window and access, with the gaps read honestly: where a difference sits inside the margin of error we say so rather than calling a winner.

best of

One question each, answered by ranking the whole field for it rather than by picking a favourite.