Every comparison here is built from the same data as the AI Coding Leaderboard: SWE-bench Verified as the ranking metric, tiebroken by the harder SWE-bench Pro, and no score printed until it is confirmed against an independent evaluation or a primary source. Figures last changed 2026-09-03.
head-to-head
Two models side by side on score, price, context window and access, with the gaps read honestly: where a difference sits inside the margin of error we say so rather than calling a winner.
- Claude Opus 5 vs GPT-5.6 Sol
- Claude Opus 5 vs Grok 4.6
- Claude Opus 5 vs GPT-5.6 Terra
- Claude Opus 5 vs GLM-5.3
- GPT-5.6 Sol vs Grok 4.6
- GPT-5.6 Sol vs GPT-5.6 Terra
- GPT-5.6 Sol vs GLM-5.3
- Grok 4.6 vs GPT-5.6 Terra
- Grok 4.6 vs GLM-5.3
- GPT-5.6 Terra vs GLM-5.3
- Claude Opus 5 vs Gemini 3.7 Flash
- GPT-5.6 Sol vs Gemini 3.7 Flash
- Claude Opus 5 vs DeepSeek V4 Flash
- GPT-5.6 Sol vs DeepSeek V4 Flash
best of
One question each, answered by ranking the whole field for it rather than by picking a favourite.