SWE-bench Verified scores of leading AI coding modelsHorizontal bars comparing the SWE-bench Verified score of each verified frontier coding model, against the late-2024 frontier baseline of about 49 percent.SWE-BENCH VERIFIED (% RESOLVED)2024 ≈ 49%Claude Opus 597% GPT-5.6 Sol96.2% Grok 4.695.6% GPT-5.6 Terra95.4% GLM-5.395.4% Claude Fable 595% Kimi K393.4% GPT-5.6 Luna93% DeepSeek V4 Flash88.8% Claude Opus 4.888.6% Grok 4.586.6% Muse Spark 1.286.6% Qwen3.8-27B86% Qwen3.8-Max85.6% Claude Sonnet 585.2% GLM 5.282.8% GPT-5.582.6% Muse Spark 1.182% Gemini 3.7 Flash80.8% DeepSeek V4 Pro80.6% Gemini 3.1 Pro80.6% MiniMax M380.5% Qwen3.7 Max80.4% Kimi K2.680.2% Gemini 3.8 Flash80% Gemini 3.6 Flash79.6% Composer 2.579.6% Gemini 3.5 Flash78.8% Gemini 3.5 Flash-Lite75%</> genztech.blog
Fig 1 · benchmark SWE-bench Verified — the share of real, human-validated GitHub issues an AI resolves end-to-end. Only models whose scores we have confirmed are charted; models still being verified are left off rather than estimated. Sources: vals.ai independent evaluations where available, otherwise maker reports — each row states which. Vendor-reported scores use the vendor's own scaffold and run high.

today's ai coding benchmark standings

#ModelSWE-bench VerifiedSWE-bench ProInputBest for
verifying Cognition SWE-2Cognition Launched Sept 10, 2026, post-trained from Kimi K3 (2.8T params) with a cost-penalized RL reward across three effort tiers (medium/high/max). Nothing to rank on yet: independent.Released Sept 10, 2026. Cognition's own numbers, all on benchmarks Cognition built or controls (FrontierCode 1.1 Main, DeepSWE 1.1, Terminal-Bench…
25 Gemini 3.8 Flash Google DeepMind 80.0% $0.75/1M Google's third Flash release in six weeks, officially launched Sept 2, 2026. Now has an independent score, landing in a tight cluster with Kimi K2.6 and Gemini 3.6 Flash rather than clearly ahead of either.Independent (vals.ai, observed 2026-09-03, mini-swe-agent bash-only harness): SWE-bench Verified 80.00% ±1.79. Leaves the verifying queue it entered… full note →
1 Claude Opus 5 Anthropic 97.0% $5/1M The highest independently measured coding score on the board at 97.0%, at half the price of Fable 5. Strongest on short and medium tasks, though GPT-5.6 Sol still edges it on multi-hour work.Independent (vals.ai, observed Jul 25 2026, mini-swe-agent bash-only harness): SWE-bench Verified 97.00% ±0.76, the highest score on the board and… full note →
2 GPT-5.6 Sol OpenAI 96.2% $5/1M The strongest model on long tasks: 98% on the 1-to-4-hour tier, ahead of Claude Opus 5, and the second-highest overall score.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 96.20% ±0.86 — now the second-highest score on the board… full note →
6 Claude Fable 5 Anthropic 95.0% 80.3% $10/1M Mythos-class flagship for long-horizon agentic runs: the model to reach for when a task spans hours and hundreds of tool calls and has to actually finish.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 95.00% ±0.98. Held the top score until GPT-5.6 Sol was… full note →
7 Kimi K3 open Moonshot AI 93.4% $3/1M The highest-scoring downloadable coding model we track, and by far the cheapest way to buy a 90%+ result at $3 per 1M input. Weights shipped Jul 27, 2026 as 96 shards totalling 1.56 TB, 4-bit MXFP4 only: roughly 1,454 GiB, which fits one eight-card 192GB node. No full-precision base checkpoint was published.Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 93.40% ±1.11. Verified Jul 18, 2026 — vals.ai had not run K3 at our Jul… full note →
8 GPT-5.6 Luna OpenAI 93.0% $0.20/1M The cost outlier of the leaders: within striking distance of the top scores while costing a fraction per solved task, and after an 80% price cut on Jul 30, 2026 it is now the cheapest ranked model on the board per token.Independent (vals.ai, eval listed Jul 17 2026, mini-swe-agent bash-only harness): SWE-bench Verified 93.00% ±1.14. Added Jul 21, 2026 — this row was… full note →
10 Claude Opus 4.8 Anthropic 88.6% 69.2% $5/1M The hardest agentic refactors and long, autonomous multi-file tasks where every point of accuracy saves a human review cycle.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 88.6% ±1.42. Corrected Jul 17, 2026: we previously printed… full note →
11 Grok 4.5 SpaceXAI (xAI) 86.6% 64.7% $2/1M The best value at the top of the board: it solves a SWE-bench Verified task for about $2.31 of input, less than half what the two models above it cost, and it is the fastest of the leaders.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 86.6% ±1.52. Verified Jul 17, 2026 — it launched Jul 8 with… full note →
3 Grok 4.6 SpaceXAI (xAI) 95.6% $2/1M A post-training refresh of Grok 4.5 that lands in the top group on the neutral harness, at a third the price of the models around it.Independent (vals.ai, benchmark updated 2026-08-12, mini-swe-agent bash-only harness): SWE-bench Verified 95.60% ±0.92. Entered ranked 2026-08-13… full note →
12 Muse Spark 1.2 Meta 86.6% $1.25/1M Bounded, well-specified changes at low cost, especially if you accept the contributor tier and let Meta train on your traffic.Independent (vals.ai, Aug 6 2026, mini-swe-agent bash-only harness): SWE-bench Verified 86.6%. Added Aug 7, 2026, the day after Meta launched it… full note →
15 Claude Sonnet 5 Anthropic 85.2% 63.2% $2/1M The best closed-model value — near-Opus scores at ~2.5× less, and the default daily driver for most developers.Vendor-reported (Anthropic), on Anthropic's own scaffold. Independent comparison: vals.ai's bash-only harness measures Sonnet 5 at 79.6% ±1.80, 5.6… full note →
16 GLM 5.2 open Z.ai (Zhipu AI) 82.8% $1.40/1M The best open-weight coder on this board that you can actually download today: an independently measured 82.8%, MIT-licensed, and roughly a third the input price of the closed models above it.Independent (vals.ai, Jul 22 2026, mini-swe-agent bash-only harness): SWE-bench Verified 82.8% ±1.69, 9th of 75 systems. Added Jul 27, 2026 after a… full note →
17 GPT-5.5 OpenAI 82.6% 58.6% $5/1M OpenAI's strongest agentic coder, with the deepest tooling and ecosystem breadth of the closed labs.Verified score from vals.ai independent eval; Pro is OpenAI-reported (rivals flag possible memorization on Pro). Price: OpenAI list $5/$30 per 1M… full note →
18 Muse Spark 1.1 Meta 82.0% $1.25/1M Meta's first paid model, and a genuine value pick: a top-10 verified score for $1.52 per solved task, cheaper per result than every model ranked above it.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 82.0% ±1.72. Verified Jul 17, 2026; Meta published no… full note →
20 DeepSeek V4 Pro open DeepSeek 80.6% 55.4% $0.435/1M The cheapest frontier-class coder — top open-weights score at ~11× less than Opus. Best pick when cost or self-hosting rules.Vendor-reported: DeepSeek's own model card, Pro-Max mode (SWE-bench Verified 80.6%, SWE-bench Pro 55.4%). Updated Aug 12, 2026: DeepSeek shipped the… full note →
21 Gemini 3.1 Pro Google DeepMind 80.6% 54.2% $2/1M Google's strongest coding model today, with deep Workspace/Cloud integration. (A 3.5 Pro is expected but not shipped.)Vendor-reported (DeepMind) pass rate. No independent eval of this exact model; vals.ai has run Gemini 3.1 Pro Preview (02/26) at 78.8%, a preview… full note →
22 MiniMax M3 open MiniMax 80.5% 59.0% $0.60/1M Open weights with 1M context, multimodal input and computer use — beats GPT-5.5 on SWE-bench Pro at 5–10% of the cost.Vendor-reported at launch (Jun 1, 2026). Independent comparison: vals.ai's bash-only harness measures MiniMax-M3 at 75.0% ±1.94, 5.5 points lower, so… full note →
23 Qwen3.7 Max Alibaba 80.4% 60.6% $1.25/1M The best non-Claude score on the hardest benchmark — 60.6% SWE-bench Pro — built for long-horizon coding agents.Vendor-reported (May 20, 2026), and the widest vendor-versus-independent gap on this board: vals.ai's bash-only harness measures Qwen 3.7 Max at… full note →
24 Kimi K2.6 open Moonshot AI 80.2% 58.6% $0.95/1M A top-three open coder whose 58.6% SWE-bench Pro beats several closed flagships.Vendor-reported (10-run average on Moonshot's own SWE-agent harness). Independent comparison: vals.ai's bash-only harness measures Kimi K2.6 at 76.2%… full note →
26 Gemini 3.6 Flash Google DeepMind 79.6% $1.50/1M Google's new general workhorse: it now outscores 3.5 Flash on an independent coding eval while emitting 17% fewer output tokens.Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 79.60% ±1.80. Verified Jul 23, 2026 — vals.ai had not run it at our Jul… full note →
19 Gemini 3.7 Flash Google DeepMind 80.8% $0.75/1M Google's volume tier at half the price of 3.6 Flash, and the first 3.7-series model with an independent coding score. A real gain over 3.6 Flash, unlike the last Flash bump.Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 80.80% ±1.76. Released Aug 13, 2026, three weeks after Gemini 3.6 Flash… full note →
27 Composer 2.5 Cursor 79.6% $0.50/1M Cheap, fast in-editor coding if you already pay for Cursor. There is no way to use it anywhere else, which is the whole catch.Independent (vals.ai, Jul 22 2026, mini-swe-agent bash-only harness): SWE-bench Verified 79.6% ±1.80, 13th of 75 systems. Added Jul 27, 2026 from a… full note →
28 Gemini 3.5 Flash Google DeepMind 78.8% $1.50/1M Frontier-ish coding at Flash speed and price, with computer use built in as a native tool.vals.ai independent eval. See our decode of its native computer-use tool. Price: Google list $1.50/$9 per 1M (cached input $0.15).
4 GPT-5.6 Terra OpenAI 95.4% $2/1M A mid-tier price with top-tier coding after vals.ai revised its score upward by 20 points. At $2/$12 per 1M it now scores within error bars of Sol, which costs 2.5x more on output.Independent (vals.ai, re-evaluated, mini-swe-agent bash-only harness): SWE-bench Verified 95.4% ±0.94, 5th of 86 systems. CORRECTION, 2026-08-20… full note →
29 Gemini 3.5 Flash-Lite Google DeepMind 75.0% 54.2% $0.30/1M The cheapest ranked model on the board: 75.0% independently measured at $0.30 per 1M input, roughly a fifth the price of the Flash tier above it.Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 75.00% ±1.94. Verified Jul 23, 2026 — vals.ai had not run it at our Jul… full note →
verifying Muse Glimmer openMeta Free/1M A 30B dense agentic model small enough to run always-on inside a 24GB consumer GPU, distilled from Muse Spark. Nothing to rank on yet.Released Aug 10, 2026 with weights on Hugging Face under a plain Apache 2.0 license, Meta's first genuinely permissive open-weight release in the…
14 Qwen3.8-Max Alibaba 85.6% $2.00/1M Alibaba's 2.4T-parameter flagship, pitched at long-horizon autonomous coding. Independently scored, and slow: the highest latency of any ranked model here.Independent (vals.ai, benchmark updated 2026-08-08, mini-swe-agent bash-only harness): SWE-bench Verified 85.6% ± 1.57. Announced Aug 3, 2026 as "a… full note →
verifying Laguna S 2.1 openPoolside 59.4% $0.10/1M Frontier-adjacent agentic coding at small-model cost: a 118B mixture-of-experts with only ~8B active per token, open weights, and INT4/GGUF builds that run on local hardware.Vendor-reported (Poolside, Jul 21 2026): Terminal-Bench 2.1 70.2% with thinking enabled (60.4% without), SWE-bench Pro 59.4%, SWE-bench Multilingual…
13 Qwen3.8-27B open Alibaba 86.0% 61.7% The best independently measured open-weight model under a true OSI licence, and small enough to serve on a single accelerator.Independent (vals.ai, 2026-08-20 sweep, mini-swe-agent bash-only harness): SWE-bench Verified 86.0% ±1.55, 14th of 86 systems. Left the verifying… full note →
verifying LongCat-2.0 openMeituan 59.5% $0.30/1M A 1.6T Mixture-of-Experts agentic coder trained entirely on domestic Chinese chips; led OpenRouter usage in stealth as "Owl Alpha".Vendor-reported (Meituan, Jun 30 2026): SWE-bench Pro 59.5%, Terminal-Bench 70.8%. No SWE-bench Verified score or independent eval published yet, so…
verifying Cohere North Mini Code openCohere Free/1M Private, self-hosted agentic coding for enterprises that cannot send source to a cloud API; runs on a single H100.Open-weight ~30B mixture-of-experts coder, free to use, runs on a single H100 with 256K context. Cohere published no SWE-bench Verified/Pro or…
verifying KAT-Coder-Pro V2.5Kwaipilot (Kuaishou) 65.2% $0.74/1M Long-horizon agentic coding on a budget: a 72B-active MoE tuned for tool use, with a cheaper Air tier at $0.15/$0.60 on the same 256K context.Vendor-reported (KAT-Coder-V2.5 technical report, arXiv 2607.05471): SWE-bench Pro 65.2%, second only to Opus 4.8 at 69.2%, plus a best-in-test…
9 DeepSeek V4 Flash open DeepSeek 88.8% $0.14/1M The cheapest model making a near-frontier agentic claim: 284B MoE with 13B active, 1M context, and native Responses API plus Codex support at $0.28 per 1M output.Independent (vals.ai, Aug 5 2026, mini-swe-agent bash-only harness): SWE-bench Verified 88.8%, for the DeepSeek-V4-Flash-0731 build, matching the… full note →
5 GLM-5.3 Z.ai (Zhipu AI) 95.4% The highest-scoring model on this board whose maker has promised open weights, and the cheapest route to a 95%-plus measured score at $0.34 per test.Independent (vals.ai, 2026-08-20 sweep, mini-swe-agent bash-only harness): SWE-bench Verified 95.4% ±0.94, 6th of 86 systems. Left the verifying… full note →

how we rank: swe-bench verified, swe-bench pro and terminal-bench

We rank by SWE-bench Verified (500 real, human-validated GitHub issues resolved end-to-end), tiebroken by the harder SWE-bench Pro. A score is only printed once confirmed against an independent evaluation or the maker's primary source — and every row states which kind it is. Where both exist, we print both: one as the ranked score, the other in that row's note. We would rather show you the gap than ask you to trust our pick.

Our independent reference is vals.ai, which runs every model itself through the same minimal bash-only harness (mini-swe-agent), so the models are compared on equal footing. That matters more than it sounds: SWE-bench scores a model and its scaffolding together, and vendors report using their own tuned scaffolds. Against vals.ai's neutral harness, the vendor claims on this board run 2.6 to 11.6 points optimistic.

So rows marked vendor-reported are best-case numbers and are not strictly comparable to the independent ones — where we know the independent figure, we print it in the row's note.

That is a deliberate choice and you should know we made it: on five rows (Claude Sonnet 5, MiniMax M3, Qwen3.7 Max, Kimi K2.6, DeepSeek V4 Pro) an independent score for that exact model exists and disagrees with the maker's, and we still rank on the maker's published figure because it is the number that model is sold and quoted on. On four of those the independent number is lower.

DeepSeek V4 Pro runs the other way: vals.ai measures the 0813 GA build 15.8 points ABOVE DeepSeek's own claim, which the disclose-don't-rebase rule was never written for, so that row is flagged in its note for a human decision on whether to rebase it. We disclose the independent one beside it instead of quietly restating the board on a single evaluator's harness choice. The honest consequence: positions that straddle the two regimes are approximate.

Qwen3.7 Max is the sharpest case — it sits at #16 on Alibaba's 80.4%, and on the neutral harness its 68.8% would put it far down the table. Note that llm-stats, which we previously miscredited as an independent tracker, labels its own SWE-bench Verified table "Verified: 0 / Self-reported: 104"; it aggregates vendor claims. Models still being checked are marked “verifying” and shown without a number rather than estimated.

Prices are per 1M input tokens on the standard API tier and can change — always confirm current pricing with the provider.

our picks: the best ai models for coding right now

Best overallClaude Opus 5

The top independently measured score on SWE-bench Verified at 97.0%, and it costs half what Fable 5 does. Treat the lead over GPT-5.6 Sol and Fable 5 as a tie, but there is no reason to pay more for the same tier.

Best value (closed)Claude Sonnet 5

85.2% SWE-bench Verified at $2/1M — near-frontier coding at a rounding-error price. The default for most work.

Best open weightsQwen3.8-27B

86.0% Verified measured independently by vals.ai, Apache 2.0 weights on Hugging Face, and a dense 27B that serves on a single accelerator. It overtook GLM 5.2 (82.8%, MIT) on the 2026-08-20 sweep. Kimi K3 scores higher at 93.4% but ships 4-bit only under a bespoke non-OSI licence, so it is open for inference rather than open outright. GLM-5.3 scores far higher still at 95.4%, but Z.ai has not released its weights yet.

Best for autonomous agentsGPT-5.6 Sol

Best on the tasks that actually run long: 98% of the 1-to-4-hour SWE-bench tier, ahead of both Claude Opus 5 (90%) and Fable 5 (93%), which is what matters for unattended repo-wide work.

Best free optionClaude Sonnet 5

It's the default model for free claude.ai users — frontier-class coding at no cost for everyday tasks.

Hardest-tasks dark horseQwen3.7 Max

Its 60.6% on SWE-bench Pro is the best non-Claude score on the benchmark that's hardest to game.

compare head-to-head

how the field got here

Top SWE-bench Verified score over time, Jul 2 to Sep 3 2026A line chart of the best confirmed SWE-bench Verified score on this leaderboard at each date the ranked data changed. It runs from 86 percent (Claude Opus 4.8) on Jul 2 to 97 percent (Claude Opus 5) on Sep 3, with 3 changes of leader. Points mark the dates we recorded a change; the axis is to scale, so gaps are periods with no movement.TOP SWE-BENCH VERIFIED SCORE OVER TIMEeach point = a date our ranked data changed · axis to scale84%92%99%Jul 2 Jul 3 Jul 17 Jul 18 Jul 19 Jul 21 Jul 23 Jul 25 Jul 27 Aug 1 Aug 6 Aug 7 Aug 10 Aug 13 Aug 17 Aug 20 Sep 3~86%Claude Opus 4.8 95.0%Claude Fable 5 96.2%GPT-5.6 Sol 96.2%GPT-5.6 Sol 96.2%GPT-5.6 Sol 96.2%GPT-5.6 Sol 96.2%GPT-5.6 Sol 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5</> genztech.blog
Fig 2 · history The frontier since we started tracking on Jul 2 2026: +11 points in 63 days, across 3 changes of leader. We plot one point per date the ranked data actually moved — not per edit — and the axis is to scale, so a flat stretch means nothing changed rather than that we stopped looking. Full series: JSON · CSV, CC BY 4.0.
  1. 2021GitHub Copilot preview Autocomplete-in-the-editor goes mainstream.
  2. 2023ChatGPT + GPT-4, then Cursor Chat-based coding and the first AI-native editor arrive.
  3. Aug 2024SWE-bench Verified launches A human-validated benchmark of real GitHub issues sets an honest bar.
  4. Oct 2024Claude 3.5 Sonnet hits ~49% Agents begin resolving real issues, not just snippets.
  5. 2025Terminal agents Claude Code and Codex CLI move AI out of the editor into the whole repo.
  6. Apr–Jun 2026Open weights close the gap DeepSeek V4, Kimi K2.6 and peers cluster at ~80% Verified — for pennies.
  7. 2026Verified saturates in the mid-80s SWE-bench Pro and Terminal-Bench become the real differentiators.
  8. Jul 2026Claude Fable 5 returns Restored after a 20-day export-control suspension; retakes SWE-bench Pro at 80.3%.
  9. Jun 30 2026Meituan open-sources LongCat-2.0 A 1.6T MoE coder trained entirely on domestic Chinese chips; vendor-reported 59.5% SWE-bench Pro, awaiting independent eval.
  10. Jul 25 2026Claude Opus 5 takes #1 vals.ai measures 97.00% ±0.76 on the neutral harness, one day after launch. A statistical tie with GPT-5.6 Sol and Claude Fable 5.
Primary sources

cite & embed

Free to cite and embed with attribution. Raw data: JSON · CSV. Embedding this live table adds it to your site and credits GENZ TECH.

CiteGENZ TECH. (2026). AI Coding Leaderboard. https://genztech.blog/ai-coding-leaderboard/
Embed<iframe src="https://genztech.blog/ai-coding-leaderboard/embed/" width="100%" height="470" loading="lazy" title="AI Coding Leaderboard by GENZ TECH" style="border:1px solid #26282b;border-radius:10px;max-width:560px"></iframe>