the ranking
| # | Model | $/1M input | Best for |
|---|---|---|---|
| 1 | Laguna S 2.1 openPoolside | $0.10/1M | Frontier-adjacent agentic coding at small-model cost: a 118B mixture-of-experts with only ~8B active per token, open weights, and INT4/GGUF builds that run on local hardware. |
| 2 | DeepSeek V4 Flash openDeepSeek | $0.14/1M | The cheapest model making a near-frontier agentic claim: 284B MoE with 13B active, 1M context, and native Responses API plus Codex support at $0.28 per 1M output. |
| 3 | GPT-5.6 LunaOpenAI | $0.20/1M | The cost outlier of the leaders: within striking distance of the top scores while costing a fraction per solved task, and after an 80% price cut on Jul 30, 2026 it is now the cheapest ranked model on the board per token. |
| 4 | Gemini 3.5 Flash-LiteGoogle DeepMind | $0.30/1M | The cheapest ranked model on the board: 75.0% independently measured at $0.30 per 1M input, roughly a fifth the price of the Flash tier above it. |
| 5 | LongCat-2.0 openMeituan | $0.30/1M | A 1.6T Mixture-of-Experts agentic coder trained entirely on domestic Chinese chips; led OpenRouter usage in stealth as "Owl Alpha". |
| 6 | DeepSeek V4 Pro openDeepSeek | $0.435/1M | The cheapest frontier-class coder — top open-weights score at ~11× less than Opus. Best pick when cost or self-hosting rules. |
the top picks, decoded
1. Laguna S 2.1 — $0.10/1M $/1M input
Frontier-adjacent agentic coding at small-model cost: a 118B mixture-of-experts with only ~8B active per token, open weights, and INT4/GGUF builds that run on local hardware. Vendor-reported (Poolside, Jul 21 2026): Terminal-Bench 2.1 70.2% with thinking enabled (60.4% without), SWE-bench Pro 59.4%, SWE-bench Multilingual 78.5%, DeepSWE v1.1 40.4%.
Poolside published no SWE-bench Verified score, so it stays unranked pending an independent eval. vals.ai has evaluated Laguna M.1 (57.6%) and Laguna XS.2 (55.2%) but not S 2.1, and a near-miss version is not a match. Price: OpenRouter $0.10/$0.20 per 1M.
Context added 2026-08-23: Nvidia is paying Poolside $6B for a non-exclusive license to the Model Factory system that builds this family, is investing $1B at a $13B post-money valuation, and has made job offers to the 109 engineers who built Laguna, per Bloomberg. Poolside stays independent and keeps shipping, so the row is unaffected, but who is training future Laguna builds is now an open question. See /p/nvidia-poolside-6b-license-model-factory-109-staff/.
Re-checked 2026-09-13: vals.ai's SWE-bench Verified board remains archived since September 1 (frozen for posterity, no new models are being scored) and still shows no evaluation for this model; all 22 previously-tracked ranked scores were re-verified this run against the frozen table and are unchanged, exact match.
2. DeepSeek V4 Flash — $0.14/1M $/1M input
The cheapest model making a near-frontier agentic claim: 284B MoE with 13B active, 1M context, and native Responses API plus Codex support at $0.28 per 1M output. Independent (vals.ai, Aug 5 2026, mini-swe-agent bash-only harness): SWE-bench Verified 88.8%, for the DeepSeek-V4-Flash-0731 build, matching the checkpoint DeepSeek own change log describes. Ranks above Claude Opus 4.8 (88.6%) and below GPT-5.6 Luna (93.0%).
Vendor-reported agent numbers for the same build (DeepSeek change log, Jul 31 2026): Terminal-Bench 2.1 82.7, Cybergym 76.7, Toolathlon verified 70.3, DeepSWE 54.4, NL2Repo 54.2, Agent Last Exam 25.2, produced with DeepSeek own unreleased "DeepSeek Harness minimal mode" and not independently reproduced. The V4-Flash-Preview model card (pre-0731 weights) separately listed SWE-bench Verified 79.0 and Terminal-Bench 2.0 56.9; those are not carried over.
Same architecture and size as the preview (284B total / 13B active, FP4+FP8), re-post-trained only. Not to be confused with DeepSeek V4 Pro (ranked, 80.6%) or the plain DeepSeek V4 (77.4%). See /p/deepseek-v4-flash-claims-82-7-terminal-bench/.
3. GPT-5.6 Luna — $0.20/1M $/1M input
The cost outlier of the leaders: within striking distance of the top scores while costing a fraction per solved task, and after an 80% price cut on Jul 30, 2026 it is now the cheapest ranked model on the board per token. Independent (vals.ai, eval listed Jul 17 2026, mini-swe-agent bash-only harness): SWE-bench Verified 93.00% ±1.14.
Added Jul 21, 2026 — this row was missing from the board even though vals.ai had already evaluated Luna, and our Kimi K3 note referenced its 93.0% score without ever listing it; adding it moves every row below it down one rank. Treat 3rd and 4th as a tie: Kimi K3's 93.40% ±1.11 is 0.4 points higher, well inside the combined margin of error (~0.25 sigma), so the ordering between them is not significant.
Like the rest of the GPT-5.6 family, OpenAI has published no SWE-bench Verified figure of its own, so we rank on the independent number per our standing rule. The striking number is cost: vals.ai measured $0.21 per test against $1.15 for GPT-5.6 Sol, $1.92 for Claude Opus 4.8 and $2.05 for Claude Fable 5, at 201s median latency. Pricing added Aug 1, 2026: OpenAI cut Luna 80% on Jul 30, 2026, from $1/$6 to $0.20/$1.20 per 1M, making it the cheapest ranked model on this board per token.
Vendor-announced (OpenAI, via its own post and consistent reporting from BleepingComputer and VentureBeat); openai.com/api/pricing returns 403 to our fetcher, so this is a vendor figure rather than one we read off the pricing page ourselves. It had been left blank since Jul 21 for exactly that reason.
how we rank
We rank by SWE-bench Verified (500 real, human-validated GitHub issues resolved end-to-end), tiebroken by the harder SWE-bench Pro. A score is only printed once confirmed against an independent evaluation or the maker's primary source — and every row states which kind it is. Where both exist, we print both: one as the ranked score, the other in that row's note. We would rather show you the gap than ask you to trust our pick.
Our independent reference is vals.ai, which runs every model itself through the same minimal bash-only harness (mini-swe-agent), so the models are compared on equal footing. That matters more than it sounds: SWE-bench scores a model and its scaffolding together, and vendors report using their own tuned scaffolds. Against vals.ai's neutral harness, the vendor claims on this board run 2.6 to 11.6 points optimistic.
So rows marked vendor-reported are best-case numbers and are not strictly comparable to the independent ones — where we know the independent figure, we print it in the row's note.
That is a deliberate choice and you should know we made it: on five rows (Claude Sonnet 5, MiniMax M3, Qwen3.7 Max, Kimi K2.6, DeepSeek V4 Pro) an independent score for that exact model exists and disagrees with the maker's, and we still rank on the maker's published figure because it is the number that model is sold and quoted on. On four of those the independent number is lower.
DeepSeek V4 Pro runs the other way: vals.ai measures the 0813 GA build 15.8 points ABOVE DeepSeek's own claim, which the disclose-don't-rebase rule was never written for, so that row is flagged in its note for a human decision on whether to rebase it. We disclose the independent one beside it instead of quietly restating the board on a single evaluator's harness choice. The honest consequence: positions that straddle the two regimes are approximate.
Qwen3.7 Max is the sharpest case — it sits at #16 on Alibaba's 80.4%, and on the neutral harness its 68.8% would put it far down the table. Note that llm-stats, which we previously miscredited as an independent tracker, labels its own SWE-bench Verified table "Verified: 0 / Self-reported: 104"; it aggregates vendor claims. Models still being checked are marked “verifying” and shown without a number rather than estimated.
Prices are per 1M input tokens on the standard API tier and can change — always confirm current pricing with the provider.
Want the raw numbers? The full dataset is public: JSON · CSV.
From our full AI Coding Leaderboard (2026-09-03). We only rank scores confirmed against primary sources.