Best AI for coding
Claude Fable 5.1 scores highest, because the 2 September update of WebDev Arena puts it first at 1765 and the next row lands seventy-eight points below it. It is also the dearest row in the table, and this page shows who that trade is worth making for and who it is not.
- eleven models with real evidence
- weights on the page
- no hand-placed positions
Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. What changed
Where each tool's score comes from
The ring below is the same set of weights printed in the table headers further down.
- WebDev leaderboard 40
- Score per dollar 20
- Tooling maturity 15
- General text leaderboard 10
- Context window 10
- Access from Iran 5
We chose these weights, and that is the only judgment call in the table. Weight them differently and the order changes.
The ranking today, from verified evidence
A model that outranks us on the leaderboard appears in the table even where we have not written its page, or first place would only mean first among the ones we had time to write.
An empty cell means we could not verify that number, not that the tool scored zero.
Behind each number
Every judged score in this table carries a written reason, and the Iran column says how it was checked. Those two are open here. The measurement trail behind each number, which row of which leaderboard and on how many votes, opens under the model it belongs to.
1 Claude Fable 5.1
Tooling maturity 4/4 It ships on all five routes and Anthropic own table prints a model ID for each. In that same table the Claude Platform on AWS cell is a dash for Opus 5 and an ID for this model.
Access from Iran blocked / no working route Iran is not on Anthropic supported-countries list. Read from Anthropic own page rather than measured on the network. Nothing changed for this version.
Measurement trail (3)
WebDev leaderboard 1,765 The claude-fable-5.1-max row, rank 1 on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes
Score per dollar 35.3 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,504 The claude-fable-5.1-max row, rank 3 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes
2 Claude Opus 5
Tooling maturity 4/4 API, official app, command line tool, agent SDK, batch execution, and shipping on three major clouds
Access from Iran blocked / no working route Iran appears on neither of Anthropic two supported-countries lists, the API one or the claude.ai one. We read that on Anthropic own page rather than measuring it: an independent network test has not been run yet, and when it is, it will be written here.
Measurement trail (3)
WebDev leaderboard 1,687 The claude-opus-5-max row, rank 3 on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes
Score per dollar 67.5 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,493 The claude-opus-5-high row, rank 9 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes
3 Claude Fable 5
Tooling maturity 4/4 It carries the same infrastructure as Opus 5 and ships on all five routes: Anthropic own API, Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry.
Access from Iran blocked / no working route Iran is not on Anthropic supported-countries list. Read from Anthropic own page rather than measured on the network.
Measurement trail (3)
WebDev leaderboard 1,628 The claude-fable-5 row, rank 8 on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes
Score per dollar 32.6 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,507 The claude-fable-5 row, rank 1 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes
4 GPT-5.6 Sol
Tooling maturity 4/4 It reaches the top step on the maker own documentation: API, SDKs and an official CLI, the Codex IDE extension, an Agents SDK, a batch lane, and a first-party guide to running on Amazon Bedrock that names the id openai.gpt-5.6-sol. This is the one column that filled up for this model without openai.com ever opening.
Access from Iran blocked / no working route This column was empty until today and is not any more. The supported-countries list lives on openai.com, which still returns 403, but the same list also sits on developers.openai.com, and that domain is open. About 195 countries and territories are named and Iran is not among them. The page itself says access from outside the list can get an account blocked. Read from the maker own page rather than measured on the network.
Measurement trail (3)
WebDev leaderboard 1,616 The gpt-5.6-sol-xhigh (codex-harness) row, rank 11 on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes
Score per dollar 80.8 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,483 The gpt-5.6-sol-xhigh row, rank 17 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes
5 Kimi K3 Max
Tooling maturity 1/4 An OpenAI-compatible API plus one official client. The platform documentation lists no official command line tool or editor extension, so we do not score it higher.
Measurement trail (3)
WebDev leaderboard 1,674 The kimi-k3-max row, rank 4 on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes
Score per dollar 111.6 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,489 The kimi-k3-max row, rank 12 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes
6 Qwen3.8 Max
Measurement trail (3)
WebDev leaderboard 1,669 The qwen3.8-max row, rank 5 on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes
Score per dollar 278.2 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,480 The qwen3.8-max row, rank 22 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes
7 DeepSeek V4 Flash
Tooling maturity 1/4 A documented API plus one official client. No official agent tooling or editor extension is listed.
Measurement trail (3)
WebDev leaderboard 1,581 The deepseek-v4-flash-high row, rank 18 on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes
Score per dollar 1,197.7 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,438 The deepseek-v4-flash-high-preview row, rank 85 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes
8 Muse Spark 1.3
Tooling maturity 2/4 The Meta Model API plus availability on OpenRouter, which is more than a bare endpoint. But Meta own documentation lists neither an official CLI nor an editor extension, so it does not go past the second step.
Measurement trail (2)
WebDev leaderboard 1,623 The muse-spark-1.3-xhigh row on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes. The rank cell on this row reads N/A rather than a number, so the score is recorded and the official rank is not yet.
Score per dollar 381.9 WebDev Arena score divided by the price per million output tokens
9 Gemini 3.8 Flash
Tooling maturity 3/4 An API, official SDKs, a web studio and a Vertex route on Google Cloud. We hold back the fourth step because Google lists no separate official coding CLI for this family, unlike Anthropic and OpenAI.
Access from Iran blocked / no working route The Gemini API available-regions page names about 195 countries and territories and Iran is not among them. Read from Google own page rather than measured on the network.
Measurement trail (3)
WebDev leaderboard 1,567 The gemini-3.8-flash-high row, rank 19 on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes
Score per dollar 417.9 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,494 The gemini-3.8-flash-high row, rank 8 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes
10 GLM-5.2 Max
Tooling maturity 1/4 A documented API plus one official client. No official command line tool or editor extension is listed in the documentation.
Measurement trail (3)
WebDev leaderboard 1,585 The glm-5.2-max row, rank 16 on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes
Score per dollar 360.2 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,472 The glm-5.2-max row, rank 37 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes
11 Grok 4.5
Tooling maturity 3/4 An OpenAI-compatible API, Python and JavaScript SDKs, and an official command line agent called Grok Build. It does not reach the top step: there is no first-party editor extension, and the batch lane rejects this particular model.
Measurement trail (3)
WebDev leaderboard 1,556 The grok-4.5 row, rank 23 on the WebDev Arena leaderboard, updated 2 September 2026 with 640,806 votes
Score per dollar 259.3 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,471 The grok-4.5 row, rank 38 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes
Not enough verified data to rank
These are real tools we track, but more than half of their criteria have no verified number yet, so ranking them would be a guess.
How this ranking is calculated
Every criterion below has a weight and a source. Change a weight and the whole table recomputes. There is no hand-placed position anywhere in this hub.
| Criterion | Weight | Evidence |
|---|---|---|
| WebDev leaderboard | 40 | WebDev ArenaReal web development tasks with tools over several steps, which is closer to daily work than a single-question exam |
| Score per dollar | 20 | computed from two numbers aboveFor a team carrying a monthly bill, a dearer model one percent ahead is a bad buy |
| Tooling maturity | 15 | defined scale, with a written reason per assignmentA good model with no command line tool and no agent SDK slows down in daily work |
| General text leaderboard | 10 | Arena (Text)Weighted low because general conversation is not a coding measure, though it is not irrelevant either |
| Context window | 10 | vendor stated specificationOn a large repository a small window means chunking, and chunking means losing context |
| Access from Iran | 5 | our access column, with its method statedWeighted low because this column has its own page under AI in Iran |
Why Fable 5.1 leads, and why price did not stop it this time
One number settled it. In the 2 September update of WebDev Arena the claude-fable-5.1-max row sits first at 1765, and the next row we carry, claude-opus-5-max, comes in at 1687. Seventy-eight points, in a column whose whole top end fitted inside twenty-two points a week ago.
And this model loses the value column outright. Its output costs $50 per million tokens, twice Opus 5, and it scores zero there. Every previous time, that column has been what knocked an expensive model down. This is the first time a score gap has outweighed a price, and the reason is plain enough: it takes the full forty weight of the leaderboard column and also carries the tooling, context and Iran columns.
So the short answer is not one answer. If you work against the API and get a monthly bill, Opus 5 is half the price and second in this same table. If your work is the kind where getting it right once is cheaper than doing it twice, seventy-eight points is what you are paying for.
Where this table falls short
Our primary criterion is a human-preference leaderboard, not an automated test. WebDev Arena gives an absolute score, a date and a vote count, which is what publishing requires, but human preference also responds to readability and answer shape rather than only to whether the code is right. SWE-bench Verified would be a sharper measure, but its leaderboard renders in JavaScript and we have not read the figures directly. We do not print a number we have not read, even where someone else has quoted it.
One candidate on this page carries no rank: Gemini 3.1 Pro. It has a real score on the general text leaderboard, but we have not read its WebDev Arena row and Google publishes no context window for it at all, so more than half of the criteria weight is empty. Its name appears below the table; its position does not. GLM-5.2 Max and DeepSeek V4 Flash sat there too for a while, and now that their price and context window are verified from their own documentation, they carry a rank.
One thing about reading this leaderboard that came up this time and is not written down anywhere else: Qwen has two separate rows. There is qwen3.8-max at 1669 and qwen3.8-max-0902 at 1688, which is second on the whole board. We record the first, because that is the id in Alibaba own documentation and a dated row is a snapshot of one serving configuration. If you see Qwen listed second somewhere else, that is why, and not a mistake on our side.
One more thing, so the numbers do not confuse you: the scores are relative. Every column is normalised between its own minimum and maximum, so adding a new candidate moves everybody's score without any model having changed. Adding Grok 4.5 on 12 August 2026 did exactly that: the order held and the numbers rose. If today's figure does not match a screenshot from last week, that is why, and not because we touched anything.
Two more data changes moved the numbers on the same day. First, we read the text leaderboard in full this time and found that GLM-5.2 Max and DeepSeek V4 Flash had been on it all along; an earlier run had read only the top of the page. Both filled their general-text column. Second, GPT-5.6 Sol became an eighth candidate: openai.com still does not answer our server, but its developer documentation publishes the price and the context window, and its leaderboard row exists. One caveat about that row, since this is the coding page: it is named gpt-5.6-sol-xhigh (codex-harness), so the score belongs to the model inside OpenAI own coding agent rather than to the bare model. GPT landed between Kimi and Qwen, and no existing row changed place relative to the others.
28 August: DeepSeek got dearer and three rows moved
This is the only number in this reference that has ever gone up. On 12 August, DeepSeek V4 Flash cost $0.14 in and $0.28 out. Its own pricing page now carries two columns, peak and off peak, and even the cheaper column is above the old price: $0.22 and $0.66 off peak, $0.44 and $1.32 at peak. We record the peak column, because DeepSeek peak hours fall in the middle of the Iranian working day.
DeepSeek itself did not lose a place, because it is still the cheapest row in the table and still holds the top of the value column. But its lead narrowed, and since every column is normalised between its own minimum and maximum, that narrowing lifted three other rows: Qwen3.8 Max from fifth to third, Kimi K3 Max from third to fourth and GPT-5.6 Sol from fourth to fifth. None of those three models changed anything at all. That is what a relative ranking means, and it is why the arithmetic is printed on this page.
3 September: the gap at the top went from 22 points to 78
Until last week the three highest scores in this column sat inside 22 points, at 1692, 1674 and 1670. This page concluded then that the models differ by less than the numbers suggest, and that what genuinely changes a working day is tooling maturity: a command line tool, batch execution, presence on a cloud your company already has a contract with. For ranks two to eleven that still holds, and it is why the criterion carries fifteen points. For rank one it no longer does.
Three other rows also moved in the same update without changing anything about themselves. Kimi K3 Max went from third to fifth and Qwen3.8 Max from fifth to sixth, while Kimi leaderboard figure stayed at exactly 1674. GPT-5.6 Sol held fourth but its score fell from 51.6 to 48.8. All three are the arithmetic of new candidates entering and the columns normalising again.
Three new candidates, and one model deliberately not in the table
Muse Spark 1.3 from Meta and Gemini 3.8 Flash from Google joined the table today, both with a real row on the same leaderboard. Muse Spark carries a point other lists rarely make: Meta spent years being the maker that publishes its weights, and this one is not published. The model at the top of Meta line today is a closed one.
And the model that is not in the table: GPT-6 Astra, OpenAI new flagship at $10 in and $50 out. Neither leaderboard on this page has a row under that name, so seventy of its hundred weight would be empty and putting it in the table would produce one line reading "not ranked". It is in the model catalogue, and we wrote down here why it is not in the table so nobody assumes we missed it.
When not to take our first pick
- If your team runs on a cloud that does not carry Opus 5, the second model with working tooling beats the first without it.
- If your work is general conversation rather than code, ignore this table. The text leaderboard orders things differently.
- If the budget is genuinely tight, the open-weight models lower down trail by less score than they trail in price.
Where it falls short
- The primary criterion is a human-preference leaderboard, not an automated test against real repositories.
- There is no SWE-bench Verified figure in the table, because we could not read that leaderboard directly.
- Anthropic own Frontier-Bench claims are relative and carry no absolute number, so that column stays empty.
- Qwen3.8 Max now has a verified price and context window, but its tool maturity and Iran columns still carry no evidence, so 20 of its 100 weight is unmeasured. That is a gap in our data, not a weakness in the model.
- The model at the top of this table is also its dearest row and scores zero in the value column. If cost decides it for you, first place here is not your answer.
- GPT-6 Astra, OpenAI new flagship, is not in the table at all, because neither leaderboard carries a row under that name. So this table does not currently measure the most expensive model OpenAI sells.
Our take
If you have to pick one and cost is not the constraint, take Fable 5.1. If you get a monthly bill, Opus 5 is half the price and seventy-eight points lower, and that is a trade you have to make rather than one we can make for you. And before changing model at all, look once at your own token usage: in the work we have seen, trimming context saves more than switching model does.
Questions people actually ask
What is the best AI for coding
Today, Claude Fable 5.1, on the strength of first place on WebDev Arena at 1765 and a seventy-eight point gap to the next row. It is also the dearest row in the table, so if cost decides it for you the practical answer is Opus 5. That answer carries a date and any new release can change it.
Are open-weight models usable for code
Yes, but the gap widened on 2 September. Kimi K3 Max at 1674 and Qwen3.8 Max at 1669 are still close to Opus 5 and about ninety points below the model at the top. One new thing is worth knowing: Meta Muse Spark, which joined the table today, does not publish its weights, so the maker best known for publishing them ships its own top model closed.
Why does this ranking differ from other lists
Because we wrote the criteria and the weights down where you can see them. Most lists never explain their order, and where there is no explanation, any order is possible.
If you would rather not spend the week wiring these tools into a project, that work is part of
our app and custom software service