Job

Best AI for coding

Claude Fable 5.1 scores highest, because the 2 September update of WebDev Arena puts it first at 1765 and the next row lands seventy-eight points below it. It is also the dearest row in the table, and this page shows who that trade is worth making for and who it is not.

  • eleven models with real evidence
  • weights on the page
  • no hand-placed positions

Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. What changed

Where each tool's score comes from

The ring below is the same set of weights printed in the table headers further down.

100Weights
  • WebDev leaderboard 40
  • Score per dollar 20
  • Tooling maturity 15
  • General text leaderboard 10
  • Context window 10
  • Access from Iran 5

We chose these weights, and that is the only judgment call in the table. Weight them differently and the order changes.

The ranking today, from verified evidence

A model that outranks us on the leaderboard appears in the table even where we have not written its page, or first place would only mean first among the ones we had time to write.

Computed score

Fixed scale, 0 to 100, higher is better

  • Claude Fable 5.1 73.1
  • GPT-6 Astra 70.2
  • Claude Opus 5 59.5
  • Claude Fable 5 51.1
  • Muse Spark 1.3 48.2
  • GPT-5.6 Sol 46.0
  • Gemini 3.8 Flash 43.4
  • Qwen3.8 Max 43.3
  • Kimi K3 Max 38.7
  • GLM-5.2 Max 31.9
  • Grok 4.5 21.8
  • Hy4 preview 20.3

Published price

USD per million tokens, lower is better

No verified price for: Hy4 preview

Best AI for coding
Rank Tool Score WebDev leaderboard 40 Score per dollar 20 Tooling maturity 15 General text leaderboard 10 Context window 10 Access from Iran 5
1 Claude Fable 5.1 Current pick 73.1 1,764 Source: arena.ai 35.3 Source: arena.ai 4/4 1,504 Source: arena.ai 1 million tokens Source: platform.claude.com blocked / no working route
2 GPT-6 Astra 70.2 1,796 Source: arena.ai 35.9 Source: arena.ai 4/4 not verified 1 million tokens Source: developers.openai.com blocked / no working route
3 Claude Opus 5 the maker has replaced it with Claude Opus 5.5 59.5 1,691 Source: arena.ai 67.6 Source: arena.ai 4/4 1,493 Source: arena.ai 1 million tokens Source: platform.claude.com blocked / no working route
4 Claude Fable 5 the maker has replaced it with Claude Fable 5.1 51.1 1,628 Source: arena.ai 32.6 Source: arena.ai 4/4 1,507 Source: arena.ai 1 million tokens Source: platform.claude.com blocked / no working route
5 Muse Spark 1.3 48.2 1,650 Source: arena.ai 388.2 Source: arena.ai 2/4 not verified 1 million tokens Source: developer.meta.com not verified
6 GPT-5.6 Sol 46.0 1,617 Source: arena.ai 80.9 Source: arena.ai 4/4 1,483 Source: arena.ai 1 million tokens Source: developers.openai.com blocked / no working route
7 Gemini 3.8 Flash 43.4 1,568 Source: arena.ai 418.1 Source: arena.ai 3/4 1,494 Source: arena.ai not verified blocked / no working route
8 Qwen3.8 Max 43.3 1,670 Source: arena.ai 278.3 Source: arena.ai not verified 1,480 Source: arena.ai 1 million tokens Source: www.alibabacloud.com not verified
9 Kimi K3 Max 38.7 1,674 Source: arena.ai 111.6 Source: arena.ai 1/4 1,489 Source: arena.ai 1 million tokens Source: platform.kimi.ai not verified
10 GLM-5.2 Max 31.9 1,589 Source: arena.ai 361.1 Source: arena.ai 1/4 1,472 Source: arena.ai 1 million tokens Source: docs.z.ai not verified
11 Grok 4.5 21.8 1,556 Source: arena.ai 259.3 Source: arena.ai 3/4 1,471 Source: arena.ai 500k tokens Source: docs.x.ai not verified
12 Hy4 preview 20.3 1,623 Source: arena.ai not verified not verified not verified 1 million tokens Source: huggingface.co not verified

An empty cell means we could not verify that number, not that the tool scored zero.

Behind each number

Every judged score in this table carries a written reason, and the Iran column says how it was checked. Those two are open here. The measurement trail behind each number, which row of which leaderboard and on how many votes, opens under the model it belongs to.

1 Claude Fable 5.1

Tooling maturity 4/4 It ships on all five routes and Anthropic own table prints a model ID for each. In that same table the Claude Platform on AWS cell is a dash for Opus 5 and an ID for this model.

Access from Iran blocked / no working route Iran is not on Anthropic supported-countries list. Read from Anthropic own page rather than measured on the network. Nothing changed for this version.

Measurement trail (3)

WebDev leaderboard 1,764 The claude-fable-5.1-max row, rank 2 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes. Arena gives this row a rank range of 1 to 2, the same range as the row above it, so the board itself does not settle first and second.

Score per dollar 35.3 WebDev Arena score divided by the price per million output tokens

General text leaderboard 1,504 The claude-fable-5.1-max row, rank 3 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

2 GPT-6 Astra

Tooling maturity 4/4 It reaches the top step on the maker own documentation: the API with a batch lane, hosted tools from web search to computer use and MCP, the Agents API that entered public beta on 10 September with a managed Codex harness, and availability on Amazon Bedrock. OpenAI own Bedrock guide says the model is served in one region so far, us-west-2.

Access from Iran blocked / no working route We re-read the supported-countries list for the OpenAI API on developers.openai.com on 11 September 2026; about 195 countries and territories are named and Iran is not among them. Read from the maker own page rather than measured on the network.

Measurement trail (2)

WebDev leaderboard 1,796 The gpt-6-astra-max row, rank 1 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes, 1,810 of them on this row. The row internal key carries codex-harness, so the model was tested inside OpenAI own coding agent rather than bare. Arena gives the row a rank range of 1 to 2, the same range as the row below it.

Score per dollar 35.9 WebDev Arena score divided by the price per million output tokens

3 Claude Opus 5

Tooling maturity 4/4 API, official app, command line tool, agent SDK, batch execution, and shipping on three major clouds

Access from Iran blocked / no working route Iran appears on neither of Anthropic two supported-countries lists, the API one or the claude.ai one. We read that on Anthropic own page rather than measuring it: an independent network test has not been run yet, and when it is, it will be written here.

Measurement trail (3)

WebDev leaderboard 1,691 The claude-opus-5-max row, rank 3 on the WebDev Arena leaderboard, read 23 September 2026 with a board total of 739,569 votes

Score per dollar 67.6 WebDev Arena score divided by the price per million output tokens

General text leaderboard 1,493 The claude-opus-5-high row, rank 10 on the Arena text leaderboard, updated 13 September 2026 with a board total of 8,146,274 votes. The score itself has not moved since 2 September; only its rank slipped one place as newer models entered above it.

4 Claude Fable 5

Tooling maturity 4/4 It carries the same infrastructure as Opus 5 and ships on all five routes: Anthropic own API, Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry.

Access from Iran blocked / no working route Iran is not on Anthropic supported-countries list. Read from Anthropic own page rather than measured on the network.

Measurement trail (3)

WebDev leaderboard 1,628 The claude-fable-5 row, rank 10 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes

Score per dollar 32.6 WebDev Arena score divided by the price per million output tokens

General text leaderboard 1,507 The claude-fable-5 row, rank 1 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

5 Muse Spark 1.3

Tooling maturity 2/4 The Meta Model API plus availability on OpenRouter, which is more than a bare endpoint. But Meta own documentation lists neither an official CLI nor an editor extension, so it does not go past the second step.

Measurement trail (2)

WebDev leaderboard 1,650 The muse-spark-1.3-max row, rank 8 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes. This row was added on 8 September. The xhigh row we recorded before now reads 1625 at rank 11; as with every other candidate, we record the higher row of the same model.

Score per dollar 388.2 WebDev Arena score divided by the price per million output tokens

6 GPT-5.6 Sol

Tooling maturity 4/4 It reaches the top step on the maker own documentation: API, SDKs and an official CLI, the Codex IDE extension, an Agents SDK, a batch lane, and a first-party guide to running on Amazon Bedrock that names the id openai.gpt-5.6-sol. This is the one column that filled up for this model without openai.com ever opening.

Access from Iran blocked / no working route This column was empty until today and is not any more. The supported-countries list lives on openai.com, which still returns 403, but the same list also sits on developers.openai.com, and that domain is open. About 195 countries and territories are named and Iran is not among them. The page itself says access from outside the list can get an account blocked. Read from the maker own page rather than measured on the network.

Measurement trail (3)

WebDev leaderboard 1,617 The gpt-5.6-sol-xhigh (codex-harness) row, rank 14 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes

Score per dollar 80.9 WebDev Arena score divided by the price per million output tokens

General text leaderboard 1,483 The gpt-5.6-sol-xhigh row, rank 17 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

7 Gemini 3.8 Flash

Tooling maturity 3/4 An API, official SDKs, a web studio and a Vertex route on Google Cloud. We hold back the fourth step because Google lists no separate official coding CLI for this family, unlike Anthropic and OpenAI.

Access from Iran blocked / no working route The Gemini API available-regions page names about 195 countries and territories and Iran is not among them. Read from Google own page rather than measured on the network.

Measurement trail (3)

WebDev leaderboard 1,568 The gemini-3.8-flash-high row, rank 22 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes

Score per dollar 418.1 WebDev Arena score divided by the price per million output tokens

General text leaderboard 1,494 The gemini-3.8-flash-high row, rank 8 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

8 Qwen3.8 Max

Measurement trail (3)

WebDev leaderboard 1,670 The qwen3.8-max row, rank 6 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes

Score per dollar 278.3 WebDev Arena score divided by the price per million output tokens

General text leaderboard 1,480 The qwen3.8-max row, rank 22 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

9 Kimi K3 Max

Tooling maturity 1/4 An OpenAI-compatible API plus one official client. The platform documentation lists no official command line tool or editor extension, so we do not score it higher.

Measurement trail (3)

WebDev leaderboard 1,674 The kimi-k3-max row, rank 5 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes

Score per dollar 111.6 WebDev Arena score divided by the price per million output tokens

General text leaderboard 1,489 The kimi-k3-max row, rank 12 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

10 GLM-5.2 Max

Tooling maturity 1/4 A documented API plus one official client. No official command line tool or editor extension is listed in the documentation.

Measurement trail (3)

WebDev leaderboard 1,589 The glm-5.2-max row, rank 18 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes

Score per dollar 361.1 WebDev Arena score divided by the price per million output tokens

General text leaderboard 1,472 The glm-5.2-max row, rank 37 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

11 Grok 4.5

Tooling maturity 3/4 An OpenAI-compatible API, Python and JavaScript SDKs, and an official command line agent called Grok Build. It does not reach the top step: there is no first-party editor extension, and the batch lane rejects this particular model.

Measurement trail (3)

WebDev leaderboard 1,556 The grok-4.5 row, rank 25 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes

Score per dollar 259.3 WebDev Arena score divided by the price per million output tokens

General text leaderboard 1,471 The grok-4.5 row, rank 38 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

12 Hy4 preview

Measurement trail (1)

WebDev leaderboard 1,623 The hy4-preview row, rank 13 on the WebDev Arena leaderboard, updated 8 September 2026 with 663,027 votes. The board prints Tencent as the maker and Apache 2.0 as the licence. The price shown beside the row was not read from any Tencent page, so it is not in this reference.

Not enough verified data to rank

These are real tools we track, but more than half of their criteria have no verified number yet, so ranking them would be a guess.

A white robotic arm typing on a black keyboard as glowing green code builds a translucent building in the air

How this ranking is calculated

Every criterion below has a weight and a source. Change a weight and the whole table recomputes. There is no hand-placed position anywhere in this hub.

Criterion Weight Evidence
WebDev leaderboard 40 WebDev ArenaReal web development tasks with tools over several steps, which is closer to daily work than a single-question exam
Score per dollar 20 computed from two numbers aboveFor a team carrying a monthly bill, a dearer model one percent ahead is a bad buy
Tooling maturity 15 defined scale, with a written reason per assignmentA good model with no command line tool and no agent SDK slows down in daily work
General text leaderboard 10 Arena (Text)Weighted low because general conversation is not a coding measure, though it is not irrelevant either
Context window 10 vendor stated specificationOn a large repository a small window means chunking, and chunking means losing context
Access from Iran 5 our access column, with its method statedWeighted low because this column has its own page under AI in Iran

Why Fable 5.1 leads, and why price did not stop it this time

One number settled it. In the 2 September update of WebDev Arena the claude-fable-5.1-max row sits first at 1765, and the next row we carry, claude-opus-5-max, comes in at 1687. Seventy-eight points, in a column whose whole top end fitted inside twenty-two points a week ago.

And this model loses the value column outright. Its output costs $50 per million tokens, twice Opus 5, and it scores zero there. Every previous time, that column has been what knocked an expensive model down. This is the first time a score gap has outweighed a price, and the reason is plain enough: it takes the full forty weight of the leaderboard column and also carries the tooling, context and Iran columns.

So the short answer is not one answer. If you work against the API and get a monthly bill, Opus 5 is half the price and second in this same table. If your work is the kind where getting it right once is cheaper than doing it twice, seventy-eight points is what you are paying for.

Where this table falls short

Our primary criterion is a human-preference leaderboard, not an automated test. WebDev Arena gives an absolute score, a date and a vote count, which is what publishing requires, but human preference also responds to readability and answer shape rather than only to whether the code is right. SWE-bench Verified would be a sharper measure, but its leaderboard renders in JavaScript and we have not read the figures directly. We do not print a number we have not read, even where someone else has quoted it.

One candidate on this page carries no rank: Gemini 3.1 Pro. It has a real score on the general text leaderboard, but we have not read its WebDev Arena row and Google publishes no context window for it at all, so more than half of the criteria weight is empty. Its name appears below the table; its position does not. GLM-5.2 Max and DeepSeek V4 Flash sat there too for a while, and now that their price and context window are verified from their own documentation, they carry a rank.

One thing about reading this leaderboard that came up this time and is not written down anywhere else: Qwen has two separate rows. There is qwen3.8-max at 1669 and qwen3.8-max-0902 at 1688, which is second on the whole board. We record the first, because that is the id in Alibaba own documentation and a dated row is a snapshot of one serving configuration. If you see Qwen listed second somewhere else, that is why, and not a mistake on our side.

One more thing, so the numbers do not confuse you: the scores are relative. Every column is normalised between its own minimum and maximum, so adding a new candidate moves everybody's score without any model having changed. Adding Grok 4.5 on 12 August 2026 did exactly that: the order held and the numbers rose. If today's figure does not match a screenshot from last week, that is why, and not because we touched anything.

Two more data changes moved the numbers on the same day. First, we read the text leaderboard in full this time and found that GLM-5.2 Max and DeepSeek V4 Flash had been on it all along; an earlier run had read only the top of the page. Both filled their general-text column. Second, GPT-5.6 Sol became an eighth candidate: openai.com still does not answer our server, but its developer documentation publishes the price and the context window, and its leaderboard row exists. One caveat about that row, since this is the coding page: it is named gpt-5.6-sol-xhigh (codex-harness), so the score belongs to the model inside OpenAI own coding agent rather than to the bare model. GPT landed between Kimi and Qwen, and no existing row changed place relative to the others.

28 August: DeepSeek got dearer and three rows moved

This is the only number in this reference that has ever gone up. On 12 August, DeepSeek V4 Flash cost $0.14 in and $0.28 out. Its own pricing page now carries two columns, peak and off peak, and even the cheaper column is above the old price: $0.22 and $0.66 off peak, $0.44 and $1.32 at peak. We record the peak column, because DeepSeek peak hours fall in the middle of the Iranian working day.

DeepSeek itself did not lose a place, because it is still the cheapest row in the table and still holds the top of the value column. But its lead narrowed, and since every column is normalised between its own minimum and maximum, that narrowing lifted three other rows: Qwen3.8 Max from fifth to third, Kimi K3 Max from third to fourth and GPT-5.6 Sol from fourth to fifth. None of those three models changed anything at all. That is what a relative ranking means, and it is why the arithmetic is printed on this page.

3 September: the gap at the top went from 22 points to 78

Until last week the three highest scores in this column sat inside 22 points, at 1692, 1674 and 1670. This page concluded then that the models differ by less than the numbers suggest, and that what genuinely changes a working day is tooling maturity: a command line tool, batch execution, presence on a cloud your company already has a contract with. For ranks two to eleven that still holds, and it is why the criterion carries fifteen points. For rank one it no longer does.

Three other rows also moved in the same update without changing anything about themselves. Kimi K3 Max went from third to fifth and Qwen3.8 Max from fifth to sixth, while Kimi leaderboard figure stayed at exactly 1674. GPT-5.6 Sol held fourth but its score fell from 51.6 to 48.8. All three are the arithmetic of new candidates entering and the columns normalising again.

Three new candidates, and one model deliberately not in the table

Muse Spark 1.3 from Meta and Gemini 3.8 Flash from Google joined the table today, both with a real row on the same leaderboard. Muse Spark carries a point other lists rarely make: Meta spent years being the maker that publishes its weights, and this one is not published. The model at the top of Meta line today is a closed one.

And the model that is not in the table: GPT-6 Astra, OpenAI new flagship at $10 in and $50 out. Neither leaderboard on this page has a row under that name, so seventy of its hundred weight would be empty and putting it in the table would produce one line reading "not ranked". It is in the model catalogue, and we wrote down here why it is not in the table so nobody assumes we missed it.

When not to take our first pick

  • If your team runs on a cloud that does not carry Opus 5, the second model with working tooling beats the first without it.
  • If your work is general conversation rather than code, ignore this table. The text leaderboard orders things differently.
  • If the budget is genuinely tight, the open-weight models lower down trail by less score than they trail in price.

Where it falls short

  • The primary criterion is a human-preference leaderboard, not an automated test against real repositories.
  • There is no SWE-bench Verified figure in the table, because we could not read that leaderboard directly.
  • Anthropic own Frontier-Bench claims are relative and carry no absolute number, so that column stays empty.
  • Qwen3.8 Max now has a verified price and context window, but its tool maturity and Iran columns still carry no evidence, so 20 of its 100 weight is unmeasured. That is a gap in our data, not a weakness in the model.
  • The model at the top of this table is also its dearest row and scores zero in the value column. If cost decides it for you, first place here is not your answer.
  • GPT-6 Astra, OpenAI new flagship, is not in the table at all, because neither leaderboard carries a row under that name. So this table does not currently measure the most expensive model OpenAI sells.

Our take

If you have to pick one and cost is not the constraint, take Fable 5.1. If you get a monthly bill, Opus 5 is half the price and seventy-eight points lower, and that is a trade you have to make rather than one we can make for you. And before changing model at all, look once at your own token usage: in the work we have seen, trimming context saves more than switching model does.

Questions people actually ask

What is the best AI for coding

Today, Claude Fable 5.1, on the strength of first place on WebDev Arena at 1765 and a seventy-eight point gap to the next row. It is also the dearest row in the table, so if cost decides it for you the practical answer is Opus 5. That answer carries a date and any new release can change it.

Are open-weight models usable for code

Yes, but the gap widened on 2 September. Kimi K3 Max at 1674 and Qwen3.8 Max at 1669 are still close to Opus 5 and about ninety points below the model at the top. One new thing is worth knowing: Meta Muse Spark, which joined the table today, does not publish its weights, so the maker best known for publishing them ships its own top model closed.

Why does this ranking differ from other lists

Because we wrote the criteria and the weights down where you can see them. Most lists never explain their order, and where there is no explanation, any order is possible.

If you would rather not spend the week wiring these tools into a project, that work is part of

our app and custom software service