Best AI for coding
Claude Opus 5 scores highest today, because it sits first on WebDev Arena at 1692 while costing half of the dearest model on the market. The table below shows the criteria and their weights, and every number links to its own source.
- six models with real evidence
- weights on the page
- no hand-placed positions
Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. What changed
The ranking today, from verified evidence
A model that outranks us on the leaderboard appears in the table even where we have not written its page, or first place would only mean first among the ones we had time to write.
| Rank | Tool | Score | WebDev leaderboard 40 | Score per dollar 20 | Tooling maturity 15 | General text leaderboard 10 | Context window 10 | Access from Iran 5 |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 Current pick | 94.0 | 1,692 Source | 67.7 Source | 4/4 API, official app, command line tool, agent SDK, batch execution, and shipping on three major clouds | 1,494 Source | 1 million tokens Source | blocked / no working route Iran appears on neither of Anthropic two supported-countries lists, the API one or the claude.ai one. We read that on Anthropic own page rather than measuring it: an independent network test has not been run yet, and when it is, it will be written here. |
| 2 | Claude Fable 5 | 36.1 | 1,628 Source | 32.6 Source | not verified | 1,506 Source | 1 million tokens Source | not verified |
| 3 | Qwen3.8 Max | 33.8 | 1,670 Source | not verified | not verified | 1,490 Source | not verified | not verified |
| 4 | Kimi K3 Max | 33.8 | 1,674 Source | not verified | not verified | 1,487 Source | not verified | not verified |
An empty cell means we could not verify that number, not that the tool scored zero.
Not enough verified data to rank
These are real tools we track, but more than half of their criteria have no verified number yet, so ranking them would be a guess.
- DeepSeek V4 Flash
- Gemini 3.1 Pro
- GLM-5.2 Max
How this ranking is calculated
Every criterion below has a weight and a source. Change a weight and the whole table recomputes. There is no hand-placed position anywhere in this hub.
| Criterion | Weight | Evidence |
|---|---|---|
| WebDev leaderboard | 40 | WebDev ArenaReal web development tasks with tools over several steps, which is closer to daily work than a single-question exam |
| Score per dollar | 20 | computed from two numbers aboveFor a team carrying a monthly bill, a dearer model one percent ahead is a bad buy |
| Tooling maturity | 15 | defined scale, with a written reason per assignmentA good model with no command line tool and no agent SDK slows down in daily work |
| General text leaderboard | 10 | Arena (Text)Weighted low because general conversation is not a coding measure, though it is not irrelevant either |
| Context window | 10 | vendor stated specificationOn a large repository a small window means chunking, and chunking means losing context |
| Access from Iran | 5 | our access column, with its method statedWeighted low because this column has its own page under AI in Iran |
Why Opus 5 leads
Two things together. First, on WebDev Arena, which measures real web development work with tools across several steps, the claude-opus-5-max row sits first at 1692, ahead of Kimi K3 Max at 1674 and Qwen3.8 Max at 1670. Second, it charges $25 per million output tokens, half of Fable 5, which itself scores lower at 1628. A model that is both higher and cheaper comes first under any sensible weighting.
Where this table falls short
Our primary criterion is a human-preference leaderboard, not an automated test. WebDev Arena gives an absolute score, a date and a vote count, which is what publishing requires, but human preference also responds to readability and answer shape rather than only to whether the code is right. SWE-bench Verified would be a sharper measure, but its leaderboard renders in JavaScript and we have not read the figures directly. We do not print a number we have not read, even where someone else has quoted it.
Two models on this page carry no rank: GLM-5.2 Max and DeepSeek V4 Flash. Both have a real WebDev Arena score, but we have not verified their price or context window from their own documentation, so more than half of the criteria weight is empty for them. Their names appear below the table; their positions do not.
Something that only shows up in daily work
The gap between first and fourth in this table is about twenty-six points, which is felt far less in practice than the number suggests. What genuinely changes a working day is tooling maturity: whether the model has a command line tool, batch execution, and presence on a cloud your company already has a contract with. That is why the criterion carries fifteen points and not five.
When not to take our first pick
- If your team runs on a cloud that does not carry Opus 5, the second model with working tooling beats the first without it.
- If your work is general conversation rather than code, ignore this table. The text leaderboard orders things differently.
- If the budget is genuinely tight, the open-weight models lower down trail by less score than they trail in price.
Where it falls short
- The primary criterion is a human-preference leaderboard, not an automated test against real repositories.
- There is no SWE-bench Verified figure in the table, because we could not read that leaderboard directly.
- Anthropic own Frontier-Bench claims are relative and carry no absolute number, so that column stays empty.
- The Chinese models here (Kimi, Qwen, GLM, DeepSeek) have no verified price in our data, and that alone holds their rank down. That is a gap in our data, not a weakness in them.
Our take
If you have to pick one and Iran is not a constraint, take Opus 5. If cost is the constraint, look at your own token usage before changing model: in the work we have seen, trimming context saves more than switching model does.
Questions people actually ask
What is the best AI for coding
Today, Claude Opus 5, on the strength of first place on WebDev Arena and a price half that of the dearest model on the market. That answer carries a date and any new release can change it.
Are open-weight models usable for code
Yes. On the same leaderboard Qwen3.8 Max scores 1670 and Kimi K3 Max 1674, both close behind the leader. We hold their rank down because we have not verified their price and specs from their own documentation, not because they perform badly.
Why does this ranking differ from other lists
Because we wrote the criteria and the weights down where you can see them. Most lists never explain their order, and where there is no explanation, any order is possible.
If you would rather not spend the week wiring these tools into a project, that work is part of our app and custom software service