For each job, which AI actually wins
We compute each ranking from stated criteria and print the arithmetic on the same page. A number with no source does not get printed, and every page carries its own last-checked date.
- computed, not opinionated
- every number sourced and dated
- an access-from-Iran column
How a ranking in this section is calculated
Three inputs go into one calculation and a table comes out. All three are visible on the page itself, and the fourth input is deliberately left unconnected.
-
1The criteria
Each category has its own criteria, and they are written on that same page.
-
2The weights
Every weight is visible. Change a weight and the table recomputes.
-
3A source per number
From the maker, or from the benchmark own leaderboard. Unverified stays unverified.
-
A cell with no source
Scoring zero and being unverified are two different things, and they stay two different things in the table. This input never enters the calculation and the cell stays empty.
-
The ranking calculation
Change a weight and the table recomputes. There is no manual position field in this section at all.
-
The ranking table
The arithmetic sits beside the table, plus an access from Iran column with the date it was checked.
The rankings here are temporary, and that is correct: any new release can move the table.
What do you want the AI to do
Each job has a ranking table whose criteria and weights are written on the page.
If you already know what you are after and just want the numbers: Every model, one table
Why these pages differ from the other lists
The arithmetic is on the page
Each job has its own criteria, each criterion a weight, and every weight is visible. Change a weight and the table recomputes. There is no "manual position" field in this section at all, so no ranking can move without a visible reason.
We write down what we do not know
If we could not verify a number from the vendor or from the benchmark own leaderboard, the cell stays empty and says "not verified". Scoring zero and being unverified are two different things, and they stay two different things in the table.
A column only we can write
For every tool we write whether it is reachable from Iran, whether a card works, and if not, what actually works instead. We check it ourselves from an Iranian connection and print the date we checked.
One page, four languages
Every page exists in Persian, English, Arabic and Turkish, and each language is written rather than translated. The facts and the numbers are identical across the four; the sentences are not.
The makers
Each maker has three axes: the products you sign in to, the model families, and the versions.
Latest changes
- Both leaderboards this reference ranks from were re-read, and the result is the largest single shift in its history. In the 2 September update of the WebDev board the claude-fable-5.1-max row went first at 1765, with the next row, claude-opus-5-max, at 1687: seventy-eight points, in a column whose whole top end fitted inside twenty-two points a week ago. So Opus 5 lost the top of the coding category after three weeks and its score fell from 77.5 to 62.7. The Fable 5.1 data file had written on 2 September that its categories were deliberately empty and would be filled the moment Arena rated it; that condition is now met, and the model entered the coding and chat categories with two evidence rows. On the chat board, though, Fable 5 held first at 1507 and 5.1 came second at 1504, three points between a retired model and its successor. All nine existing candidates were updated too: Qwen3.8 Max fell from 1490 to 1480 on the text board and dropped from fourth to eighth in chat, and Kimi K3 Max went from third to fifth in coding on an untouched 1674, because new candidates landed above it. Coding
- GPT-5.6 Sol got cheaper: input from $5 to $4 and output from $30 to $20 per million tokens, read from OpenAI own pricing page. It is the only model in this reference whose price has ever fallen rather than risen. Its Iran column filled at the same time, and the reason was a mistake of ours that this run corrected: we had written that the supported-countries list lives only on openai.com and that the domain returns 403 to our server. The domain still returns 403, but the same list also sits on developers.openai.com, which is open; about 195 countries and territories are named and Iran is not among them. Those ten Iran weights plus the price drop took the model from 54.2 to 65.2 in the chat category, though it still moved down a place to fourth because Fable 5.1 entered above it.
- Thirteen new models joined the catalogue and two of them entered the ranking tables. GPT-6 Astra is OpenAI new flagship at $10 in and $50 out with a 1.05M token window, and it is deliberately in no table because neither leaderboard carries a row under that name. Gemini 3.8 Flash is the first Google model to take a rank here, at 1494 on the text board, above the Opus 5 row. Muse Spark 1.3 and 1.2 from Meta were added too, and the point about them is that their weights are not published: the maker best known for years for publishing weights now ships its own top model closed. The rest: Kimi K2.7 Code, Grok 4.20, Veo 3.1 Lite, Lyria 3.5 Pro, Gemini 3.5 Transcribe, Eleven Music v2 and Scribe v2. The two transcription models are deliberately out of the audio category, because 65 of that rubric weight sits on Persian quality and voice control and both are about making speech rather than hearing it.
- The Claude Fable 5.1 page was written, for a model released on 1 September 2026. The finding: the base price did not move at all, $10 in and $50 out exactly as on Fable 5, and the only row that changed on Anthropic pricing table is the cache read, from one dollar to twenty-five cents. Anthropic calls that an exception in a footnote under that same table, because the cache multiplier is 0.1x the input price on every Claude model and 0.025x on Fable 5.1 and Mythos 5.1 alone. The benchmark jump is not uniform either, and that became the argument of the page: agentic scientific research goes from 24.7 to 52.6, terminal coding from 42.0 to 55.8, but CursorBench only from 70.5 to 73.4. The model appears in no ranking table on this site and that is a decision: the coding rubric puts 70 of its 100 points on two Arena leaderboards and a one-day-old model has not been rated on either, so placing it in the category would only print it below the table as unranked. Mythos 5.1 was added to the catalogue and Fable 5 moved to legacy. Claude Fable 5.1
A blunt word about this page
We built this because our own clients keep asking which tool to buy for a given job, and the answer changes every quarter. So the rankings here are temporary, and that is correct: any new release can move the table.
What you will not find here is an account for sale. We sell no subscriptions and no placement in any table. If something needs doing with these tools, the service for it is on the site and says so plainly.
Measured demand for this topic in the United States
AI image generation runs 823,000 searches a month and the click costs $0.85. Enormous attention, almost no commercial value per visitor. That single ratio explains why so many AI tools struggle to convert traffic into revenue.
| Query | Searches a month | Cost per click | Difficulty | Intent |
|---|---|---|---|---|
| ai image generator | 823,000 | $0.85 | 93 | Informational |
| ai video generator | 165,000 | $1.71 | 91 | Commercial |
| ai chatbot | 90,500 | $0.68 | 95 | Commercial |
| ai tools | 27,100 | $3.89 | 91 | Informational + Commercial |
| best ai tools | 6,600 | $4.57 | 73 | Commercial |
Source: Semrush, US database, pulled 27 August 2026. Volumes age; treat anything older than a quarter as a direction, not a number.