Best AI chatbot
For open conversation, Claude Fable 5 scores highest today: 1506 on the Arena text leaderboard, first of 389 models, with Opus 5 twelve points behind at half the price. But this table ranks models rather than apps, and that single decision is why ChatGPT has no row in it.
- nine models, all of them ranked
- weights on the page
- models ranked, not apps
Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. What changed
The ranking today, from a leaderboard that publishes a number, a date and a vote count
No candidate sits below this table unranked. An empty cell means the vendor published nothing, not that the model scored zero.
| Rank | Tool | Score | General text leaderboard 45 | Score per dollar 20 | Context window 15 | Tooling maturity 10 | Access from Iran 10 |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 Current pick | 78.6 | 1,506 Source | 30.1 Source | 1 million tokens Source | 4/4 It carries the same infrastructure as Opus 5 and ships on all five routes: Anthropic own API, Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. | blocked / no working route Iran is not on Anthropic supported-countries list. Read from Anthropic own page rather than measured on the network. |
| 2 | Claude Opus 5 | 71.1 | 1,494 Source | 59.8 Source | 1 million tokens Source | 4/4 API, official app, command line tool, agent SDK, batch execution, and shipping on three major clouds | blocked / no working route Iran appears on neither of Anthropic two supported-countries lists, the API one or the claude.ai one. We read that on Anthropic own page rather than measuring it: an independent network test has not been run yet, and when it is, it will be written here. |
| 3 | GPT-5.6 Sol | 54.2 | 1,481 Source | 49.4 Source | 1 million tokens Source | 4/4 It reaches the top step on the maker own documentation: API, SDKs and an official CLI, the Codex IDE extension, an Agents SDK, a batch lane, and a first-party guide to running on Amazon Bedrock that names the id openai.gpt-5.6-sol. This is the one column that filled up for this model without openai.com ever opening. | not verified |
| 4 | Qwen3.8 Max | 49.4 | 1,490 Source | 248.3 Source | 1 million tokens Source | not verified | not verified |
| 5 | Kimi K3 Max | 48.2 | 1,487 Source | 99.1 Source | 1 million tokens Source | 1/4 An OpenAI-compatible API plus one official client. The platform documentation lists no official command line tool or editor extension, so we do not score it higher. | not verified |
| 6 | Gemini 3.1 Pro | 42.7 | 1,486 Source | 123.8 Source | not verified | not verified | blocked / no working route The Gemini API available-regions page lists the countries where the API works, and Iran is not on that list. We read it there rather than measuring it: an independent test has not been run yet. |
| 7 | GLM-5.2 Max | 37.0 | 1,470 Source | 334.1 Source | 1 million tokens Source | 1/4 A documented API plus one official client. No official command line tool or editor extension is listed in the documentation. | not verified |
| 8 | DeepSeek V4 Flash | 33.6 | 1,435 Source | 5,125.0 Source | 1 million tokens Source | 1/4 A documented API plus one official client. No official agent tooling or editor extension is listed. | not verified |
| 9 | Grok 4.5 | 29.1 | 1,469 Source | 244.8 Source | 500k tokens Source | 3/4 An OpenAI-compatible API, Python and JavaScript SDKs, and an official command line agent called Grok Build. It does not reach the top step: there is no first-party editor extension, and the batch lane rejects this particular model. | not verified |
An empty cell means we could not verify that number, not that the tool scored zero.
How this ranking is calculated
Every criterion below has a weight and a source. Change a weight and the whole table recomputes. There is no hand-placed position anywhere in this hub.
| Criterion | Weight | Evidence |
|---|---|---|
| General text leaderboard | 45 | Arena (Text)Human preference in open conversation. For chat it is the closest thing to a right measure that exists, which is why it carries the most weight |
| Score per dollar | 20 | computed from two numbers aboveIf you are building chat on an API, twice the price for one percent more score is a bad trade. If you are buying an app subscription, this column does not touch you |
| Context window | 15 | vendor stated specificationA long conversation loses its memory first and gets expensive second, once the window fills |
| Tooling maturity | 10 | defined scale, with a written reason per assignmentA model with no official app, no SDK and no cloud route falls behind a slightly weaker one in daily work |
| Access from Iran | 10 | our access column, with its method statedTen of a hundred, because this column has a page of its own. Today only three of the nine rows carry any evidence in it |
Why Fable 5 leads, and why that is arguable
One number and one condition. The number: Fable 5 sits first of 389 models on the Arena text leaderboard at 1506, and takes the whole forty-five-weight column. The condition: the same model loses the score-per-dollar column outright. Its output costs $50 per million tokens, the most expensive row in the table, which puts its score-per-dollar at 30.1, the floor.
Opus 5 is twelve points lower and half the price. On a leaderboard whose entire top tier fits inside about fifty points, twelve is not much. So the real answer splits by question: if you are building chat on an API and paying a monthly bill, Opus 5; if you are buying an app subscription, the price column does not apply to you and the top of the table is what you want.
One thing against our own leader, which no column here measures: Fable 5 is the only model Anthropic own table marks as slower, and its adaptive thinking is always on, so there is no lever to turn it down. In conversation, latency is felt. Our table does not see that and you will inside ten minutes of using it.
This table ranks models, not apps
That is a deliberate choice with a cost, so let us be direct about it. All five columns here are numbers published about a model: a leaderboard score, a price per million tokens, a context window. No vendor publishes any of the three for its chat app. Rank apps instead and three of the five columns go empty for most rows, which stops being a table.
The second reason matters more: one model runs under several apps, and an app swaps the model underneath it without telling you. Opus 5 sits under the Claude app, the API and Claude Code at the same time. Make "Claude" a row and we count the same number again in the coding table. It is the same reason the video category ranks versions rather than tools.
We made the opposite call for studying and it was right there, because Google publishes the notebook own limits as figures: fifty sources, five hundred thousand words. For a chat app no vendor publishes anything comparable. The cost is that Claude, Gemini and ChatGPT as apps have no row here, and their own pages are where those questions get answered.
Why ChatGPT has no row, and what we do have from OpenAI
The short answer is that ChatGPT is a product and this table ranks models, so it was never going to be a row. The longer answer counts against us and matters more: even if we wanted to, we could not describe it. Every path on openai.com returns 403 to this server, re-tested on 12 August 2026, while that domain own robots.txt does not forbid crawling.
One thing is open, though, and until today we had not used it: developers.openai.com answers in full. GPT-5.6 Sol is listed there at $5 input and $30 output, a 1.05M token context window, 128K max output and a knowledge cutoff of 16 February 2026. Its Arena rows exist too. So today it is the third row of this very table.
The summary is one sentence: the model can be ranked, the app cannot be described. What we still cannot tell you is which model runs under ChatGPT today, what its plans cost, and whether it opens from Iran. All three live on the domain that will not answer. The OpenAI page writes that boundary out in full.
What reading the leaderboard properly taught us
Two rows we had recorded as unverified had been on the board the whole time. GLM-5.2 Max is 33rd at 1470 and DeepSeek V4 Flash is 83rd at 1435. An earlier run read the top of the page; this one read the JSON the page itself ships. So both candidates that could not be ranked in this category now can be, and the coding table moved as well. That gap was ours, not the vendors.
The second concerns what a row is named. This reference already argued that an Arena row names a serving configuration rather than a model. For OpenAI it goes one step further: the WebDev row reads, literally, gpt-5.6-sol-xhigh (codex-harness). That score belongs to the model inside OpenAI own coding agent, not to the bare model. Price and context attach to the base model; the score attaches to whatever was actually tested.
The third was a trap we did not fall into. The same board carries a row called deepseek-v4-pro at 1458, sitting above Flash. That is a different model in the same family, not a setting, so we did not borrow its number. Had we done so, DeepSeek would have climbed two places and nobody would have noticed.
Where our own method exaggerates
All nine rows fit inside 71 Elo points, from 1506 down to 1435. Min-max normalisation stretches those 71 points into a spread from 78.6 to 29.1. So the score column makes the distance look larger than it is, and the last row is not weak: it is last in a table of high-end models.
A sharper example, because it surprised us too. GPT-5.6 Sol wins the context column with 1,050,000 tokens and Kimi K3 Max holds 1,048,576. The gap is 1,424 tokens, about a tenth of a percent, and in practice it changes nothing. Both take essentially the full column, which is correct. But add a model with a two-million window tomorrow and both lose that column together, without anything about either having changed.
When not to take our first pick
- If you want an app rather than a model, this table is not your answer. The Claude and Gemini product pages answer that question at the right level.
- If your work is code, the order here differs from the coding table and the criteria there are entirely different.
- If your workload is high volume and short answers, ignore the first column and read score per dollar, where DeepSeek leads by a wide margin.
- If you are reaching chat from Iran with no official route, the order of these models is the last thing that bears on your problem.
Where it falls short
- The main criterion is a human preference leaderboard: it measures whose answers people prefer, not which answers are correct.
- The whole table fits inside 71 Elo points and normalisation stretches that into 78.6 down to 29.1, so the score column overstates the real distance.
- Six of the nine rows carry no evidence in the Iran column, because their vendors publish no country list at all. Only Anthropic and Google do.
- ChatGPT, the Claude app and the Gemini app have no row here, because the table ranks models rather than apps.
- Qwen3.8 Max and Gemini 3.1 Pro each have two empty columns. That is a gap in what the vendor publishes rather than a weakness in the model, but it genuinely holds their scores down.
- DeepSeek at $0.28 output stretches the score-per-dollar column so hard that the fourfold gap between Opus 5 and Qwen turns into less than one point in it.
Our take
If you are building chat on an API and have to pick one, Opus 5: twelve points below Fable 5 at half its price. If you are buying an app subscription, this table price column does not apply to you and the top of the table is what to read. And if you are asking which of them will answer your customers, our honest answer is none of them, because that is a different job.
Questions people actually ask
What is the best AI chatbot
On the Arena text leaderboard today, Claude Fable 5 leads at 1506. But Opus 5 sits at 1494 for half the price, and on a board whose top tier fits inside fifty points those twelve are hard to feel in practice.
Why is ChatGPT not in this table
Because the table ranks models and ChatGPT is a product. Its model is here: GPT-5.6 Sol is the third row. The app itself we cannot describe, because every path on openai.com returns 403 to our server.
Why does this order differ from the coding table
Because the criteria differ. There the WebDev leaderboard carries forty of a hundred; here the general text leaderboard carries forty-five. A model that writes better code does not necessarily give answers people prefer, and Opus 5 and Fable 5 swap places between exactly those two tables.
Which of these work from Iran
None of them has an official route. Anthropic and Google publish country lists and Iran is on neither; the other vendors publish no list at all, which is why those cells are empty. We neither sell accounts nor recommend circumvention.
The model at the top of this table answers you, not your customers. For the second job we send a person rather than a bot, and that is exactly what our virtual assistant service