Job

Best AI chatbot

For open conversation, Claude Fable 5 still holds the highest Arena text score at 1507, and that is the odd part today: Anthropic moved that same model to legacy on 1 September, and its successor sits directly below it three points back. The table ranks models rather than apps, which is why ChatGPT has no row in it.

  • twelve models, all of them ranked
  • weights on the page
  • models ranked, not apps

Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. What changed

Where each tool's score comes from

The ring below is the same set of weights printed in the table headers further down.

100Weights
  • General text leaderboard 45
  • Score per dollar 20
  • Context window 15
  • Tooling maturity 10
  • Access from Iran 10

We chose these weights, and that is the only judgment call in the table. Weight them differently and the order changes.

The ranking today, from a leaderboard that publishes a number, a date and a vote count

No candidate sits below this table unranked. An empty cell means the vendor published nothing, not that the model scored zero.

Best AI chatbot
Rank Tool Score General text leaderboard 45 Score per dollar 20 Context window 15 Tooling maturity 10 Access from Iran 10
1 Claude Fable 5 Current pick the maker has replaced it with Claude Fable 5.1 78.6 1,507 Source: arena.ai 30.1 Source: arena.ai 1 million tokens Source: platform.claude.com 4/4 blocked / no working route
2 Claude Fable 5.1 76.7 1,504 Source: arena.ai 30.1 Source: arena.ai 1 million tokens Source: platform.claude.com 4/4 blocked / no working route
3 Claude Opus 5 70.1 1,493 Source: arena.ai 59.7 Source: arena.ai 1 million tokens Source: platform.claude.com 4/4 blocked / no working route
4 GPT-5.6 Sol 65.2 1,483 Source: arena.ai 74.2 Source: arena.ai 1 million tokens Source: developers.openai.com 4/4 blocked / no working route
5 Muse Spark 1.2 62.8 1,499 Source: arena.ai 352.7 Source: arena.ai 1 million tokens Source: developer.meta.com 2/4 not verified
6 Gemini 3.8 Flash 60.1 1,494 Source: arena.ai 398.4 Source: arena.ai not verified 3/4 blocked / no working route
7 Kimi K3 Max 49.5 1,489 Source: arena.ai 99.3 Source: arena.ai 1 million tokens Source: platform.kimi.ai 1/4 not verified
8 Qwen3.8 Max 45.1 1,480 Source: arena.ai 246.7 Source: arena.ai 1 million tokens Source: www.alibabacloud.com not verified not verified
9 Gemini 3.1 Pro 43.7 1,487 Source: arena.ai 123.9 Source: arena.ai not verified not verified blocked / no working route
10 GLM-5.2 Max 41.6 1,472 Source: arena.ai 334.5 Source: arena.ai 1 million tokens Source: docs.z.ai 1/4 not verified
11 DeepSeek V4 Flash 33.6 1,438 Source: arena.ai 1,089.4 Source: arena.ai 1 million tokens Source: api-docs.deepseek.com 1/4 not verified
12 Grok 4.5 32.2 1,471 Source: arena.ai 245.2 Source: arena.ai 500k tokens Source: docs.x.ai 3/4 not verified

An empty cell means we could not verify that number, not that the tool scored zero.

Behind each number

Every judged score in this table carries a written reason, and the Iran column says how it was checked. Those two are open here. The measurement trail behind each number, which row of which leaderboard and on how many votes, opens under the model it belongs to.

1 Claude Fable 5

Tooling maturity 4/4 It carries the same infrastructure as Opus 5 and ships on all five routes: Anthropic own API, Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry.

Access from Iran blocked / no working route Iran is not on Anthropic supported-countries list. Read from Anthropic own page rather than measured on the network.

Measurement trail (2)

General text leaderboard 1,507 The claude-fable-5 row, rank 1 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 30.1 Arena (Text) score divided by the price per million output tokens

2 Claude Fable 5.1

Tooling maturity 4/4 It ships on all five routes and Anthropic own table prints a model ID for each. In that same table the Claude Platform on AWS cell is a dash for Opus 5 and an ID for this model.

Access from Iran blocked / no working route Iran is not on Anthropic supported-countries list. Read from Anthropic own page rather than measured on the network. Nothing changed for this version.

Measurement trail (2)

General text leaderboard 1,504 The claude-fable-5.1-max row, rank 3 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 30.1 Arena (Text) score divided by the price per million output tokens

3 Claude Opus 5

Tooling maturity 4/4 API, official app, command line tool, agent SDK, batch execution, and shipping on three major clouds

Access from Iran blocked / no working route Iran appears on neither of Anthropic two supported-countries lists, the API one or the claude.ai one. We read that on Anthropic own page rather than measuring it: an independent network test has not been run yet, and when it is, it will be written here.

Measurement trail (2)

General text leaderboard 1,493 The claude-opus-5-high row, rank 9 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 59.7 Arena (Text) score divided by the price per million output tokens

4 GPT-5.6 Sol

Tooling maturity 4/4 It reaches the top step on the maker own documentation: API, SDKs and an official CLI, the Codex IDE extension, an Agents SDK, a batch lane, and a first-party guide to running on Amazon Bedrock that names the id openai.gpt-5.6-sol. This is the one column that filled up for this model without openai.com ever opening.

Access from Iran blocked / no working route This column was empty until today and is not any more. The supported-countries list lives on openai.com, which still returns 403, but the same list also sits on developers.openai.com, and that domain is open. About 195 countries and territories are named and Iran is not among them. The page itself says access from outside the list can get an account blocked. Read from the maker own page rather than measured on the network.

Measurement trail (2)

General text leaderboard 1,483 The gpt-5.6-sol-xhigh row, rank 17 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 74.2 Arena (Text) score divided by the price per million output tokens

5 Muse Spark 1.2

Tooling maturity 2/4 The same routes as 1.3: the Meta Model API and OpenRouter, with no official CLI and no editor extension in Meta own documentation.

Measurement trail (2)

General text leaderboard 1,499 The muse-spark-1.2 (xHigh) row, rank 5 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 352.7 Arena (Text) score divided by the price per million output tokens

6 Gemini 3.8 Flash

Tooling maturity 3/4 An API, official SDKs, a web studio and a Vertex route on Google Cloud. We hold back the fourth step because Google lists no separate official coding CLI for this family, unlike Anthropic and OpenAI.

Access from Iran blocked / no working route The Gemini API available-regions page names about 195 countries and territories and Iran is not among them. Read from Google own page rather than measured on the network.

Measurement trail (2)

General text leaderboard 1,494 The gemini-3.8-flash-high row, rank 8 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 398.4 Arena (Text) score divided by the price per million output tokens

7 Kimi K3 Max

Tooling maturity 1/4 An OpenAI-compatible API plus one official client. The platform documentation lists no official command line tool or editor extension, so we do not score it higher.

Measurement trail (2)

General text leaderboard 1,489 The kimi-k3-max row, rank 12 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 99.3 Arena (Text) score divided by the price per million output tokens

8 Qwen3.8 Max

Measurement trail (2)

General text leaderboard 1,480 The qwen3.8-max row, rank 22 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 246.7 Arena (Text) score divided by the price per million output tokens

9 Gemini 3.1 Pro

Access from Iran blocked / no working route The Gemini API available-regions page lists the countries where the API works, and Iran is not on that list. We read it there rather than measuring it: an independent test has not been run yet.

Measurement trail (2)

General text leaderboard 1,487 The gemini-3.1-pro-preview row, rank 15 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 123.9 Arena (Text) score divided by the price per million output tokens

10 GLM-5.2 Max

Tooling maturity 1/4 A documented API plus one official client. No official command line tool or editor extension is listed in the documentation.

Measurement trail (2)

General text leaderboard 1,472 The glm-5.2-max row, rank 37 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 334.5 Arena (Text) score divided by the price per million output tokens

11 DeepSeek V4 Flash

Tooling maturity 1/4 A documented API plus one official client. No official agent tooling or editor extension is listed.

Measurement trail (2)

General text leaderboard 1,438 The deepseek-v4-flash-high-preview row, rank 85 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 1,089.4 Arena (Text) score divided by the price per million output tokens

12 Grok 4.5

Tooling maturity 3/4 An OpenAI-compatible API, Python and JavaScript SDKs, and an official command line agent called Grok Build. It does not reach the top step: there is no first-party editor extension, and the batch lane rejects this particular model.

Measurement trail (2)

General text leaderboard 1,471 The grok-4.5 row, rank 38 on the Arena text leaderboard, updated 2 September 2026 with 7,999,020 votes

Score per dollar 245.2 Arena (Text) score divided by the price per million output tokens

How this ranking is calculated

Every criterion below has a weight and a source. Change a weight and the whole table recomputes. There is no hand-placed position anywhere in this hub.

Criterion Weight Evidence
General text leaderboard 45 Arena (Text)Human preference in open conversation. For chat it is the closest thing to a right measure that exists, which is why it carries the most weight
Score per dollar 20 computed from two numbers aboveIf you are building chat on an API, twice the price for one percent more score is a bad trade. If you are buying an app subscription, this column does not touch you
Context window 15 vendor stated specificationA long conversation loses its memory first and gets expensive second, once the window fills
Tooling maturity 10 defined scale, with a written reason per assignmentA model with no official app, no SDK and no cloud route falls behind a slightly weaker one in daily work
Access from Iran 10 our access column, with its method statedTen of a hundred, because this column has a page of its own. Today only three of the nine rows carry any evidence in it

Why Fable 5 leads, and why that is stranger than usual this time

One number and one contradiction. The number: in the 2 September update of the Arena text leaderboard Fable 5 sits first at 1507 and takes the whole forty-five-weight column. The contradiction: Anthropic moved that same model to legacy on 1 September. The top of this table today is a model its own maker no longer recommends.

Its successor is directly underneath it. Fable 5.1 is second at 1504, three points back, at the same price and the same context window. Three points on this board is not something anyone feels in conversation. So the ranking is telling the truth and the practical advice is a different thing, and we show both: under the Fable 5 name in this very table we print which model replaces it. Being superseded is a fact, not a score, and the rubric has no way to score it and should not have one.

Both of them lose the score-per-dollar column outright. Output on either costs $50 per million tokens, the dearest rows in the table, which puts their score-per-dollar at 30.1, the floor. Opus 5 is fourteen points lower at 1493 and half the price. So the real answer splits by question: if you are building chat on an API and paying a monthly bill, Opus 5; if you are buying an app subscription the price column does not apply to you and the top of the table is what you want, with the caveat that you should take 5.1 rather than 5.

One thing against our own leader, which no column here measures: Fable 5 is the only model Anthropic own table marks as slower, and its adaptive thinking is always on, so there is no lever to turn it down. In conversation, latency is felt. Our table does not see that and you will inside ten minutes of using it.

This table ranks models, not apps

That is a deliberate choice with a cost, so let us be direct about it. All five columns here are numbers published about a model: a leaderboard score, a price per million tokens, a context window. No vendor publishes any of the three for its chat app. Rank apps instead and three of the five columns go empty for most rows, which stops being a table.

The second reason matters more: one model runs under several apps, and an app swaps the model underneath it without telling you. Opus 5 sits under the Claude app, the API and Claude Code at the same time. Make "Claude" a row and we count the same number again in the coding table. It is the same reason the video category ranks versions rather than tools.

We made the opposite call for studying and it was right there, because Google publishes the notebook own limits as figures: fifty sources, five hundred thousand words. For a chat app no vendor publishes anything comparable. The cost is that Claude, Gemini and ChatGPT as apps have no row here, and their own pages are where those questions get answered.

Why ChatGPT has no row, and what we do have from OpenAI

The short answer is that ChatGPT is a product and this table ranks models, so it was never going to be a row. The longer answer counts against us and matters more: even if we wanted to, we could not describe it. Every path on openai.com returns 403 to this server, re-tested on 12 August 2026, while that domain own robots.txt does not forbid crawling.

One thing is open, though: developers.openai.com answers in full. GPT-5.6 Sol is listed there at $4 input and $20 output, a 1.05M token context window, 128K max output and a knowledge cutoff of 16 February 2026. That price was $5 and $30 as recently as 12 August, which makes it the only model in this table that has got cheaper.

And one thing we corrected in this run, at the expense of something we had written ourselves. We had said the OpenAI supported-countries list lives only on openai.com and that the domain will not answer us. Half of that was right: the domain still returns 403, but the same list also sits on developers.openai.com, and that one is open. It names about 195 countries and territories, Iran is not among them, and the page itself says access from outside the list can get an account blocked. So the Iran column on this row is filled as of today.

Those ten weights, plus the price drop, took GPT-5.6 Sol from 54.2 to 65.2. It still moved down a place, to fourth, because Fable 5.1 entered the table above it. That is exactly what a relative score means, and it is why the arithmetic for every row is printed on this page.

The summary is one sentence: the model can be ranked, the app cannot be described. What we still cannot tell you is which model runs under ChatGPT today and what its plans cost. Both live on the domain that will not answer. The OpenAI page writes that boundary out in full.

What reading the leaderboard properly taught us

Two rows we had recorded as unverified had been on the board the whole time. GLM-5.2 Max is 37th at 1472 and DeepSeek V4 Flash is 85th at 1438. An earlier run read the top of the page; this one read the JSON the page itself ships. So both candidates that could not be ranked in this category now can be, and the coding table moved as well. That gap was ours, not the vendors.

The second concerns what a row is named. This reference already argued that an Arena row names a serving configuration rather than a model. For OpenAI it goes one step further: the WebDev row reads, literally, gpt-5.6-sol-xhigh (codex-harness). That score belongs to the model inside OpenAI own coding agent, not to the bare model. Price and context attach to the base model; the score attaches to whatever was actually tested.

The third was a trap we did not fall into. The same board carries a row called deepseek-v4-pro at 1458, sitting above Flash. That is a different model in the same family, not a setting, so we did not borrow its number. Had we done so, DeepSeek would have climbed two places and nobody would have noticed.

Where our own method exaggerates

All twelve rows fit inside 69 Elo points, from 1507 down to 1438. Min-max normalisation stretches those 69 points into a spread from 78.6 to 32.2. So the score column makes the distance look larger than it is, and the last row is not weak: it is last in a table of high-end models.

Three new candidates joined the table today and not one of them had changed anything about itself. Fable 5.1 second, Meta Muse Spark 1.2 fifth and Google Gemini 3.8 Flash sixth. Their arrival pushed Qwen3.8 Max from fourth to eighth, while the only real thing that happened to Qwen is that its figure on that same leaderboard moved from 1490 to 1480. If you noted the order down somewhere yesterday, that is the explanation for the difference.

A sharper example, because it surprised us too. GPT-5.6 Sol wins the context column with 1,050,000 tokens and Kimi K3 Max holds 1,048,576. The gap is 1,424 tokens, about a tenth of a percent, and in practice it changes nothing. Both take essentially the full column, which is correct. But add a model with a two-million window tomorrow and both lose that column together, without anything about either having changed.

When not to take our first pick

  • If you want an app rather than a model, this table is not your answer. The Claude and Gemini product pages answer that question at the right level.
  • If your work is code, the order here differs from the coding table and the criteria there are entirely different.
  • If your workload is high volume and short answers, ignore the first column and read score per dollar, where DeepSeek leads by a wide margin.
  • If you are reaching chat from Iran with no official route, the order of these models is the last thing that bears on your problem.

Where it falls short

  • The main criterion is a human preference leaderboard: it measures whose answers people prefer, not which answers are correct.
  • The whole table fits inside 69 Elo points and normalisation stretches that into 78.6 down to 32.2, so the score column overstates the real distance.
  • Six of the twelve rows carry no evidence in the Iran column, because their vendors publish no country list at all. Only Anthropic, Google and OpenAI do.
  • ChatGPT, the Claude app and the Gemini app have no row here, because the table ranks models rather than apps.
  • Qwen3.8 Max and Gemini 3.1 Pro each have two empty columns. That is a gap in what the vendor publishes rather than a weakness in the model, but it genuinely holds their scores down.
  • DeepSeek at $1.32 output stretches the score-per-dollar column so hard that the fourfold gap between Opus 5 and Qwen turns into less than one point in it.
  • The top of the table is a retired model. The ranking measures score and does not measure succession, so we show that as a note under the model name rather than as points.

Our take

If you are building chat on an API and have to pick one, Opus 5: fourteen points below the top of the table at half its price. If you are buying an app subscription the price column does not apply to you and the top of the table is what to read, but take Fable 5.1 rather than Fable 5; three points is a bad trade against a model its maker has retired. And if you are asking which of them will answer your customers, our honest answer is none of them, because that is a different job.

Questions people actually ask

What is the best AI chatbot

On the Arena text leaderboard Claude Fable 5 leads at 1507, but Anthropic has retired that model. Its successor Fable 5.1 is second at 1504 and is the one to take. If cost matters, Opus 5 at 1493 is half the price of either.

Why is ChatGPT not in this table

Because the table ranks models and ChatGPT is a product. Its model is here: GPT-5.6 Sol is the fourth row. The app itself we cannot describe, because every path on openai.com returns 403 to our server. The OpenAI country list, though, was read in this run, from the developer domain.

Why does this order differ from the coding table

Because the criteria differ. There the WebDev leaderboard carries forty of a hundred; here the general text leaderboard carries forty-five. A model that writes better code does not necessarily give answers people prefer, and Opus 5 and Fable 5 swap places between exactly those two tables.

Which of these work from Iran

None of them has an official route. Anthropic, Google and, as of this run, OpenAI all publish country lists and Iran is on none of them; the other vendors publish no list at all, which is why those cells are empty. We neither sell accounts nor recommend circumvention.

The model at the top of this table answers you, not your customers. For the second job we send a person rather than a bot, and that is exactly what

our virtual assistant service