What changed
Every change to this hub is logged here with its date and its source, including changes to the ranking weights.
-
ranking moved Coding
Both leaderboards this reference ranks from were re-read, and the result is the largest single shift in its history. In the 2 September update of the WebDev board the claude-fable-5.1-max row went first at 1765, with the next row, claude-opus-5-max, at 1687: seventy-eight points, in a column whose whole top end fitted inside twenty-two points a week ago. So Opus 5 lost the top of the coding category after three weeks and its score fell from 77.5 to 62.7. The Fable 5.1 data file had written on 2 September that its categories were deliberately empty and would be filled the moment Arena rated it; that condition is now met, and the model entered the coding and chat categories with two evidence rows. On the chat board, though, Fable 5 held first at 1507 and 5.1 came second at 1504, three points between a retired model and its successor. All nine existing candidates were updated too: Qwen3.8 Max fell from 1490 to 1480 on the text board and dropped from fourth to eighth in chat, and Kimi K3 Max went from third to fifth in coding on an untouched 1674, because new candidates landed above it.
Source: arena.ai
-
price
GPT-5.6 Sol got cheaper: input from $5 to $4 and output from $30 to $20 per million tokens, read from OpenAI own pricing page. It is the only model in this reference whose price has ever fallen rather than risen. Its Iran column filled at the same time, and the reason was a mistake of ours that this run corrected: we had written that the supported-countries list lives only on openai.com and that the domain returns 403 to our server. The domain still returns 403, but the same list also sits on developers.openai.com, which is open; about 195 countries and territories are named and Iran is not among them. Those ten Iran weights plus the price drop took the model from 54.2 to 65.2 in the chat category, though it still moved down a place to fourth because Fable 5.1 entered above it.
Source: developers.openai.com
-
new page
Thirteen new models joined the catalogue and two of them entered the ranking tables. GPT-6 Astra is OpenAI new flagship at $10 in and $50 out with a 1.05M token window, and it is deliberately in no table because neither leaderboard carries a row under that name. Gemini 3.8 Flash is the first Google model to take a rank here, at 1494 on the text board, above the Opus 5 row. Muse Spark 1.3 and 1.2 from Meta were added too, and the point about them is that their weights are not published: the maker best known for years for publishing weights now ships its own top model closed. The rest: Kimi K2.7 Code, Grok 4.20, Veo 3.1 Lite, Lyria 3.5 Pro, Gemini 3.5 Transcribe, Eleven Music v2 and Scribe v2. The two transcription models are deliberately out of the audio category, because 65 of that rubric weight sits on Persian quality and voice control and both are about making speech rather than hearing it.
Source: developers.openai.com
-
new page Claude Fable 5.1
The Claude Fable 5.1 page was written, for a model released on 1 September 2026. The finding: the base price did not move at all, $10 in and $50 out exactly as on Fable 5, and the only row that changed on Anthropic pricing table is the cache read, from one dollar to twenty-five cents. Anthropic calls that an exception in a footnote under that same table, because the cache multiplier is 0.1x the input price on every Claude model and 0.025x on Fable 5.1 and Mythos 5.1 alone. The benchmark jump is not uniform either, and that became the argument of the page: agentic scientific research goes from 24.7 to 52.6, terminal coding from 42.0 to 55.8, but CursorBench only from 70.5 to 73.4. The model appears in no ranking table on this site and that is a decision: the coding rubric puts 70 of its 100 points on two Arena leaderboards and a one-day-old model has not been rated on either, so placing it in the category would only print it below the table as unranked. Mythos 5.1 was added to the catalogue and Fable 5 moved to legacy.
Source: www.anthropic.com
-
new page Translation
The translation category was written, the first of tranche two. The finding: two things have come apart here, the model that names your language in its own list cannot be bought, and the model you can download and run does not name your language. Command A Translate is the only model whose maker both calls it a translation model and names all four of this site languages in one list of 23; Llama 4, which does carry a commercial licence, gives a closed list of twelve with neither Persian nor Turkish on it. Licence openness carries 30 of 100 here and nowhere else in this reference, because in Iran the hosted API is the link that does not work. One criterion was tried and removed: the context window, because min-max normalisation with a low of 8,000 and a high of ten million rendered a thirty-two fold difference as 0.4 against 0.2.
Source: docs.cohere.com
-
new page Cohere
The Cohere page was written, a maker that until today was only a few rows in the catalogue. The finding: Cohere is the only maker in this reference that names Persian item by item in a model language list, in Command A Translate and in Aya Expanse 32B, and all three routes to that Persian are shut separately. The translation model has no published price, the open-weight model is CC BY-NC so no commercial use, and the model that is Apache 2.0 and usable commercially has no Persian at all. Separately, Cohere own commercial agreement writes Iran by name into its Restricted Location definition, which differs from every other maker here: the others leave Iran off a list, Cohere writes it in.
Source: docs.cohere.com
-
new page Amazon
The Amazon page was written. Two things it exists for: first, the Nova specification table prints 200+ in the Supported Languages cell and names fifteen languages in that cell footnote, with Arabic and Turkish among the fifteen and Persian not. The same pattern as ElevenLabs, this time in a footnote. Second, Nova access is a list of data centres rather than a list of countries, and not one of those lists contains a Middle East region; the nearest place a Nova model runs is Mumbai or Frankfurt. No Nova price was printed either, because the Bedrock pricing page fills its numbers in with JavaScript.
Source: docs.aws.amazon.com
-
new page Voice
The voice category was written, with six ranked models and one finding the whole page exists for: at ElevenLabs, Persian is on the v3 language list only. Flash v2.5 carries 32 languages and Multilingual v2 carries 29, and Persian is on neither. Flash v2.5 is the model people drift toward because it is faster, cheaper and has eight times the character limit, so they feed it Persian, get a poor result and conclude ElevenLabs cannot do Persian. In this category Persian carries a weight of 35 of 100, because when the output is itself a voice, a model without Persian is zero rather than slightly worse.
Source: elevenlabs.io
-
new page ElevenLabs
The ElevenLabs page was written. Apart from the Persian question, there is a condition on their own pricing page that is rarely written down: the free plan has neither a commercial licence nor voice cloning. So you can test a whole project for free, pull the audio out, and only then discover you had no right to use that output in paid work. The commercial licence starts on the Starter plan at $6.
Source: elevenlabs.io
-
new page ElevenLabs v3
The ElevenLabs v3 page was written, the model that came first in the voice table. A condition to know before a project starts: its character limit is 5,000, an eighth of Flash v2.5. So the model that reads Persian is the one you have to cut long text up for, then stitch the pieces back together yourself.
Source: elevenlabs.io
-
new page Llama 4 Scout
Llama 4 Scout and Mistral Large 3 got pages, the two models their maker pages could not link to. Scout has the largest context window in this reference, ten million tokens, ten times its nearest rival. Large 3 is the largest open-weight model recorded here, 675 billion parameters under Apache 2.0. Both pages make the same point, which is rarely written down: the active parameter count tells you the compute cost and the total tells you the memory, and anyone reading only the first orders the wrong hardware.
Source: huggingface.co
-
new page Mistral AI
The Mistral and Meta pages were written, two makers that were data-only until yesterday. The Mistral finding: open means four things at this company and one of them forbids commercial use. Large 3, Small 4 and Ministral 3 are Apache 2.0, Medium 3.5 is a modified MIT licence, and Voxtral TTS is CC BY-NC 4.0. Anyone who read that Mistral is the open one and built their product voice on Voxtral is in breach. Also, Le Chat and La Plateforme no longer appear in their own documentation; Vibe and Studio have replaced them.
Source: docs.mistral.ai
-
spec correction Meta
Meta lists exactly twelve languages for Llama 4 and Persian is not one of them. Arabic is. The list is closed and does not say including, so unlike the Mistral case the absence here is not an inference. Two more numbers were recorded: both models sit on an August 2024 knowledge cutoff, the oldest among the live models in this reference, and Scout at 109 billion parameters has a ten million token window while Maverick at 400 billion has one million.
Source: huggingface.co
-
new page
A full model list was added to this reference, and with it ninety two new entities, including four makers that were simply absent until yesterday: Mistral, Cohere, Amazon and Meta. The catalogue went from eighteen models to eighty nine and every cell in it was read from the maker own page today. These new models are deliberately in no category, which means they moved no ranking table: pushing thirty new rows into six published rankings with one command would have made the prose on those pages false the moment it ran.
-
price DeepSeek V4 Flash
DeepSeek V4 Flash got dearer, and this is the only number in this reference that has ever gone up. On 12 August it was $0.14 in and $0.28 out; its own pricing page now carries two columns, $0.22 and $0.66 off peak and $0.44 and $1.32 at peak. We record the peak column, because DeepSeek peak hours fall in the middle of the Iranian working day. DeepSeek itself lost no place and is still the cheapest row, but because its lead narrowed three other rows rose in the coding table: Qwen from fifth to third, Kimi from third to fourth and GPT from fourth to fifth, none of which changed anything.
Source: api-docs.deepseek.com
-
spec correction
Sora came out of the we have no facts state. This section decided twice to keep Sora data only, because openai.com returns 403 to our server and its developer surface did not carry Sora either. Today it does: the model page and the price list both answer, so the id, the resolutions, the input and output modalities and the price per second are now recorded. Sora 2 Pro still has no price, because OpenAI prints it as a range and a range is not a number.
Source: developers.openai.com
-
new page Veo
The Veo family page was written and it records a contradiction inside Google own documentation: the model table marks Veo 3 and Veo 2 Stable, and the pricing page on the same site declares both deprecated with a shutdown date of 30 June 2026, which is already past as we write. The only version worth using, Veo 3.1, is marked Preview. We are not saying the models were switched off; we are saying the Stable label in this family does not tell you what to build on.
Source: ai.google.dev
-
new page Seedance
The Seedance family page was written, and the resolution regression in this family now has a third and firmest piece of evidence: on the Runway price list Seedance 2.0 has four resolution rows going up to 4K and Seedance 2.5 has only two, 480p and 720p. A price table is not a marketing page. The family timeline comes from ByteDance own model ids, which carry the release date inside them.
Source: docs.dev.runwayml.com
-
new page Runway
The Runway, ByteDance and Kuaishou pages were written, closing the last stubs in this section. The Runway finding: its own Gen-4.5 is twelve credits a second and the Seedance it resells is a hundred and fifty at 4K, so the most expensive video in its shop costs twelve times the cheapest model it makes. And Runway carries a silent price for Veo 3.1 that Google own pricing page does not have at all.
Source: docs.dev.runwayml.com
-
new page Music generation
The music category was written, with five candidates and four of them ranked. One of its columns we have seen in no other comparison: file delivery. The question of which one sings better has an answer and Suno wins it, but the second question is whether you end up with the file, and there the order changes. This is the only category here whose quality numbers do not come from Arena, because Arena has no music board.
Source: artificialanalysis.ai
-
new page Udio
The Udio page was written, and one sentence in their own help centre is the reason it exists: since 29 October 2025, alongside the Universal deal, audio, video and stem downloads have been disabled. So the paid subscription makes songs and hands you no file. We have no dollar price, because udio.com renders entirely in JavaScript and its body is empty to curl, so that column stayed empty.
Source: help.udio.com
-
new page Suno
The Suno page was written and it took the top of the music table: 1171 on vocals, 1188 on instrumental, and the longest track at eight minutes. The thing buyers miss, and it is on the page: credits carry over neither from one day to the next nor from one month to the next, and the download cap is a separate ceiling from the credit cap. Pro gives 2,500 credits and twenty downloads a month.
Source: suno.com
-
new page Chat and assistants
The chat category was written, with nine candidates and all of them ranked. The decision argued on the page: this table ranks models rather than apps, because all five of its columns are numbers published only about a model. The cost is that ChatGPT, the Claude app and the Gemini app carry no row.
Source: arena.ai
-
new page
GPT-5.6 Sol was added to the data as a candidate in both the chat and coding tables: $5 and $30, a 1.05M token window, 128K max output. It also corrects an assumption of ours. We had written that nothing from OpenAI is readable; the accurate version is that openai.com is not readable while developers.openai.com is. So the model can be ranked and the app cannot be described. It has no page, because a full page is impossible without talking about ChatGPT.
Source: developers.openai.com
-
ranking moved Coding
Two data corrections moved the coding table. First, the text leaderboard was read in full and the GLM-5.2 Max row (1470) and the DeepSeek V4 Flash row (1435) turned out to have been there all along, with an earlier run having read only the top of the page. Second, GPT-5.6 Sol became an eighth candidate. Opus 5 went 76.0 to 77.5, Kimi 49.9 to 52.4, Qwen 49.3 to 51.3, GLM 20.1 to 25.0, Grok 10.8 to 15.6. The only change of order: GPT landed between Kimi and Qwen.
-
new page Grok 4.5
xAI, the Grok 4 family and Grok 4.5 were added. The number most worth writing down: from 200,000 tokens up, every token in that request bills at double, not just the tokens above the line. And the newest member of the family carries half the context window of the one before it at more than double the output price.
Source: docs.x.ai
-
ranking moved Coding
Adding Grok 4.5 as a seventh candidate raised every score in the coding table without any model having changed: Opus 5 from 64.1 to 76.0, Fable 5 from 46.1 to 60.6, Kimi from 44.1 to 49.9, Qwen from 34.6 to 49.3, DeepSeek from 20.0 to 38.1 and GLM from 2.3 to 20.1. The cause is relative normalisation: Grok became the floor of three columns, and the floor of a column scores zero. No row changed position.
-
new page Claude Fable 5
The Fable 5 page was written. The number that stood out: the knowledge cutoff of Anthropic dearest widely released model is January 2026, while Opus 5, at half the price, reads through May 2026. The two are 64 points apart on the code leaderboard and 12 apart on the text one, in opposite directions.
Source: arena.ai
-
new page Kimi K3 Max
The Kimi K3 Max page was written, with the point that K3 Max is not a model in Moonshot catalogue. Its context window is 1,048,576 tokens, two to the twentieth, and that sub-five-percent difference hands it the whole ten-point context column. We say so on the page, because it is not ten points of real advantage.
Source: platform.kimi.ai
-
new page GLM-5.2 Max
The GLM-5.2 Max page was written, and it holds the cleanest evidence for the claim that max is a parameter: Z.ai own documentation calls the model glm-5.2 and lists max as a value of reasoning_effort. Its table total is 2.3, with fifteen of a hundred points carrying no evidence.
Source: docs.z.ai
-
new page DeepSeek V4 Flash
The DeepSeek V4 Flash page was written. Output at $0.28, close to ninety times cheaper than Opus 5, and a 384,000 token output ceiling, three times the Anthropic models. Its table total is 20 and every point comes from one column; it scores zero in the forty-point leaderboard column.
Source: api-docs.deepseek.com
-
new page Gemini 3.1 Pro
The Gemini 3.1 Pro page was written, with an Iran column read from the Gemini API available-regions list, which does not include Iran. It is the only candidate the coding table cannot rank: Google publishes no context window and we have not read its web-dev row. Its price is tiered, with the boundary at 200,000 tokens.
Source: ai.google.dev
-
spec correction Coding
Three sentences on the coding page had stopped being true and were corrected in all four languages. GLM-5.2 Max and DeepSeek V4 Flash now carry a rank and the only unranked candidate is Gemini 3.1 Pro; and of the four Chinese models in the table only Qwen3.8 Max still has no verified price. No weight and no rank changed.
Source: arena.ai
-
new page Video generation
The video category went live with four models and it ranks versions rather than products, because one model is resold across several tools. Weights: clip length 30, resolution 25, reference inputs 20, scene control 15, Iran 10. Seedance 2.5 came first. Generation time is not a criterion because no vendor publishes it, and neither is price, because Kling is absent from the only dollar-denominated list there is.
Source: www.volcengine.com
-
new page Seedance 2.5
Seedance 2.5 was recorded: up to 30 seconds in one generation and up to 50 reference inputs, being 30 images, 10 videos and 10 audio clips. The thing the product page omits and the API docs state: this model resolution ceiling is 720p, where Seedance 2.0 went to 4K.
Source: www.volcengine.com
-
new page Kling VIDEO 3.0
Kling VIDEO 3.0 was recorded: three to fifteen seconds, native 4K since 23 April, and native audio in five speech languages. Persian is not one of the five, which for a Persian reader is the most important fact on the page.
Source: kling.ai
-
new page Kling
The Kling page was written. Five plans with prices from their own credit guide, plus the point that the advertised prices are first-subscription prices: Standard costs $6.99 and renews at $8.80.
Source: kling.ai
-
new page Veo 3.1
Veo 3.1 was recorded: four, six or eight seconds, audio always on, and up to three reference images. 4K and 1080p exist on the eight second clip only, and an extended output is 720p.
Source: ai.google.dev
-
new page Runway Gen-4.5
Runway Gen-4.5 was recorded, with numbers read from their own OpenAPI schema: two to ten seconds, a 720p ceiling, no audio parameter. No reference-input ceiling is published, so that cell stayed unverified and cost it twenty points of weight.
Source: docs.dev.runwayml.com
-
new page Studying and learning
The studying category went live with four candidates and, unlike coding, it ranks products rather than model versions. Weights: digestion formats 35, your own material 20, workspace 20, output languages 10, Iran 15. Gemini Notebook came first. ChatGPT is in the table without a rank, because OpenAI pages return 403 to our server.
Source: support.google.com
-
new page Gemini Notebook
Gemini Notebook was recorded, and recording it showed that Google has dropped the NotebookLM name: the official page and the help centre both write Gemini Notebook. Free plan: 100 notebooks, 50 sources each, nine output formats, and Persian on the output-language list.
Source: notebooklm.google
-
new page Gemini
The Gemini page was written around Guided Learning. The ten files per prompt ceiling and the 70 languages come from Google own pages, and so does the boundary with Gemini Notebook: the Studio outputs cannot be generated in Gemini.
Source: blog.google
-
new page Claude
The Claude page was written. Learning mode was confirmed from Anthropic own education page and the Pro price ($20 monthly, $17 on the annual plan) from the pricing page. Project capacity has no published figure, so that column stays empty in the study table.
Source: claude.com
-
new page Coding
The coding category went live with six candidates and real evidence. The primary criterion is the WebDev Arena leaderboard at weight forty, and Opus 5 came first at 1692. No SWE-bench figure was entered, because we could not read that leaderboard directly.
Source: arena.ai
-
new version Claude Opus 5
Opus 5, released 24 July 2026, was recorded with full specs at $5 and $25. The announcement Frontier-Bench and ARC-AGI 3 claims are relative and carry no absolute number, so they did not enter the table.
Source: www.anthropic.com
