The DeepSeek V4 family
This family has two members, and choosing between them inverts what a flash-and-pro ladder normally implies: Flash costs a third as much, carries five times the concurrency, and the Responses API still does not work on Pro. Make Flash the default.
- Pro costs three times more
- Flash concurrency is five times higher
- a 1M token window on both
Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. What changed
Which member to take
Flash. DeepSeek own table gives both models the same one million token context window and the same 384K maximum output, so capacity is not what separates them. Price is. Flash charges $0.14 for uncached input and $0.28 for output; Pro charges $0.435 and $0.87, roughly three times as much. The concurrency limit is 2500 requests on Flash and 500 on Pro.
So the cheaper model carries five times the parallel capacity. For a service whose traffic arrives in bursts that occasionally matters more than the price itself, and anyone who assumes Pro is simply the higher tier has usually not looked at that column.
What the vendor itself wrote, and few people quote
The change log entry dated 31 July 2026 says the official Flash release arrived with sharply enhanced agent capability and that its results are far beyond V4-Pro-Preview. The figures printed beside it: Terminal Bench 2.1 at 82.7, NL2Repo at 54.2, Cybergym at 76.7, DeepSWE at 54.4 and Toolathlon verified at 70.3.
Precision matters here, because the distinction changes the conclusion. That comparison is against the preview of Pro, not against the Pro sitting in today price table, and the same change log says the official Pro release will follow soon. So the correct claim is not that Flash beats Pro. It is that the vendor has nowhere shown Pro to be better, while charging three times as much for it. For a purchasing decision that is enough.
One technical detail from the same entry: Flash 0731 keeps the model architecture and size of the preview and was only re-post-trained. The jump in those numbers came from training, not from more hardware.
Two features only Flash has
The feature table carries a column that rarely makes it into summaries: the Responses API works on Flash and not on Pro. The footnote to that same table says Pro support will be added in early August 2026. We read the page on 12 August 2026 and that sentence was still sitting there, so either it has not shipped or it shipped and the documentation was not updated. Both readings mean the same thing for you: do not plan around it until you have seen it yourself.
Flash is also natively adapted for Codex, per the same change log. If your tool is Codex, Claude Code or OpenCode, this family can be dropped in as the backend model without writing any code.
The timeline, which is the clearest in this group
24 April 2026: V4 arrives and the API accepts both names, V4 Pro and V4 Flash. That same entry announces that the two legacy names, deepseek-chat and deepseek-reasoner, will be discontinued three months later on 24 July 2026, and that until then they point at the non-thinking and thinking modes of Flash. 31 July 2026: the official Flash release, which describes itself as being in public beta.
This level of dating is unmatched among the Chinese vendors in this reference. Moonshot and Z.ai date no version at all, which is exactly why the GLM-5 family page has an order and no timeline.
Release timeline
- DeepSeek V4 Flash current The cheapest price in the coding table, and the longest output ceiling in it.
Using it from Iran
This column carries a date and says how it was checked, because most listicles guess it. Where we have only read a vendor policy page, the note below says exactly that.
- Reachable
- not checked
- Payment
- not checked
- Free tier
- no
How we checked: DeepSeek publishes no supported-countries list, so the access column stays unchecked. We have not tested from an Iranian connection either.
What it is good at
- The cheapest row in our coding table, even at the uncached input rate
- A one million token window and up to 384K output on both members
- A 2500 request concurrency limit on Flash, five times that of Pro
- A dated change log, down to announcing the shutdown day of the legacy model ids
Where it falls short
- The pricing page says overall API pricing will rise in the near future with a significant increase expected, so do not fix a long-term budget to today figure.
- The official V4 Pro release has not landed, and the vendor benchmark comparison is against the Pro preview rather than the Pro being sold.
- The Responses API does not work on Pro, and the documentation still shows the passed early-August date.
- Flash itself is called public beta in the change log, which matters for a sensitive workload.
- FIM Completion works in non-thinking mode only, on both members.
- Access from Iran is unchecked, and DeepSeek publishes no country list.
Our take
Make Flash the default and reach for Pro only once you have measured your own workload and seen a difference, because the vendor has not shown one. And keep today price out of any long-term contract, since they have written that it is going up.
Questions people actually ask
Is V4 Pro stronger than V4 Flash
DeepSeek has published no benchmark figure putting Pro ahead. The only comparison it offers places Flash far beyond the Pro preview. Pro costs three times more, carries a fifth of the concurrency, and the Responses API does not work on it.
What happened to the deepseek-chat and deepseek-reasoner ids
They were discontinued on 24 July 2026, announced three months in advance in the 24 April change log entry. Until that day they pointed at the non-thinking and thinking modes of deepseek-v4-flash.
Sources
- DeepSeek models and pricing vendor source 12 August 2026
- DeepSeek change log vendor source 12 August 2026