Terminal-Bench goes from 61.8 to 82.7. DeepSWE goes from 7.3 to 54.4. And the small model beats its own house flagship on all nine published benchmarks. So much for the press release.
What interests me here is not the score race, which expires in three weeks. It is what those twenty-four hours say about the market: price per million tokens no longer means anything, and what counts now is cost per finished task. On that ground the gap is not a gap, it is a cliff.
V4 Flash costs about $0.03 per task, where GPT-5.6 Sol costs $1.86. For a large share of agentic steps, that is enough. Model choice becomes a budget decision before it becomes a technical one. Figures current as of August 7, 2026.
What changed on July 31
One clarification first, because it changes how you read the rest. DeepSeek did not ship a new model. The V4 family has existed since April 24, 2026, under an MIT licence, with weights available on Hugging Face. What landed at the end of July is a new build called V4-Flash-0731. Same model, same architecture. Only the post-training was redone.
So the entire gain comes from the agentic training phase. The base did not move, and that is exactly what makes the jump interesting: it does not cost a fresh pre-training run.
| Benchmark | Preview build | 0731 build |
|---|---|---|
| Terminal-Bench 2.1 | 61.8 | 82.7 |
| DeepSWE | 7.3 | 54.4 |
| Cybergym | 38.7 | 76.7 |
| NL2Repo | 39.4 | 54.2 |
| Toolathlon-Verified | — | 70.3 |
| Agent Last Exam | — | 25.2 |
| Automation Bench (public) | — | 25.1 |
| DSBench-FullStack | — | 68.7 |
| DSBench-Hard | — | 59.6 |
DeepSWE is multiplied by 7.5. Cybergym doubles. And the small model moves ahead of the house flagship on all nine published benchmarks.
Three caveats before getting carried away
These numbers come from DeepSeek, not from an independent third party. Three things to keep in mind.
The harness matters as much as the model. Evaluation ran on DeepSeek’s in-house harness, in minimal mode, max tier, top_p at 0.95 and temperature at 1.0. DeepSeek itself states that agent scores are highly sensitive to the harness. Your numbers will differ from theirs.
Two of the nine benchmarks are internal. DSBench-FullStack and DSBench-Hard belong to DeepSeek. They compare to nothing else.
The top spot is not taken. On Terminal-Bench 2.1, Opus 4.8 is still ahead at 85.0 against 82.7. GLM-5.2 sits at 81.0. My point is not that DeepSeek is better. It is that DeepSeek is almost as good, for far less money.
Why a small model pulls it off
284 billion parameters is not small. But it is not what you pay for.
V4 Flash is a Mixture-of-Experts model. Out of the 284 billion stored parameters, only 13 billion activate per token. Compute cost follows what activates, not what is stored.
On top of that sits a hybrid attention scheme combining two compression mechanisms. The claimed result: 27% of V3.2’s FLOPs per token, and 10% of its KV cache at one million tokens of context.
In practice, long context stops being a tax. An agent dragging 300,000 tokens of history, documentation and tool output no longer blows up the bill. That is exactly the profile of a real agent, as opposed to a chatbot.
Throughput follows: 81.3 tokens per second against 35.6 for V4 Pro on the official endpoint. Across a forty-step agent loop, that latency gap often matters more than a benchmark point.
One last practical note. Quantised to 4 bits, the model fits in 150 to 190 GB. A single H100 or H200 node is enough. So is a machine with 192 GB of unified memory. This is the first model at this level that a mid-sized company can genuinely host itself.
V4 Flash
Tool calls, extraction, classification, log reading, retry loops.
Mid-tier model
Aggregating outputs, writing for a human reader, checking a chain of steps.
Frontier model
Architecture calls, cross-cutting refactors, decisions where an error is expensive.
What it actually costs
Here are the official rates, after OpenAI’s July 30 cut.
| Model | Input / M tokens | Output / M tokens |
|---|---|---|
| DeepSeek V4 Flash | $0.14 ($0.0028 on cache hits) | $0.28 |
| GPT-5.6 Luna | $0.20 | $1.20 |
| Claude Opus 4.8 | $5.00 | $25.00 |
| GPT-5.6 Sol | $5.00 | $30.00 |
First thing to fix: stop comparing V4 Flash to Sol. Nobody runs forty agent steps on a model at $30 per million output tokens. That would be picking the easiest opponent.
Since July 30, the real competitor to V4 Flash is Luna. Against Luna, the gap is no longer 100×. It is 1.4× on input and 4.3× on output. Less spectacular. Also far more solid, because output is what costs money in an agent loop.
Price per token does not tell the whole story
A cheap but verbose model that needs three times as many steps ends up costing more than an expensive, efficient one. That is why Artificial Analysis measures something else: the average cost of a finished task.
| Model | Average cost per task |
|---|---|
| DeepSeek V4 Flash | $0.03 |
| Kimi K3 | $0.86 |
| GPT-5.6 Sol | $1.86 |
| Claude Fable 5 | $3.15 |
On their Intelligence Index, which aggregates nine benchmarks, V4 Flash scores 50 out of 100 — level with Gemini 3.6 Flash, one point below Muse Spark 1.1 and GLM-5.2. Their conclusion: it is the cheapest well-known model to run.
A worked example
Take a support agent. Forty steps per ticket, 200,000 input tokens and 40,000 output tokens, 10,000 tickets a month. That gives 2 billion input tokens and 400 million output tokens. I am showing the assumptions on purpose: adjust them to your own traffic, the structure of the calculation stays the same.
| Model | Per month | Per year |
|---|---|---|
| V4 Flash · DeepSeek API | $392 | $4,704 |
| V4 Flash · DeepSeek API, 70% cache hits | $200 | $2,400 |
| V4 Flash · OpenRouter alias ($0.09 / $0.18) | $252 | $3,024 |
| GPT-5.6 Luna | $880 | $10,560 |
| Claude Opus 4.8 | $20,000 | $240,000 |
| GPT-5.6 Sol | $22,000 | $264,000 |
The gap between Sol and Flash exceeds $259,000 a year, on a single use case.
Caching deserves a mention. At $0.0028 per million, an agent with a stable system prompt and fixed tools halves its bill again. And since real agentic traffic is overwhelmingly dominated by input, this line weighs more in practice than my deliberately conservative 5:1 ratio suggests.
One detail for budget owners. Before July 30, the Luna line read $4,400 a month. OpenAI’s cut just saved this fictional company $3,520 a month. It asked for nothing, renegotiated nothing, migrated nothing. Somebody else obtained that discount on its behalf. It is the same mechanism I described in my piece on the FinOps wake-up call on coding agent bills: the budget line moves without the engineering team touching anything.
The price hike DeepSeek announced
A slightly comical aside. While I was writing this article, I received this email.
The publicly stated reason is fairly ironic: the very low price attracted so many users that the platform is saturating and becoming unstable. DeepSeek is a victim of its own selling point.
The timing is intriguing too. Five days after an update that puts its small model ahead of its flagship. Six days after OpenAI’s 80% cut.
Does this invalidate the article? No. And the reason is precisely the point of the whole piece.
DeepSeek can raise the price of its API. DeepSeek cannot raise the price of its model.
— what an MIT licence entails
The weights are published. Anyone can host them and resell them at whatever price they choose. This is not a hypothesis: on OpenRouter, the alias that always points to the latest build of the family is listed at $0.09 per million input tokens and $0.18 output, roughly 35% below DeepSeek’s official rate.
The model is also the most heavily used one on the platform. In other words, the lab that built the model no longer sets its price. If the official API gets expensive, traffic moves to the other hosts.
So what to watch over the coming weeks is not DeepSeek’s announcement. It is how providers on OpenRouter react. My hypothesis, and I own it as such: prices there will not move much, because none of those hosts has any reason to follow a competitor’s increase. I will update this section if the facts prove me wrong.
Treat $0.14 and $0.28 as a price of the moment. DeepSeek had already announced doubled rates during two daily peak windows in Beijing time, without ever communicating a start date. The right architecture is not “I plug in DeepSeek”. It is “I plug in V4 Flash, and I change host by editing an environment variable”.
What Flash does well, and what it doesn’t
You often read that a model like this covers 80% of agentic use cases. Let us be honest about that number: it is a field estimate, not a measurement. No study establishes it, and it should not be confused with the 80% of open-source-stack startups running a Chinese model, which is a different statistic.
What is established is the mechanism. Agentic workloads are dominated by narrow, repetitive steps. Parsing a tool call, extracting data, classifying, rewriting. Those steps do not need a frontier model.
The practice taking hold in 2026 is tiered routing. A small model handles the majority of steps and escalates to a large one only for genuinely hard reasoning. Published measurements point the same way: roughly 98% of the quality for half the cost at Orq.ai, up to 85% inference savings according to IBM.
What V4 Flash absorbs well:
- tool calls and processing their returns;
- data extraction and normalisation;
- ticket classification and routing;
- reading and summarising logs;
- test generation and localised fixes;
- error-recovery loops;
- navigating a large context.
What I would not put it on:
- architecture decisions;
- refactors touching several modules;
- high-stakes legal or financial reasoning;
- anything where a single error costs more than ten thousand model calls.
Hence my rule: route by cost of error, not by perceived difficulty. A task can be simple and still deserve the most expensive model, because a false positive turns into a customer incident. Another can look complicated and run perfectly well on Flash, because the error gets caught at the next step.
To place the models relative to each other before making that call, my 2026 coding LLM comparison gives the wider landscape.
The 24 hours of July 30–31
Here is the timeline for the month. It speaks for itself.
| Date | Event |
|---|---|
| July 7 | CNBC documents that Chinese models capture up to 46% of US enterprise tokens on OpenRouter |
| July 9 | OpenAI launches the GPT-5.6 family: Sol, Terra, Luna |
| July 24 | More than 50 companies, including Meta and Microsoft, sign a letter against banning open-weight models. Neither OpenAI nor Anthropic signs it |
| July 29 | Mark Zuckerberg says banning Chinese models is not an effective solution, and raises the risk of regulatory capture, naming OpenAI and Anthropic |
| July 30 | OpenAI cuts Luna by 80% and Terra by 20%. Sol does not move |
| July 31 | DeepSeek publishes V4-Flash-0731 |
| August 5 | DeepSeek announces a “significant” increase to its API pricing |
Nothing supports the claim that OpenAI cut prices because of DeepSeek. What can be said is that the two decisions land twenty-four hours apart, and that they answer the same market reality.
One detail deserves attention. OpenAI did not defend Sol, it defended Luna. You do not cut the price of a three-week-old model by 80% out of comfort. And you do not do it in the segment where you are safe. The move shows where the pressure sits: on the volume of agentic steps, not on frontier reasoning.
The shift has already happened
OpenRouter’s weekly ranking for July 27 to August 2 is unambiguous.
| Rank | Model | Tokens processed |
|---|---|---|
| 1 | DeepSeek V4 Flash | 7.22T |
| 2 | Xiaomi MiMo-V2.5 | 5.1T |
| 3 | Tencent Hunyuan Hy3 | 5.01T |
| 4 | DeepSeek V4 Flash (second listing) | 3.45T |
| 5 | GPT-5.6 Luna | 2.99T |
The top four places are Chinese. V4 Flash holds two of them. Over the week, Chinese models processed 28.13 trillion tokens — the fourteenth consecutive week ahead of US models.
One line is worth pausing on. GPT-5.6 Luna jumps 738% in a single week. The July 30 cut worked immediately. That is the best possible confirmation: in this segment, volume follows price, not benchmarks.
The rest of the picture:
- Chinese models went from under 2% of token consumption in late 2024 to more than 50% in June 2026;
- the share of open-weight models on OpenRouter went from 34% in January to 65% in June;
- on Vercel’s AI gateway in June, open Chinese models accounted for 29% of tokens for less than 4% of spend;
- Flo Crivello, founder of Lindy, moved 100% of his traffic to DeepSeek V4 in June;
- Martin Casado, at Andreessen Horowitz, estimates that around 80% of startups on an open-source AI stack run a Chinese model.
What this does to margins
OpenAI runs at roughly $25 billion in annualised revenue, for an estimated 33% gross margin. Anthropic saw its inference margin move from 38% to more than 70% in 2026, carried by Claude Code.
The problem is not temporary, it is structural. Once anyone can host a model, nobody controls its price. Inference drifts toward the cost of the hardware.
And the rule makes no exceptions. It applies to DeepSeek too, which is learning it live: the lab announces an increase while others sell its model a third cheaper. Publishing your weights under MIT buys adoption and forfeits pricing power.
It is also the best lens for reading Zuckerberg’s statement. Without ascribing intent, two things can be observed. A company that built its strategy on open models has no interest in the market closing. And the two players who did not sign the July 24 letter are the ones operating the dominant closed models.
For a European company, the consequence is simple. Your negotiating power just went up, even if you do not switch providers. Having a credible, tested alternative changes the contractual conversation.
The compliance caveat
The reservation has to be stated, and it is a serious one.
The API hosted by DeepSeek stores data on servers in China. European authorities have taken it up. Italy blocked access in late January 2025. Belgium opened an investigation. France’s CNIL announced an analysis of DeepSeek’s tools. The complaints concern the legal basis, retention periods, identification of sub-processors and the absence of a representative in the EU.
For a European company handling customer data, plugging a production agent into the public DeepSeek API is not reasonable.
But the weights are MIT-licensed. You can download, host, modify and commercially exploit them. The model and the service are two different things: only the service raises a data-residency problem.
Three routes, in increasing order of infrastructure cost:
- The DeepSeek API, to prototype and measure on non-sensitive data. Fast, nearly free, and to be ruled out in production on customer data.
- Inference operated inside the European Union. Several providers offer V4 Flash on European GPUs, with a DPA, a published sub-processor list and no CLOUD Act exposure. Observed orders of magnitude run from €0.10 input and €0.20 output for a dedicated offer, up to €1.20 and €3.40 for a shared router. Check these rates when you decide: they move fast.
- Self-hosting. 150 to 190 GB at 4 bits, one H100 or H200 node. A real entry cost, full control, and quick payback above a certain volume.
Chinese pricing without Chinese risk is achievable. But you pay for it in infrastructure work, not with an API key. It is the same trade-off I describe in do not choose a model, choose a strategy.
Where to start
Five steps, in order.
1. Instrument before you migrate. Measure your cost per finished task, not per token. Without that measurement you will never know whether a model change made or lost you money. Most teams do not have this number.
2. Identify your three most verbose loops. In almost every agent architecture I see, two or three loops concentrate most of the spend. Often rewriting or retry steps nobody has reviewed since they were written.
3. Replay those loops on V4 Flash, with the same harness. That is the condition for an honest comparison, since DeepSeek itself acknowledges harness sensitivity. Compare three things: success rate, cost, end-to-end latency.
4. Route what passes, keep the large model for the rest. Do not try to move everything. The goal is not to leave your provider. It is to stop paying frontier-reasoning rates for tool-call parsing.
5. Choose your hosting based on data sensitivity. That decision belongs with the DPO, not with the engineering team alone.
Above all, do not tie yourself to any provider. The model name and the endpoint URL must be configuration variables, never constants in the code. Since V4 Flash is served by several hosts through an OpenAI-compatible API, switching provider should stay an environment-variable change:
import os
from openai import OpenAI
# Switching host = changing these three variables, never the code.
# LLM_BASE_URL=https://api.deepseek.com LLM_MODEL=deepseek-v4-flash
# LLM_BASE_URL=https://openrouter.ai/api/v1 LLM_MODEL=deepseek/deepseek-v4-flash
# LLM_BASE_URL=http://my-h200-node:8000/v1 LLM_MODEL=DeepSeek-V4-Flash (vLLM)
client = OpenAI(
base_url=os.environ["LLM_BASE_URL"],
api_key=os.environ["LLM_API_KEY"],
)
resp = client.chat.completions.create(
model=os.environ["LLM_MODEL"],
messages=[
{"role": "system", "content": "You extract structured data. Reply with strict JSON."},
{"role": "user", "content": "Ticket 4821: the Fabric sync has been failing since 03:12, code 429."},
],
temperature=0,
)
usage = resp.usage
print(resp.choices[0].message.content)
# The number to track is not the sticker price, it is the cost per finished task.
print(f"in={usage.prompt_tokens} out={usage.completion_tokens}")
The token counter at the end of the script is not decorative. It is the starting point of step 1: without a per-task measurement, comparing two models stays an opinion.
My take
The debate about Chinese models hides the real shift. The question is no longer who builds the most intelligent model. On that front, the gap between the top and the middle of the range narrows every quarter. The question is how much a finished task costs.
Something changed at the end of July. A model with 13 billion active parameters reaches solid agentic performance for three cents a task. A US lab cuts its entry-level price by 80% the day before. And more than half of the tokens consumed worldwide come from open Chinese models.
Five days later, DeepSeek announces a price increase, unable to stop its own model from being sold a third cheaper elsewhere. That is probably the most durable lesson of the episode. Once the weights are published, no lab regains control of the price — not even the one that trained them.
My advice: test V4 Flash this week on a non-sensitive case, measure the cost per task, and keep an architecture able to switch hosts. The question to ask in the meeting is no longer “which model is best”. It is “which is the cheapest model that passes my tests, and can I swap it next week”.
August 22, 2026 update. DeepSeek has also closed the biggest gap in its API lineup with V4 Flash Vision Exp, its new multimodal variant for agents.
Sources
- DeepSeek — V4 Flash GA announcement and agent benchmarks.
- DeepSeek API Docs — official models and pricing.
- Artificial Analysis — DeepSeek V4 Flash, cost per task and Intelligence Index.
- vLLM Recipes — 284B parameters, 13B active, 1M context.
- MarkTechPost — agentic gains in the 0731 build.
- Bloomberg — DeepSeek plans a significant price increase.
- TechNode — DeepSeek plans significant API price increases.
- TechNode — V4 Flash tops the OpenRouter weekly ranking.
- OpenRouter — DeepSeek V4 Flash Latest alias page.
- CNBC — OpenAI cuts prices on two GPT-5.6 models (July 30, 2026).
- VentureBeat — AI price wars, competition shifts toward cost.
- CNBC — Chinese models gain ground with US enterprises.
- Forbes — plunging token costs and the new margin squeeze.
- TNW — Zuckerberg against blocking Chinese AI models (July 29, 2026).
- Euractiv — DeepSeek and European data protection authorities.
- CNPD Luxembourg — recommendations on DeepSeek.



