DeepSeek could reason, code, call tools and hold a massive context. Yet it was missing something fairly basic for a modern agent. It could not see.

That gap forced teams into awkward pipelines. DeepSeek could handle the reasoning, but a screenshot, an invoice or a chart had to be routed to another model. That meant another provider, another routing rule, another data policy, another bill and a few extra seconds of latency. On August 21, 2026, that anomaly finally started to disappear.

Key takeaway

deepseek-v4-flash-vision-exp keeps V4 Flash text capabilities, accepts text and images, retains a one-million-token context and bills images at the same input price as V4 Flash.

The gap is finally closed

This story did not begin yesterday. When DeepSeek V3 launched in December 2024, the team already told users to look forward to multimodal support. DeepSeek then moved the market on reasoning, cost and open models, while its official lineup remained strangely blind.

One nuance matters. DeepSeek was not absent from multimodal research. Janus, Janus-Pro and DeepSeek-VL could already understand images, with weights available for self-hosting. What was missing was a Vision variant in the main API lineup, compatible with familiar formats and aligned with V4 Flash agent capabilities.

That is why I see this launch as more than one extra feature. This is finally the model DeepSeek had been missing all along. It does not simply fill a cell in a product matrix. It lets DeepSeek cover a much larger part of an agent workflow.

An agent can now receive a UI screenshot, spot an error, read a table, connect what it sees to its code knowledge and call a tool. Perception and reasoning live in the same model. The architecture becomes simpler and more coherent.

What DeepSeek actually announced

DeepSeek describes V4 Flash Vision Exp as an experimental multimodal model aligned with V4 Flash on text. It therefore retains the Flash branch's agent, reasoning and general knowledge capabilities.

The company also says multimodal agent performance comes close to Opus 4.8. That is a strong signal, but it needs careful wording. The launch post does not provide detailed scores or a complete protocol. I see a frontier-level ambition, not yet a proven win across every workload.

CapabilityDeepSeek V4 Flash Vision Exp
InputsText and images
Context1 million tokens
Maximum output384K tokens
APIsChat Completions, Responses and Anthropic Messages
FormatsJPEG, PNG, GIF and WebP
StatusExperimental

Why agents gain a new dimension

Vision becomes genuinely useful once it moves beyond “describe this image.” Combined with tools, visual elements can become actions.

  • Support and QA. An agent receives a screenshot, identifies the UI state, matches the issue with a knowledge base and opens a detailed ticket.
  • Business documents. It reads an invoice, purchase order or form, checks the fields and sends structured data to a workflow.
  • Data and monitoring. It analyses a chart, detects a visual break and then queries the underlying metrics or logs.
  • Development. It compares a mockup with a rendered page, sees a responsive bug and proposes a code fix.
  • Desktop agents. It understands the screen before deciding which tool to call or what step to execute.

DeepSeek says the model works across agent frameworks and released DeepSeek Harness 0.1.1 with built-in support. This is probably where the release delivers the most value. Vision is not isolated. It joins a model that already knows how to reason and use tools.

For agentic work with Hermes Agent, the impact is very practical. Until now, DeepSeek could handle the reasoning, but a screenshot or image forced the workflow to switch to a vision-capable model such as Codex. With V4 Flash Vision Exp, the same provider can now read the image, reason about it and continue the action. There is no longer a need to combine DeepSeek and Codex solely to give Hermes visual understanding. Routing, data governance, observability and billing all become much simpler.

An almost negligible cost

The pricing decision is aggressive. Each image is resized and converted into tokens, with a maximum of 384 tokens per image. Those tokens are billed exactly like V4 Flash input.

Price on August 22, 2026Off-peakPeak
Cached input$0.007 / 1M$0.014 / 1M
Uncached input$0.22 / 1M$0.44 / 1M
Output$0.66 / 1M$1.32 / 1M

At the 384-token ceiling, reading one image alone costs about $0.000084 off-peak or $0.000169 at peak, excluding text and generated output. Even a batch of 600 images at the ceiling is around $0.05 or $0.10 for visual input. This theoretical calculation is not the full cost of a workflow, but it shows how cheap visual perception has become.

Small images are scaled up to roughly 384 × 384 pixels. Larger ones are scaled down toward an envelope close to 800 × 800 pixels. Sending a 5000 × 5000 file does not buy more billed precision than a 2000 × 2000 file after resizing. For a dense document, crop useful regions instead of blindly uploading a huge image.

The API in a few lines

The integration uses the OpenAI client. You can send an image as base64 data, an external URL or a file_id. The last option is the cleanest choice when the same document is reused across multiple requests.

Python · image URL
from openai import OpenAI

client = OpenAI(
    api_key="<DEEPSEEK_API_KEY>",
    base_url="https://api.deepseek.com",
)

response = client.chat.completions.create(
    model="deepseek-v4-flash-vision-exp",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Analyse this screenshot and list the visible issues."},
            {
                "type": "image_url",
                "image_url": {
                    "url": "https://example.com/screenshot.webp",
                    "detail": "high",
                },
            },
        ],
    }],
)

print(response.choices[0].message.content)

The detail="low" setting resizes an image to 512 × 512 for faster and cheaper processing. The high, original and, for now, auto modes keep the original image before internal tokenisation.

Limits you need to know

The specification looks impressive, but the exp suffix is not decorative. Before placing the model in a critical workflow, I would test response stability, tiny text, dense tables, skewed documents, ambiguous charts and actual tool usage after visual analysis.

LimitOfficial value
Request body48 MiB
Base64 or URL image32 MiB maximum
Files API image64 MiB maximum
Number of images600 per request
Total image size64 MiB without file_id, 200 MiB with it
Dimensions8192 px per side, 4096 px from 15 images onward

In Chat Completions, images must appear in user messages. Putting one in a system or assistant message returns a 400 error. You must also call the Vision model explicitly. Other DeepSeek models reject image blocks.

!
Production

OpenAI compatibility makes testing easy, but it does not magically turn an experimental model into a reliable component. Add evaluation sets, business validation and a fallback model for sensitive cases.

My verdict

This is finally the model DeepSeek had been missing all along. DeepSeek already had the price, reasoning, code, tools and context. Vision finally gives it eyes.

The most interesting use is not asking what a photo contains. It is letting the model observe a business element, reason about it and act within the same loop. At this price, many multimodal routing layers that automatically reserved vision for a more expensive provider deserve to be reconsidered.

I would not call the race won yet. The claimed proximity to Opus 4.8 needs independent tests and real workloads. Still, DeepSeek has closed the biggest gap in its lineup without sacrificing the pricing that made it powerful. This release deserves an early test.

Quick questions

What is the exact DeepSeek Vision model name?

The model is named deepseek-v4-flash-vision-exp. Use that identifier in API requests.

How much does image analysis cost?

An image accounts for at most 384 tokens. It is billed at the V4 Flash input rate, plus the text you send and the output the model generates.

Which image input method should you use?

An external URL is convenient for a quick test. Base64 works well for a small local file. The Files API is better for a large image or one reused across requests.

Sources