DeepSeek Releases V4 Flash Vision Exp for Multimodal Agent Work

DeepSeek has released V4 Flash Vision Exp, an experimental API model that adds image understanding to the V4 Flash family while keeping its text capabilities.
Synthesized via Fish Audio S2.1 Pro in British English. Natural editorial summary, not verbatim reading.
Key Takeaways
- check_circleDeepSeek dated the release 21 August 2026 and made the experimental model available through its API as deepseek-v4-flash-vision-exp.
- check_circleThe model accepts image and text input through DeepSeek's Chat Completions, Messages and Responses formats. The documentation lists JPEG, PNG, GIF and WebP support.
- check_circleDeepSeek reports results on coding and visual-agent benchmarks, but those figures come from the maker's own evaluation setup and are not independent validation.
- check_circleAIMI first saw the official DeepSeek API route at 10:19:21 SAST, one hour, 47 minutes and 21 seconds after OpenCode announced Ox Alpha Free.
DeepSeek adds vision to V4 Flash
DeepSeek's API change log dated 21 August describes V4 Flash Vision Exp as a new experimental multimodal vision-understanding model. The release uses the model ID deepseek-v4-flash-vision-exp, so developers can select it without guessing at a provider-specific alias.
The official vision guide says the model accepts images alongside text. It can describe pictures, read text from screenshots and analyse charts. Supported image formats are JPEG, PNG, GIF and WebP. DeepSeek documents the same image input through its OpenAI-compatible Chat Completions format and its Responses API, while the platform page also lists Messages format support.
What DeepSeek reports
DeepSeek reports 83.9 on Terminal Bench 2.1, 57.7 on NL2Repo and 64.3 on Chartography. It also says the model is on par with the official V4 Flash on pure-text tasks and improves visual agent work compared with the text-only model.
Those numbers are vendor-reported results. The change log says the code-agent tests used DeepSeek Harness minimal mode with max effort, topp 0.95 and temperature 1.0. That context matters when comparing the figures with results from another lab or serving stack. The article does not treat DeepSeek's claim about being close to Opus-4.8 as an independent comparison.
The AIMI route timeline
The maker announcement has a calendar date but no publication time. AIMI first saw the official DeepSeek API route at 08:19:21 UTC, or 10:19:21 SAST. OpenRouter added the same canonical model at 11:48:48 UTC, or 13:48:48 SAST, which was 3 hours, 29 minutes and 27 seconds after the official API route. OpenRouter recorded the provider-created timestamp as 11:26:03 UTC, but that is provider metadata, not the maker announcement time.
This release followed OpenCode's Ox Alpha Free announcement at 08:32:00 SAST by exactly 1 hour, 47 minutes and 21 seconds. That makes DeepSeek the next major release signal in the same-day run. The timing is useful context, but neither primary source says that DeepSeek was responding to OpenCode or that OpenCode was responding to DeepSeek.
How to try the model
DeepSeek's API uses the base URL https://api.deepseek.com. Set the model to deepseek-v4-flash-vision-exp and send an image as an image content block alongside the text prompt. The official vision guide documents inline base64 images, image URLs and uploaded files.
AIMI also recorded an OpenRouter route under deepseek/deepseek-v4-flash-vision-exp. At the time of the observation, OpenRouter listed paid pricing of $0.22 per million input tokens and $0.66 per million output tokens. Provider pricing can change, so check the live route before building a cost estimate.
What has not been announced
The official release is an API launch. The sources checked for this article do not announce open weights, a downloadable checkpoint or a separate app and web rollout. The model should therefore be described as an API release, not as a new open-weight model.
The DeepSeek API page welcomes testing and feedback. Teams evaluating it should use their own screenshots, charts and browser-agent tasks, then compare successful task completion, review time and cost against the text-only V4 Flash route.
Frequently Asked Questions
Is DeepSeek V4 Flash Vision Exp a real new release?
Yes. DeepSeek's official API change log dated 21 August 2026 names it as a new release and gives the model ID deepseek-v4-flash-vision-exp.
What can the model see?
DeepSeek documents image input for JPEG, PNG, GIF and WebP files. Its examples cover image descriptions, screenshot text extraction and chart analysis.
Are the published benchmark numbers independent?
No. The benchmark results in the release note are DeepSeek's own reported results and use the evaluation settings described in that note.
Does this release include open weights?
No open-weight release was announced in the official sources checked for this article. The confirmed release is access through the DeepSeek API.
Is the OpenRouter route a separate model?
AIMI found the same canonical model ID on OpenRouter. That confirms provider availability, but it does not create a second maker release or change DeepSeek's 21 August announcement date.
Related Frontier Models & Releases
Explore verified specifications, benchmark results, and route pricing across alternative models in this class.
DeepSeek releases V4.1 Flash with native multimodal vision, 552B MoE and lower API rates
DeepSeek has launched DeepSeek-V4.1-Flash, a 552B-parameter mixture-of-experts model featuring a novel Causal Encoder–Decoder architecture with just 8B active input and 16B active output parameters. The release brings native vision, compresses KV cache storage by up to 8x, and slashes off-peak API pricing to $0.15 per million input tokens.
DeepSeek V4 Pro Reaches GA with Adjustable Reasoning and Responses API Support
DeepSeek has released the GA version of V4 Pro for its app, web service and API. The 0813 model adds adjustable reasoning, native Responses API support, and a later hosted route on NVIDIA NIM.
OpenAI Releases GPT-6 Astra: Next-Generation Flagship with 1.05M Context and Deep Multimodal Reasoning
OpenAI has officially launched GPT-6 Astra, its frontier flagship model featuring a 1,050,000-token context window, 128,000 max output tokens, and native tool-use for autonomous agent workflows.