Published by AICodeNews Editorial Team | August 21, 2026
In a major multimodal upgrade to its low-cost model lineup, DeepSeek has officially released DeepSeek-V4-Flash-Vision-Exp, adding high-speed visual reasoning and chart analysis while matching the full text and coding performance of the base V4-Flash model.
Announced directly on the official DeepSeek API Platform, DeepSeek-V4-Flash-Vision-Exp elevates open-weights multimodal agent performance, closing the capability gap with closed models on benchmark suites like Terminal Bench (83.9) and ApexBench (36.5).
1. Multimodal Benchmarks & Vision Capabilities for DeepSeek-V4-Flash-Vision-Exp

While previous lightweight vision models struggled with dense UI layouts and technical schematics, DeepSeek-V4-Flash-Vision-Exp is engineered for autonomous agent computer use and document parsing:
- Terminal Bench Score (83.9): Excels at interpreting CLI terminal outputs, error stack traces, and multi-window developer workflows.
- ApexBench Visual Accuracy (36.5): High precision in extracting data from multi-column PDF tables, system architecture flowcharts, and financial charts.
- Zero Compromise on Text & Code: Retains 100% of the reasoning, coding logic, and world knowledge of the standard V4-Flash model.
2. API Architecture, 384-Token Pricing & The New Files API
Developers can begin using DeepSeek-V4-Flash-Vision-Exp immediately across production developer endpoints using standard Chat Completions, Messages, and Responses APIs:
| Feature / Parameter | Specification / Capability |
|---|---|
| Model Identifier | deepseek-v4-flash-vision-exp |
| Image Tokenization & Pricing | Capped at max 384 tokens per image at standard V4-Flash pricing rates |
| Input Formats | Mixed text + image via Base64, external image URLs, or the new Files API |
| New Files API | 100% Free: Upload an image once, reference by file_id to save request payload bandwidth |
| Ecosystem Support | Day-one support in DeepSeek Harness 0.1.1, OpenCode, and OpenRouter |
The companion release of DeepSeek Harness 0.1.1 provides out-of-the-box support, allowing autonomous agents to stream visual desktop frames and inspect local code artifacts without manual token chunking.
3. Key Takeaways on DeepSeek-V4-Flash-Vision-Exp
- Experimental Multimodal Release: DeepSeek-V4-Flash-Vision-Exp brings image comprehension to the low-cost V4-Flash tier.
- Predictable Image Billing: Flat 384-token billing cap per image makes high-volume visual agent pipelines economical.
- Free Files API: Eliminates redundant image re-uploads by referencing persistent file IDs across multi-turn chat sessions.
Bookmark AICodeNews.com for daily coverage on multimodal AI models, agent frameworks, and developer API releases.


Comments
One response to “DeepSeek-V4-Flash-Vision-Exp Launches: Multimodal Vision Reasoning & Free Files API”
[…] DeepSeek-V4-Flash […]