DeepSeek-V4-Flash-Vision-Exp Launches: Multimodal Vision Reasoning & Free Files API

DeepSeek-V4-Flash-Vision-Exp

Published by AICodeNews Editorial Team | August 21, 2026

In a major multimodal upgrade to its low-cost model lineup, DeepSeek has officially released DeepSeek-V4-Flash-Vision-Exp, adding high-speed visual reasoning and chart analysis while matching the full text and coding performance of the base V4-Flash model.

Announced directly on the official DeepSeek API Platform, DeepSeek-V4-Flash-Vision-Exp elevates open-weights multimodal agent performance, closing the capability gap with closed models on benchmark suites like Terminal Bench (83.9) and ApexBench (36.5).

1. Multimodal Benchmarks & Vision Capabilities for DeepSeek-V4-Flash-Vision-Exp

DeepSeek-V4-Flash-Vision-Exp benchmarks

While previous lightweight vision models struggled with dense UI layouts and technical schematics, DeepSeek-V4-Flash-Vision-Exp is engineered for autonomous agent computer use and document parsing:

  • Terminal Bench Score (83.9): Excels at interpreting CLI terminal outputs, error stack traces, and multi-window developer workflows.
  • ApexBench Visual Accuracy (36.5): High precision in extracting data from multi-column PDF tables, system architecture flowcharts, and financial charts.
  • Zero Compromise on Text & Code: Retains 100% of the reasoning, coding logic, and world knowledge of the standard V4-Flash model.

2. API Architecture, 384-Token Pricing & The New Files API

Developers can begin using DeepSeek-V4-Flash-Vision-Exp immediately across production developer endpoints using standard Chat Completions, Messages, and Responses APIs:

Feature / ParameterSpecification / Capability 
Model Identifierdeepseek-v4-flash-vision-exp
Image Tokenization & PricingCapped at max 384 tokens per image at standard V4-Flash pricing rates
Input FormatsMixed text + image via Base64, external image URLs, or the new Files API
New Files API100% Free: Upload an image once, reference by file_id to save request payload bandwidth
Ecosystem SupportDay-one support in DeepSeek Harness 0.1.1, OpenCode, and OpenRouter

The companion release of DeepSeek Harness 0.1.1 provides out-of-the-box support, allowing autonomous agents to stream visual desktop frames and inspect local code artifacts without manual token chunking.

3. Key Takeaways on DeepSeek-V4-Flash-Vision-Exp

  • Experimental Multimodal Release: DeepSeek-V4-Flash-Vision-Exp brings image comprehension to the low-cost V4-Flash tier.
  • Predictable Image Billing: Flat 384-token billing cap per image makes high-volume visual agent pipelines economical.
  • Free Files API: Eliminates redundant image re-uploads by referencing persistent file IDs across multi-turn chat sessions.

Bookmark AICodeNews.com for daily coverage on multimodal AI models, agent frameworks, and developer API releases.

Comments

One response to “DeepSeek-V4-Flash-Vision-Exp Launches: Multimodal Vision Reasoning & Free Files API”