深潮TechFlow
深潮TechFlow|Aug 21, 2026 09:30
[V4-Flash-Vision-Exp Launched, Enabling Multimodal API Services] According to TechFlow, on August 21, DeepSeek announced that the new multimodal visual understanding model, DeepSeek-V4-Flash-Vision-Exp, has been launched on the DeepSeek API platform. This model is experimental in nature, and users can invoke it by setting model="deepseek-v4-flash-vision-exp". Its pure text capabilities are on par with the official version of DeepSeek-V4-Flash, while demonstrating significant improvements in visual understanding benchmarks for Agents, with multimodal Agent capabilities approaching Opus-4.8. The model supports three invocation formats: Chat Completions, Messages, and Responses. It can be integrated with various Agent tools, supports mixed text and image input, and allows image input through three methods: base64 inline, external URL, and Files API. DeepSeek has also opened the Files API, enabling users to upload images and reference them via file_id to reduce redundant uploads and bandwidth usage.
+2
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads