律动BlockBeats
律动BlockBeats|8月 21, 2026 09:13
[DeepSeek Finally Moving to Multimodal? New Vision Model Added to Official Harness] Beating AI News Flash: A new vision model, deepseek-v4-flash-vision-exp, from DeepSeek is starting to surface. Some community members have already tested it via the API, allowing direct image-based queries. A more direct signal is that DeepSeek Harness has added this model to the default model directory today, explicitly marking it as supporting image input. The corresponding code PR is titled 'publish the vision model.' Currently, DeepSeek has not officially announced this, and the official API documentation has yet to reflect it. DeepSeek has actually been capable of image recognition for some time. Both the web and app versions have already launched image recognition modes. Previously, DeepSeek's vision team publicly shared research based on V4-Flash, enabling the model to insert points and bounding boxes directly into the reasoning chain—essentially 'pointing' at the image while completing visual CoT (Chain-of-Thought) reasoning. This release has been hinted at for a while. Two days ago, DeepSeek Harness RC.8 added native image input support to the official DeepSeek adapter. The implementation notes at the time already mentioned deepseek-v4-flash-vision-exp but deliberately excluded it from the default model directory, citing the need to wait for the model endpoint to be ready. Two days later, this restriction has been officially lifted. Harness v0.1.1-rc.1 has now listed deepseek-v4-flash-vision-exp as the default vision model. Previously, it was 'waiting for the endpoint to be ready,' but now it has directly entered the official model list. The new vision model is about to be released. [Original Link]
+3
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads