DeepSeek has made the weights for DeepSeek-V4-Flash-Vision-Exp available through Hugging Face, according to a PANews report on Aug. 31. The multimodal model (AI handling text and images) builds on the V4-Flash architecture by adding image understanding, including the ability to interpret screenshots and charts. It can also work with tools to perform agent tasks. On text-only agent benchmarks such as Terminal Bench, its performance was roughly comparable to V4-Flash. The model series had previously been available only through an API, but developers can now download the weights for local deployment and integrate them through Transformers, vLLM or SGLang.