DeepSeek Open-Sources V4-Flash-Vision-Exp: 305B Native Multimodal MoE on Hugging Face


On August 31, 2026, Chinese AI research lab DeepSeek officially published the open-weight checkpoints for DeepSeek-V4-Flash-Vision-Exp, its flagship multimodal foundation model, under a permissive MIT license on Hugging Face.
The release transitions the 305-billion-parameter Mixture-of-Experts (MoE) model from API-only availability to full self-hosted enterprise deployment. Unlike earlier iterations that relied on modular external visual projectors, V4-Flash-Vision-Exp incorporates a natively trained vision encoder integrated directly into the core transformer backbone. This architecture allows the model to process high-resolution screenshots, financial charts, and UI wireframes with minimal token latency while retaining the agentic text performance of the base V4-Flash checkpoint.
DeepSeek released both full-precision weights and optimized FP8 and GGUF quantized builds on Hugging Face (deepseek-ai/DeepSeek-V4-Flash-Vision-Exp). The model supports DSpark speculative decoding and integrates into production inference engines including SGLang, vLLM, and TensorRT-LLM, allowing teams to run multimodal agentic loops on dual GPU accelerator nodes like the Nvidia DGX Spark.
The MIT-licensed weight drop intensifies competition across open-source visual reasoning, offering developers a zero-royalty alternative to proprietary vision systems like OpenAI's GPT-5.6 Terra and Google's Gemini 3.6 Flash.


