Aakib Ansari.
Back to articles
News Brief

Alibaba Releases Qwen3.8-Flash-Next: 125B MoE Previews Qwen4 Architecture with 6B Active Parameters

Md Aakib Ansari
Md Aakib AnsariWeb Developer & AI Tools Reviewer
•3 min read•Model: Qwen3.8-Flash-Next•Company: Alibaba
Alibaba Releases Qwen3.8-Flash-Next: 125B MoE Previews Qwen4 Architecture with 6B Active Parameters

On August 26, 2026, Alibaba Cloud and the Qwen research team officially released Qwen3.8-Flash-Next, an open-weight 125-billion-parameter multimodal Mixture-of-Experts (MoE) model that serves as the architectural preview for the upcoming Qwen4 generation.

Published under the Qwen Community License 1.0 on Hugging Face and ModelScope, the model features an efficient sparse routing design that activates just 6 billion parameters per token alongside a 51-billion-parameter N-gram embedding table. The architecture replaces standard dense transformers with a hybrid design combining Gated DeltaNet linear attention and Qwen Sparse Attention (QSA), designed to slash pre-training compute to roughly one-ninth of previous Qwen generations while preserving high-throughput inference on enterprise and consumer GPUs.

On vendor-reported benchmarks published in Alibaba's technical release notes, Qwen3.8-Flash-Next achieved 81.0 on SWE-bench Multilingual and 91.9 on LiveCodeBench v6, matching the performance of much larger models like Qwen3.7-Plus on repository-level code generation. The checkpoint ships with native multimodal vision encoding and a 262,144-token context window that extends up to 1 million tokens for large codebase analysis.

The launch achieved immediate day-zero support across open-source serving runtimes, including SGLang, vLLM, Ollama, and Unsloth, alongside FP8 and NVFP4 quantized weights for local deployment on NVIDIA hardware. This release continues Alibaba's aggressive cadence of open-weight drops following the earlier Qwen 3.8-27B dense release and Qwen 3.8 Max open weights.

By offering 6B-scale execution speeds backed by 125B total capacity, Qwen3.8-Flash-Next provides developer teams with a practical preview of next-generation MoE efficiency for high-frequency coding agents and multimodal tool loops.

Frequently Asked Questions

What is Qwen3.8-Flash-Next and how does its architecture work?
Qwen3.8-Flash-Next is Alibaba's open-weight MoE model previewing the Qwen4 architecture. It pairs 125B total parameters with a 6B active parameter routing budget using hybrid Gated DeltaNet and Qwen Sparse Attention.
Where can developers download and run Qwen3.8-Flash-Next?
Weights are available on Hugging Face and ModelScope under the Qwen Community License 1.0, with day-zero support in vLLM, SGLang, Ollama, and GGUF quantization formats.

Related Articles