Qwen 3.8-Flash-Next

Released on August 26, 2026

Changes from previous version

Unlike dense Qwen3.8-27B, Flash-Next is a sparse hybrid-attention MoE that previews Qwen4: QSA micro-block attention, gated residuals, and n-gram embeddings for cheaper long-context agents.

Release Summary

Open-weight preview of the Qwen4 architecture: a 125B MoE (6B active) plus 51B n-gram embeddings, hybrid Gated DeltaNet and Qwen Sparse Attention, and native image and video input. 262K native context (extensible to 1M). Qwen reports 62.5 on SWE-bench Pro, 58.7 on DeepSWE v1.1, and 81.0 on SWE-bench Multilingual, ahead of Qwen3.8-27B on long-horizon coding and agent tasks.

Timeline

August 26, 2026

Qwen3.8-Flash-Next released

Qwen publishes Flash-Next weights on Hugging Face and ModelScope as an early look at the Qwen4 architecture.