Qwen3.8-Flash-Next is a 125B multimodal MoE with only 6B active parameters and a new architecture built around QSA, Gated Residual, N-gram embeddings, and Muon. Its open weights give an early look at the architecture Qwen is building toward Qwen4.
Qwen has done this once before. Qwen3-Next gave everyone an early look at the architecture that later showed up in Qwen3.5.
Qwen3.8-Flash-Next is doing the same for Qwen4.
It’s a 125B model with only 6B active parameters, and a lot of the new design is about getting more capability without dragging compute up with it. QSA makes long-context retrieval cheaper, Gated Residual gives information more paths through the network, and the new N-gram memory adds capacity with very little per-token compute.
Qwen keeps putting these architecture previews out as real models people can actually run, well before the next generation arrives.
the gated residual + n-gram memory combo is the part i'd want to poke at. changing how information flows through the network usually means old LoRA adapters and fine-tunes built for Qwen3 don't transfer cleanly to the new architecture. is that the tradeoff here, or did you design it so existing Qwen3 tooling and adapters still mostly work on top of this
Congrats on the launch! love that these previews are actual downloadable models.
the active param count is spot on — 125B total is what kills it for local though. any chance of a ~35-40B/A3-6B version on this architecture? 35B-A3B was the local goat and it has no successor yet 🙏
About Qwen3.8-Flash-Next on Product Hunt
“The open-weight preview of Qwen4”
Qwen3.8-Flash-Next launched on Product Hunt on August 27th, 2026 and earned 107 upvotes and 4 comments, placing #15 on the daily leaderboard. Qwen3.8-Flash-Next is a 125B multimodal MoE with only 6B active parameters and a new architecture built around QSA, Gated Residual, N-gram embeddings, and Muon. Its open weights give an early look at the architecture Qwen is building toward Qwen4.
Qwen3.8-Flash-Next was featured in Open Source (68.8k followers) and Artificial Intelligence (477.8k followers) on Product Hunt. Together, these topics include over 134.2k products, making this a competitive space to launch in.
Who hunted Qwen3.8-Flash-Next?
Qwen3.8-Flash-Next was hunted by Zac Zuo. A “hunter” on Product Hunt is the community member who submits a product to the platform — uploading the images, the link, and tagging the makers behind it. Hunters typically write the first comment explaining why a product is worth attention, and their followers are notified the moment they post. Around 79% of featured launches on Product Hunt are self-hunted by their makers, but a well-known hunter still acts as a signal of quality to the rest of the community. See the full all-time top hunters leaderboard to discover who is shaping the Product Hunt ecosystem.
Want to see how Qwen3.8-Flash-Next stacked up against nearby launches in real time? Check out the live launch dashboard for upvote speed charts, proximity comparisons, and more analytics.
Hi everyone!
Qwen has done this once before. Qwen3-Next gave everyone an early look at the architecture that later showed up in Qwen3.5.
Qwen3.8-Flash-Next is doing the same for Qwen4.
It’s a 125B model with only 6B active parameters, and a lot of the new design is about getting more capability without dragging compute up with it. QSA makes long-context retrieval cheaper, Gated Residual gives information more paths through the network, and the new N-gram memory adds capacity with very little per-token compute.
Qwen keeps putting these architecture previews out as real models people can actually run, well before the next generation arrives.
And the weights are already up :) (with a license worth reading)