Runtime as featured inForbesRead the article
All open source models
Qwen (Alibaba) logo

Qwen3.8 Flash Next

Qwen (Alibaba) · #10 in our ranking · Most efficient long context

Qwen3.8 Flash Next is an open-weight mixture-of-experts model from Qwen (Alibaba), released Aug 2026, with 180B parameters, a 256K-token context window, and the Qwen Community License.

Qwen3.8 Flash Next pairs Gated DeltaNet with Qwen Sparse Attention to cut long-context latency, and adds n-gram embeddings for cheaper parameter scaling. Alibaba positions it as a preview of the Qwen4 architecture.

Qwen3.8 Flash Next specs

Context256K tokens
Parameters180B
Provider
Qwen (Alibaba)
Architecture
Mixture-of-experts
Parameters
180B total
Context
256K tokens | 197K words
License
Qwen Community License
Input
Text, Image
Released
Aug 2026

Sourced from the model card on Hugging Face. Last checked 2026-10-01.

Highlights

  • Hybrid attention for low long-context latency
  • Image input
  • Preview of the Qwen4 architecture

Where it fits for payment teams

Worth testing for agents that read long ledgers or logs where latency adds up across many runs.

Harnesses in Runtime

Run Qwen3.8 Flash Next through any of these harnesses on Runtime.

OpenCodeCline

How to run Qwen3.8 Flash Next

01

Download the weights

Pull Qwen/Qwen3.8-Flash-Next from Hugging Face. Check the Qwen Community License before you deploy.

02

Serve it in your cloud

Run it with vLLM or SGLang on your own GPUs, or use a managed provider that hosts the model.

03

Put it to work in Runtime

Point an agent or a single skill at the model through OpenCode, then compare it to your current model with evals.

Qwen3.8 Flash Next FAQ

What is the context window of Qwen3.8 Flash Next?

+

Qwen3.8 Flash Next supports a 256K-token context window (262,144 tokens), roughly 197K words of English text.

How many parameters does Qwen3.8 Flash Next have?

+

Qwen3.8 Flash Next is a mixture-of-experts model with 180B total parameters.

Is Qwen3.8 Flash Next free for commercial use?

+

Qwen3.8 Flash Next is released under the Qwen Community License, a custom license. Commercial use may come with conditions, so read the license on Hugging Face before you deploy or fine-tune it.

What can Qwen3.8 Flash Next take as input?

+

Qwen3.8 Flash Next accepts text, image input and generates text.

Can I run Qwen3.8 Flash Next in my own cloud with Runtime?

+

Yes. Download the weights from Hugging Face (Qwen/Qwen3.8-Flash-Next) and serve them with vLLM or SGLang in your cloud, or use a managed provider that hosts the model. Runtime agents can then use Qwen3.8 Flash Next through a harness like OpenCode, with your data staying in your environment.