Ali Qianwen released the Qwen3.8-Flash model
Ali Qianwen released the Qwen3.8-Flash model, which is described as a multimodal MoE model and an early preview of the Qwen4 architecture.
The production version of Qwen3.8-Flash will soon be available through the Qwen Cloud API, with a price of only $0.16 per million input tokens and $0.47 per million output tokens. The model has 125 billion parameters + 51 billion N-gram embedding parameters, but only activates 6 billion parameters per token, achieving extremely high cost-effectiveness.
Related tags






