On February 26, Alibaba (BABA) announced the open source of its base model Wanxiang 2.1 (Wan). In the evaluation set, it surpassed Sora, Luma and other models and ranked first.
The strongest open source video model debuted
It is learned that Wanxiang 2.1 has two parameter scales. The 14 billion parameter model is suitable for professionals with higher requirements for generation effects. The 1.3 billion parameter model has a faster generation speed and is compatible with all consumer-grade GPUs. All inference codes and weights of the two models have been open sourced.

In terms of video generation, Wanxiang 2.1 enhances the spatiotemporal context modeling capabilities through the self-developed efficient VAE and DiT architecture, supports efficient encoding and decoding of infinite length 1080P videos, and realizes the Chinese text video generation function for the first time. It also supports multiple tasks such as text-generated video, image-generated video, video editing, text-generated image, and video-generated audio.
According to previous introductions, Wanxiang 2.1 supports Chinese and English videos, can generate artistic words with one click, and also provides a variety of video special effects options to enhance visual expression, such as transitions, particle effects, simulations, etc.
Analysts said that with the open source of the Wanxiang 2.1 model, Alibaba Cloud has achieved full-modal and full-size open source. This means that more developers will be able to obtain and use the underlying code of the model at a low cost, and then use it to carry out various video generation applications related to their own business.

Opening a new era of full-modal open source
Since 2025, the open source trend has gradually become the standard in the field of large models around the world. Domestically, in February, many companies have launched their own open source models, including ByteDance’s Doubao and Baidu’s Wenxin Yiyan, which have jointly set off a new round of open source craze.
Internationally, with the complete open source of Wanxiang 2.1, competitors such as OpenAI and Google will also face the challenge of commercialization: better models have been open sourced, and the pricing of AI-generated videos will also face challenges. Google’s Veo 2 model recently disclosed pricing, and it costs $0.5 to generate 1 second of video, which is equivalent to $1,800 to generate one hour of video.
WIMI open source multimodal application scenario expansion

Public information shows that WiMi Hologram Cloud Inc. (NASDAQ: WIMI) has a significant layout in the field of AI video generation, covering large language, multimodal and other fields. Facing the open source video generation large model track, from large language models to visual generation models, from basic models to diversified derivative models, it has achieved full-modal and full-size open source, and the development of WIMIAI open source ecology is constantly injected with powerful momentum.
In fact, in recent years, WIMI has focused on the research and development of multimodal AIGC (generative AI). The core technology lies in combining large-scale pre-training with multimodal algorithm optimization to improve the coherence and physical rationality of generated content. At the same time, in the industry ecology, WIMI has gradually realized the ability to generate videos from text and images, and supports scenarios such as plot creation and short video generation. In the future, it may accelerate the technical iteration of AI’s ability to quickly generate videos through APIs or industry solutions.
Conclusion
In the future, AI models will enter a watershed. Institutions generally believe that Alibaba’s move will accelerate the commercialization of AI video technology and promote the upgrading of the entire industry chain, including computing power, cloud computing, and content creation. Therefore, the second half of AI is not a simple technology competition, but a comprehensive game about resources, efficiency, and cost. This new revolution is accelerating.




