Alibaba launches AI model that turns PDFs into videos
What's the story
Alibaba has launched its latest AI video-generation model, Wan 3.0. The new system is capable of creating clips of up to 30 seconds from a range of inputs including text, images, audio, videos and documents like PDFs and presentations. The company claims that Wan 3.0 can generate a 30-second video in a single pass, double the capacity of its predecessor, Wan 2.7 which had a limit of 15 seconds per clip.
Technological advancements
Wan 3.0 ensures consistency across video content
The new model supports output up to 1080p and is designed to ensure consistency across a video, including characters, objects, backgrounds and spatial layouts.
It can also create faces with synchronized micro-expressions and supports multilingual voice output.
One of the standout features of Wan 3.0 is its ability to convert existing information into video content.
Users can input a PDF, PowerPoint presentation or spreadsheet into the model, which will then generate a corresponding video from that material.
Market positioning
The new model is competitively priced against US counterparts
Alibaba has priced its video generation model competitively against US counterparts.
The company charges $0.05 per second for 480p video, $0.10 per second for 720p and $0.20 per second for 1080p output from the model. For context, Google's Veo 3.1 standard tier costs $0.40 per second of usage.
Wan 3.0 was made available to public beta users on August 6 through Alibaba Cloud's Model Studio and Qwen Cloud, with wider access now being rolled out by the company.