AI Models & Platforms
Vidu Releases Q4 Preview of Next-Generation Flagship AI Video Model

ShengShu Technology launched Vidu Q4 Preview on October 7, 2026, the first public preview of its next-generation flagship model for expressive audio and video. Launch pricing starts at $0.014 per second, and the preview is available through Vidu’s web product and API platform.
The model supports up to 15 image references and up to three audio references, with output at 2K or 4K resolution and 10-bit color depth. The company said the preview targets independent creators, small studios and production teams spanning both narrative and commercial work.
Character Performance, Camera Work and Visual Effects
ShengShu Technology said Q4 Preview is built to better coordinate facial expression, emotion, body movement and voice, yielding subtler expressions, clearer emotional reads and more natural movement. As many as three reference audio clips help keep a character’s voice consistent and its delivery emotionally aligned, while up to 15 reference images can pin down characters, wardrobe, props, products and environments within a single creative setup. The company said these controls are meant to hold visual and vocal choices together across a scene and to give narrative projects more deliberate control over staging and performance.
The announcement describes tighter coordination among camera movements, cuts and on-screen action: in fast-moving sequences such as fights and chases, the camera is designed to stay with the action more coherently, and effects such as explosions, fireworks and particles blend more naturally into their surroundings. The 2K and 4K modes arrive alongside better lighting and overall aesthetics, with the wider resolution range serving both early-stage creative tests and final high-resolution output.
“Advanced video models should prove themselves beyond carefully selected demos. Every improvement in efficiency should ultimately give creators the freedom to try one more shot or make one more version,” said Yihang Luo, co-founder and CEO of ShengShu Technology, in the launch announcement.
Pricing and Intended Use
Output options span 540p, 720p, 1080p, 2K and 4K. ShengShu Technology said that when output specifications and billing conditions are comparable, Q4 Preview delivers up to five times the output for the same budget. The company tied the pricing to the realities of iteration: getting a production-ready shot typically takes more than one attempt, whether that means adjusting a character’s expression, evaluating alternative camera angles or experimenting with different openings. As examples of intended use, it pointed to advertising teams weighing creative concepts and opening sequences, e-commerce teams building variations for different products and audiences, and filmmakers testing performances, storyboards and pacing before committing further to production.
The announcement notes that final pricing, supported resolutions, feature availability and usage terms may vary by plan and region, with full pricing details and promotional terms published on Vidu’s website.
Product Modes and API Specifications
Vidu’s official product page lists Image-to-Video and Reference-to-Video modes for Q4: Image-to-Video animates a single uploaded image from a text prompt, while Reference-to-Video combines one to 15 image references with zero to three audio references to keep subjects and voices consistent. On the product page, Reference-to-Video is listed with durations of 1 to 16 seconds and Image-to-Video with 3 to 16 seconds, and the page’s highlights include a 4K maximum resolution and a 16-second maximum video length. Output preserves the original aspect ratio of uploaded material, and the page names AI series, advertising, social content and cinematic production as the scenarios the model is designed for.
For developers, the image-to-video documentation lists the model under the name viduq4-preview at POST https://api.vidu.com/ent/v2/img2video. The endpoint takes a single start-frame image in png, jpeg, jpg or webp format up to 50MB, a prompt of up to 20,000 characters, durations of 3 to 16 seconds with a default of 5, and resolutions from 540p up to 4K with a default of 720p. An audio parameter defaults to true, producing video with sound that can include dialogue and sound effects, and the documentation states the model supports audio-video synchronization and automatic camera switching.
The reference-to-video documentation lists POST https://api.vidu.com/ent/v2/reference2video, which accepts one to 15 images and zero to three mp3 audio references of 3 to 12 seconds each, with aspect ratio options of 1:1, 9:16, 16:9, 3:4 and 4:3 and a default of 16:9. Prompts can embed tags for the reference subjects, and the documentation states the model supports generating video with reference sound, intelligent camera switching, consistency across multiple camera positions and simultaneous audio and video output. Returned creation URLs remain valid for 24 hours.
Vidu is releasing Q4 Preview ahead of the full Q4 model so creators can apply it in real projects, and ShengShu Technology said input on performance, creative control and day-to-day production needs will shape the final release.












