Generate multi-minute videos from text or images with a 13.6B parameter open-source model
You can generate long videos that maintain color consistency and quality using simple prompt inputs. This video generation tool integrates text, image, and existing video processing into a single model. Recently, an avatar feature that matches lip movements to audio has been added, expanding its range of applications.
Latest Update: Avatar 1.5 Release
LongCat-Video-Avatar-1.5 was released on May 21, 2026. Unlike previous versions, it uses Whisper-Large instead of Wav2Vec2 to improve lip-sync accuracy. It supports stylized domains such as animation and animals, and handles both single and multiple audio inputs. By applying step distillation technology, it reduces inference steps to 8, significantly improving speed.
Long Video Generation and Efficient Inference
Pre-trained on video continuation tasks, this model creates videos several minutes long without color fading or quality degradation. It employs a strategy that generates from coarse to fine along temporal and spatial axes, producing 720p, 30fps videos in a few minutes. Block Sparse Attention is applied to maintain efficiency even at high resolutions. The documentation states that internal and public benchmark evaluations show performance comparable to major open-source models and latest commercial solutions.
Installation and Execution
You must create a virtual environment with conda in a Python 3.10 environment and then install torch and flash-attn. Model weights are downloaded from Hugging Face using huggingface-cli. Scripts corresponding to each task, such as text-to-video, image-to-video, and video continuation, are executed with torchrun. It can be run in both single and multi-GPU environments through context parallel processing. A Streamlit-based web interface is also provided, allowing testing without writing code directly.
Tips for Using the Avatar Feature
For audio-to-video generation, setting the audio CFG value between 3 and 5 yields the best lip-sync accuracy. Including detailed descriptions of character appearance, actions, and scene context in the prompt improves consistency and naturalness. To reduce repetitive motions, adjust --ref_img_index between 0 and 24 and --mask_frame_range appropriately. In version 1.5, the --use_distill option must be added, and --use_int8 can be used together to reduce VRAM usage.
Pre-Application Checks
Model weights are released under the MIT license, allowing commercial use. However, since not all sub-applications have been comprehensively evaluated, you must directly verify accuracy and safety before deploying in sensitive or high-risk scenarios. It is the developer's responsibility to comply with laws regarding data protection and content safety. It is recommended to determine the scope of use considering general limitations of large language models, such as performance differences by specific language.