Automate video production from simple ideas to final output
Imagine entering a single scene from your mind and having the system handle script writing, storyboard design, character consistency, and final video synthesis on its own. ViMax is a tool that bundles these complex video production steps into a single agent system. Users can receive a finished video by specifying only their desired style and requirements, without needing to know the technical details.
v1.2.0 Update and Web-Based Workspace
The recent v1.2.0 version introduces a web-based user interface, allowing project management, agent conversation, storyboard preview, and rendering checkpoint verification to be handled on a single screen. It also adds support for OpenRouter GPT Image 2 image generation and Seedance 2.0 Fast video generation, broadening the range of selectable models. Previous versions provided session reuse and context compression features through a terminal-based interface (TUI), and this update allows the same agent runtime and configuration files to be shared and used in a browser environment.
Four Paths from Idea to Video
ViMax provides four main workflows depending on the input format. Idea2Video converts short concepts into structured stories, characters, scripts, storyboards, shots, and final videos. Script2Video controls multiple scenes and shots based on an already written script while preserving creative intent. Novel2Video compresses long novels into episodic visual narratives and performs character tracking and scene planning. AutoCameo inserts people or pets from reference photos into generated stories with consistent appearances.
Consistency Maintenance and Parallel Generation Mechanism
To address the issue where existing AI video tools only generate short clips and characters or scenes change frame by frame, ViMax features a pipeline that adjusts reference images, first frames, camera continuity, and final assembly to the end. It improves visual quality and consistency by generating two candidate key frames in parallel during image generation and having a vision-capable chat model select one. It also accelerates multi-shot video production by generating compatible shots and media assets simultaneously.
Execution Environment and Configuration
ViMax is written in Python and runs on Linux and Windows. It uses uv for environment management, and dependencies are installed with the uv sync command after cloning the repository. To use the TUI, copy configs/agent.example.yaml to configs/agent.local.yaml and configure the model information and API keys for the LLM, image generator, and video generator respectively. To use the web UI, Node.js 18 or higher is required; run npm install and npm run dev in the web directory, then connect to the local server in a browser. When running on a remote server, port forwarding allows access from a local machine.
Costs and Constraints to Check Before Use
ViMax defaults to a method of creating two candidates for selection during image generation, which increases visual quality but may raise API call costs. To disable this feature, change the num_candidates value to 1 in the configuration file. Also, since both the web UI and TUI use the same agent runtime and configuration files, it is important to accurately set the API keys and endpoints for each model provider. Complex video generation processes involve multiple stages of AI model calls, so expected costs and processing times for large-scale generation should be considered in advance.