Video Generation Nodes
Create videos from text prompts, images, or existing videos using AI models. All video generation uses the WaveSpeed API via the generate-image edge function with a 6-minute polling timeout.
Text To Video
Generate videos from text descriptions
Creates videos directly from text prompts using AI video generation models. Optionally accepts image conditioning for models that support it.
| Property | Description |
|---|---|
| Inputs | input (text), image (image, optional), images (image, optional) |
| Outputs | output (video) |
Configuration:
| Setting | Description |
|---|---|
| Model | AI model for generation (see table below) |
| Duration | Video length in seconds (varies by model) |
| Resolution | Output resolution (e.g., 720p, 768p, 1080p) |
| Aspect Ratio | Output dimensions (e.g., 16:9, 9:16, 1:1) |
| Generate Audio | Enable audio generation (model dependent) |
| Guidance Scale | Prompt adherence strength (some models) |
| Negative Prompt | What to avoid (some models) |
Text-to-Video Models
| Model | Duration | Audio | Notes |
|---|---|---|---|
google/veo3 | 4-8s | Yes | Highest quality |
google/veo3-fast | 4-8s | Yes | Balanced speed/quality |
google/veo3.1/text-to-video | 4-8s | Yes | Latest Veo generation |
kwaivgi/kling-v3.0-pro/text-to-video | 3-15s | Yes | Kling 3.0 Pro, supports cfg_scale + negative prompt |
kwaivgi/kling-v3.0-std/text-to-video | 3-15s | Yes | Kling 3.0 Standard |
kwaivgi/kling-v2.6-pro/text-to-video | 5 or 10s | Yes | Kling 2.6 Pro with sound |
kwaivgi/kling-video-o1/text-to-video | 5 or 10s | No | Kling O1 |
minimax/hailuo-02/pro | default 6s | No | Hailuo Pro, prompt expansion |
minimax/hailuo-2.3/t2v-standard | default 6s | No | Hailuo 2.3 Standard (default T2V model) |
minimax/video-01 | -- | No | Minimax Video 01 |
Some T2V models also accept optional image conditioning:
| Model | Description |
|---|---|
bytedance/waver-1.0 | Image + text conditioning |
minimax/hailuo-02/fast | Fast mode with image input |
google/veo3.1-fast/image-to-video | Veo 3.1 Fast I2V (in T2V payload builder) |
google/veo3.1/reference-to-video | Multi-image reference (accepts images array) |
Veo 3/3.1 and Kling v2.6+/v3.0 models generate synchronized audio when generateAudio is enabled (default: true).
Image To Video
Animate still images into video
Transforms a static image into a moving video with AI-generated motion. Supports single-image and multi-image reference modes.
| Property | Description |
|---|---|
| Inputs | image (image), input (text, optional), images (image, optional for multi-reference) |
| Outputs | output (video) |
Configuration:
| Setting | Description |
|---|---|
| Model | AI model for generation (see table below) |
| Duration | Video length in seconds |
| Resolution | Output resolution |
| Aspect Ratio | Output dimensions |
| Guidance Scale | Prompt adherence (some models) |
| Generate Audio | Enable audio (model dependent) |
| Camera Fixed | Maintain static camera (Seedance) |
| Negative Prompt | What to avoid (Kling models) |
Image-to-Video Models
| Model | Duration | Audio | Notes |
|---|---|---|---|
google/veo3-fast/image-to-video | 4, 6, or 8s | Yes | Default I2V model |
google/veo3/image-to-video | 4, 6, or 8s | Yes | Higher quality Veo 3 |
kwaivgi/kling-v3.0-pro/image-to-video | 3-15s | Yes | Kling 3.0 Pro with cfg_scale |
kwaivgi/kling-v3.0-std/image-to-video | 3-15s | Yes | Kling 3.0 Standard |
kwaivgi/kling-v2.6-pro/image-to-video | 5 or 10s | Yes | Kling 2.6 Pro with sound |
kwaivgi/kling-video-o1/image-to-video | 5 or 10s | No | Kling O1 (supports end frame) |
kwaivgi/kling-v2.5-turbo-std/image-to-video | configurable | No | Kling 2.5 Turbo |
minimax/hailuo-2.3/i2v-standard | default 5s | No | Hailuo 2.3 Standard |
minimax/hailuo-2.3/i2v-pro | default 5s | No | Hailuo 2.3 Pro |
minimax/hailuo-2.3/i2v-fast | default 5s | No | Hailuo 2.3 Fast |
minimax/hailuo-02/fast | default 6s | No | Hailuo Fast with prompt expansion |
alibaba/wan-2.5/image-to-video | default 5s | No | Wan 2.5 |
bytedance/seedance-v1-pro-i2v-480p | default 5s | No | Seedance with camera_fixed option |
bytedance/waver-1.0 | -- | No | Waver 1.0 |
wavespeed-ai/wan-2.2-spicy/image-to-video | default 5s | No | Wan 2.2 Spicy |
wavespeed-ai/ltx-2.3/image-to-video | default 5s | No | LTX 2.3 |
Multi-Image Reference (Kling O1)
The kwaivgi/kling-video-o1/reference-to-video model accepts multiple images via the images handle. This enables scene transitions and character-consistent video generation from multiple reference images.
Video Extend
Lengthen existing videos
Extends a video beyond its original duration, generating new content that continues the scene.
| Property | Description |
|---|---|
| Inputs | video (video), prompt (text, optional), negative_prompt (text, optional) |
| Outputs | video (video) |
Configuration:
| Setting | Description |
|---|---|
| Model | Extension model (e.g., Wan 2.5 Video Extend, Kling Video Extend) |
| Duration | Extension length in seconds |
| Resolution | Output resolution |
Video Upscale
Increase video resolution
Upscales video to higher resolution while maintaining quality.
| Property | Description |
|---|---|
| Inputs | video (video) |
| Outputs | video (video) |
Configuration:
| Setting | Description |
|---|---|
| Scale | Target scale (e.g., 2x, 4x) |
| Denoise | Noise reduction level (0-1) |
| Frame Interpolation | Smooth motion (off, 2x, 4x) |
Video Eraser
Remove objects from video
Removes unwanted objects or areas from video while maintaining temporal consistency.
| Property | Description |
|---|---|
| Inputs | video (video) |
| Outputs | video (video) |
Use cases:
- Remove watermarks
- Remove unwanted objects
- Clean up footage
Video Watermark Remover
Specialized watermark removal
Detects and removes watermarks from video.
| Property | Description |
|---|---|
| Inputs | video (video) |
| Outputs | video (video) |
Video Face Swap
Replace faces in video
Swaps faces in video with a source face image while maintaining expressions and movement.
| Property | Description |
|---|---|
| Inputs | video (video), faceImage (image), sourceImages (image, optional), targetImages (image, optional), input (text, optional) |
| Outputs | video (video) |
Use responsibly. Face swapping should only be used with appropriate permissions.
Video Translate
Translate spoken content in video
Translates audio/speech in video to another language with lip sync.
| Property | Description |
|---|---|
| Inputs | video (video) |
| Outputs | video (video) |
Supported Languages: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, and more.
Model Comparison
Audio-Capable Models
| Model | Type | Duration | Notes |
|---|---|---|---|
| Veo 3 / 3.1 | T2V + I2V | 4-8s | Highest quality with audio |
| Veo 3 Fast | T2V + I2V | 4-8s | Faster with audio |
| Kling v3.0 Pro/Std | T2V + I2V | 3-15s | Flexible duration, cfg_scale |
| Kling v2.6 Pro | T2V + I2V | 5 or 10s | Sound enabled by default |
Non-Audio Models
| Model | Type | Duration | Notes |
|---|---|---|---|
| Kling O1 | T2V + I2V + Reference | 5 or 10s | Multi-image reference support |
| Hailuo 2.3 | T2V + I2V | ~5-6s | Standard/Pro/Fast tiers |
| Wan 2.5 | T2V + I2V | ~5s | Budget option |
| Waver 1.0 | T2V + I2V | -- | Simple image+text conditioning |
| Seedance v1 Pro | I2V | ~5s | Camera fixed option |
Tips for Video Generation
Prompts
- Describe motion -- "walking", "flowing", "zooming"
- Specify camera -- "static shot", "slow pan", "tracking"
- Include timing -- "slow motion", "time-lapse"
Quality
- Start with high-quality images for Image to Video
- Use appropriate aspect ratio for your target platform
- Choose model based on quality vs speed needs
- Enable audio for models that support it
Workflow Example
Complete video creation pipeline:
Text Prompt --> Text to Image --> Image to Video --> Video Extend --> Library