Skip to main content

Video Generation Nodes

Create videos from text prompts, images, or existing videos using AI models. All video generation uses the WaveSpeed API via the generate-image edge function with a 6-minute polling timeout.

Text To Video​

Generate videos from text descriptions

Creates videos directly from text prompts using AI video generation models. Optionally accepts image conditioning for models that support it.

PropertyDescription
Inputsinput (text), image (image, optional), images (image, optional)
Outputsoutput (video)

Configuration:

SettingDescription
ModelAI model for generation (see table below)
DurationVideo length in seconds (varies by model)
ResolutionOutput resolution (e.g., 720p, 768p, 1080p)
Aspect RatioOutput dimensions (e.g., 16:9, 9:16, 1:1)
Generate AudioEnable audio generation (model dependent)
Guidance ScalePrompt adherence strength (some models)
Negative PromptWhat to avoid (some models)

Text-to-Video Models​

ModelDurationAudioNotes
google/veo34-8sYesHighest quality
google/veo3-fast4-8sYesBalanced speed/quality
google/veo3.1/text-to-video4-8sYesLatest Veo generation
kwaivgi/kling-v3.0-pro/text-to-video3-15sYesKling 3.0 Pro, supports cfg_scale + negative prompt
kwaivgi/kling-v3.0-std/text-to-video3-15sYesKling 3.0 Standard
kwaivgi/kling-v2.6-pro/text-to-video5 or 10sYesKling 2.6 Pro with sound
kwaivgi/kling-video-o1/text-to-video5 or 10sNoKling O1
minimax/hailuo-02/prodefault 6sNoHailuo Pro, prompt expansion
minimax/hailuo-2.3/t2v-standarddefault 6sNoHailuo 2.3 Standard (default T2V model)
minimax/video-01--NoMinimax Video 01

Some T2V models also accept optional image conditioning:

ModelDescription
bytedance/waver-1.0Image + text conditioning
minimax/hailuo-02/fastFast mode with image input
google/veo3.1-fast/image-to-videoVeo 3.1 Fast I2V (in T2V payload builder)
google/veo3.1/reference-to-videoMulti-image reference (accepts images array)
Audio Support

Veo 3/3.1 and Kling v2.6+/v3.0 models generate synchronized audio when generateAudio is enabled (default: true).


Image To Video​

Animate still images into video

Transforms a static image into a moving video with AI-generated motion. Supports single-image and multi-image reference modes.

PropertyDescription
Inputsimage (image), input (text, optional), images (image, optional for multi-reference)
Outputsoutput (video)

Configuration:

SettingDescription
ModelAI model for generation (see table below)
DurationVideo length in seconds
ResolutionOutput resolution
Aspect RatioOutput dimensions
Guidance ScalePrompt adherence (some models)
Generate AudioEnable audio (model dependent)
Camera FixedMaintain static camera (Seedance)
Negative PromptWhat to avoid (Kling models)

Image-to-Video Models​

ModelDurationAudioNotes
google/veo3-fast/image-to-video4, 6, or 8sYesDefault I2V model
google/veo3/image-to-video4, 6, or 8sYesHigher quality Veo 3
kwaivgi/kling-v3.0-pro/image-to-video3-15sYesKling 3.0 Pro with cfg_scale
kwaivgi/kling-v3.0-std/image-to-video3-15sYesKling 3.0 Standard
kwaivgi/kling-v2.6-pro/image-to-video5 or 10sYesKling 2.6 Pro with sound
kwaivgi/kling-video-o1/image-to-video5 or 10sNoKling O1 (supports end frame)
kwaivgi/kling-v2.5-turbo-std/image-to-videoconfigurableNoKling 2.5 Turbo
minimax/hailuo-2.3/i2v-standarddefault 5sNoHailuo 2.3 Standard
minimax/hailuo-2.3/i2v-prodefault 5sNoHailuo 2.3 Pro
minimax/hailuo-2.3/i2v-fastdefault 5sNoHailuo 2.3 Fast
minimax/hailuo-02/fastdefault 6sNoHailuo Fast with prompt expansion
alibaba/wan-2.5/image-to-videodefault 5sNoWan 2.5
bytedance/seedance-v1-pro-i2v-480pdefault 5sNoSeedance with camera_fixed option
bytedance/waver-1.0--NoWaver 1.0
wavespeed-ai/wan-2.2-spicy/image-to-videodefault 5sNoWan 2.2 Spicy
wavespeed-ai/ltx-2.3/image-to-videodefault 5sNoLTX 2.3

Multi-Image Reference (Kling O1)​

The kwaivgi/kling-video-o1/reference-to-video model accepts multiple images via the images handle. This enables scene transitions and character-consistent video generation from multiple reference images.


Video Extend​

Lengthen existing videos

Extends a video beyond its original duration, generating new content that continues the scene.

PropertyDescription
Inputsvideo (video), prompt (text, optional), negative_prompt (text, optional)
Outputsvideo (video)

Configuration:

SettingDescription
ModelExtension model (e.g., Wan 2.5 Video Extend, Kling Video Extend)
DurationExtension length in seconds
ResolutionOutput resolution

Video Upscale​

Increase video resolution

Upscales video to higher resolution while maintaining quality.

PropertyDescription
Inputsvideo (video)
Outputsvideo (video)

Configuration:

SettingDescription
ScaleTarget scale (e.g., 2x, 4x)
DenoiseNoise reduction level (0-1)
Frame InterpolationSmooth motion (off, 2x, 4x)

Video Eraser​

Remove objects from video

Removes unwanted objects or areas from video while maintaining temporal consistency.

PropertyDescription
Inputsvideo (video)
Outputsvideo (video)

Use cases:

  • Remove watermarks
  • Remove unwanted objects
  • Clean up footage

Video Watermark Remover​

Specialized watermark removal

Detects and removes watermarks from video.

PropertyDescription
Inputsvideo (video)
Outputsvideo (video)

Video Face Swap​

Replace faces in video

Swaps faces in video with a source face image while maintaining expressions and movement.

PropertyDescription
Inputsvideo (video), faceImage (image), sourceImages (image, optional), targetImages (image, optional), input (text, optional)
Outputsvideo (video)
caution

Use responsibly. Face swapping should only be used with appropriate permissions.


Video Translate​

Translate spoken content in video

Translates audio/speech in video to another language with lip sync.

PropertyDescription
Inputsvideo (video)
Outputsvideo (video)

Supported Languages: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, and more.


Model Comparison​

Audio-Capable Models​

ModelTypeDurationNotes
Veo 3 / 3.1T2V + I2V4-8sHighest quality with audio
Veo 3 FastT2V + I2V4-8sFaster with audio
Kling v3.0 Pro/StdT2V + I2V3-15sFlexible duration, cfg_scale
Kling v2.6 ProT2V + I2V5 or 10sSound enabled by default

Non-Audio Models​

ModelTypeDurationNotes
Kling O1T2V + I2V + Reference5 or 10sMulti-image reference support
Hailuo 2.3T2V + I2V~5-6sStandard/Pro/Fast tiers
Wan 2.5T2V + I2V~5sBudget option
Waver 1.0T2V + I2V--Simple image+text conditioning
Seedance v1 ProI2V~5sCamera fixed option

Tips for Video Generation​

Prompts​

  1. Describe motion -- "walking", "flowing", "zooming"
  2. Specify camera -- "static shot", "slow pan", "tracking"
  3. Include timing -- "slow motion", "time-lapse"

Quality​

  1. Start with high-quality images for Image to Video
  2. Use appropriate aspect ratio for your target platform
  3. Choose model based on quality vs speed needs
  4. Enable audio for models that support it

Workflow Example​

Complete video creation pipeline:

Text Prompt --> Text to Image --> Image to Video --> Video Extend --> Library