Audio Nodes
Generate audio content including speech, music, and sound effects. All audio nodes output a video artifact (video with audio track) except Text to Audio which outputs an audio artifact.
Text to Audio
Generate audio from text prompts
Generates audio (speech, music, sound effects) from text using the Kling Text-to-Audio model.
| Property | Description |
|---|---|
| Inputs | text (text) |
| Outputs | audio (audio) |
Configuration:
| Setting | Description |
|---|---|
| Model | kwaivgi/kling-text-to-audio (default) |
| Duration | Audio duration in seconds (default: 10) |
| Prompt Override | Custom prompt (overrides flow context promptText) |
How it works:
- Validates prompt text from flow context
- Builds payload with duration and prompt
- Submits to
generate-imageedge function withmediaType: 'audio' - Polls for completion (2-minute timeout)
- Returns audio artifact
Use cases:
- Voiceovers for videos
- Sound effects generation
- Music generation from text descriptions
- Narration for content
Video Foley
Generate foley sound effects for video
Automatically generates synchronized foley sound effects for a video using the Hunyuan Video Foley model.
| Property | Description |
|---|---|
| Inputs | video (video), ctx (text, optional prompt) |
| Outputs | video (video -- with audio) |
Configuration:
| Setting | Description |
|---|---|
| Prompt Override | Custom prompt for sound guidance (overrides flow context) |
| Seed | Reproducibility (-1 for random) |
Model: wavespeed-ai/hunyuan-video-foley
How it works:
- Analyzes the video content and generates appropriate foley sounds synchronized with the visuals
- Prompt can guide the type of sounds to generate
- Returns a video artifact with the generated audio track
Use cases:
- Add environmental sounds to silent video
- Generate footstep, impact, and action sounds
- Create ambient audio matching video content
Kling Video To Audio
Generate SFX and background music for video
Generates synchronized sound effects and background music for videos using Kling's specialized video-to-audio model. Supports separate prompts for SFX and BGM.
| Property | Description |
|---|---|
| Inputs | video (video), sfx (text, optional SFX prompt), bgm (text, optional BGM prompt) |
| Outputs | video (video -- with audio) |
Configuration:
| Setting | Description |
|---|---|
| ASMR Mode | Enhanced micro-detail audio for immersive experience |
Model: kwaivgi/kling-video-to-audio
Payload parameters:
video-- Source video URLsound_effect_prompt-- Text prompt for sound effects (fromsfxinput)bgm_prompt-- Text prompt for background music (frombgminput)asmr_mode-- Boolean for enhanced audio detail
Example:
Video: Nature scene
SFX: "Gentle rustling leaves, distant bird calls"
BGM: "Soft ambient piano with nature atmosphere"
ASMR Mode: enabled
Use cases:
- Add atmospheric audio to AI-generated videos
- Create immersive soundscapes with ASMR mode
- Combine distinct SFX and BGM layers
MMAudio
Multi-modal audio generation with fine control
Advanced audio generation using MMAudio v2 that analyzes video content to create synchronized audio with negative prompt support and configurable parameters.
| Property | Description |
|---|---|
| Inputs | video (video), ctx (text, prompt), neg (text, optional negative prompt) |
| Outputs | video (video -- with audio) |
Configuration:
| Setting | Default | Description |
|---|---|---|
| Duration | 8s | Audio duration (1-30s) |
| Guidance Scale | 4.5 | Prompt adherence strength (1.0-10.0) |
| Inference Steps | 25 | Quality vs speed tradeoff (10-50) |
| Mask Away Clip | false | Mask away video clip features |
Model: wavespeed-ai/mmaudio-v2
Negative prompt support:
Connect a Text Prompt node to the neg input to specify what to avoid:
Prompt: "Gentle acoustic guitar with soft ambient sounds"
Negative: "No drums, no electronic sounds, no vocals"
How it works:
- Extracts prompt from flow context and negative prompt from
neginput - Builds payload with duration, guidance scale, inference steps, and mask settings
- Submits to
generate-imageedge function - Polls for completion (2-minute timeout)
- Returns video artifact with generated audio
Parameter guidance:
- Guidance scale -- Higher values (4.5-7.0) follow prompts more closely
- Inference steps -- More steps (25-50) produce better quality but run slower
- Mask away clip -- Enable to ignore video features for pure audio generation
Audio Workflow Examples
Voiceover Pipeline
Add narration to video:
Text Prompt --> Text to Audio --> Lip Sync
^
Upload Video
Automatic Video Audio
Let AI generate appropriate audio:
Upload Video --> MMAudio --> Library
^
Text Prompt: "Natural ambient sounds
with subtle background music"
Kling Video with Layered Audio
Add both SFX and BGM to Kling-generated video:
Text Prompt (SFX) --v
Kling Video --> Kling Video To Audio --> Library
Text Prompt (BGM) --^
Foley Sound Pipeline
Upload Video --> Video Foley --> Library
Tips for Audio Quality
Text to Audio
- Be descriptive -- "Warm male narrator with calm energy" works better than "male voice"
- Specify duration -- Set duration to match your video length
Video Foley
- Clear video content -- Videos with distinct actions produce better foley
- Guide with prompt -- Use the prompt to specify the type of sounds desired
Kling Video To Audio
- ASMR mode -- Enable for immersive, detailed audio experiences
- Separate prompts -- Use distinct SFX and BGM text inputs for independent control
- Best with Kling videos -- Optimized for Kling-generated video characteristics
MMAudio
- Guidance scale -- Start at 4.5, increase for stronger prompt following
- Inference steps -- Use 25 for speed, 40+ for quality
- Negative prompts -- Effective for excluding unwanted audio elements
- Mask away clip -- Enable to generate audio independent of video features