Skip to main content

Audio Nodes

Generate audio content including speech, music, and sound effects. All audio nodes output a video artifact (video with audio track) except Text to Audio which outputs an audio artifact.

Text to Audio​

Generate audio from text prompts

Generates audio (speech, music, sound effects) from text using the Kling Text-to-Audio model.

PropertyDescription
Inputstext (text)
Outputsaudio (audio)

Configuration:

SettingDescription
Modelkwaivgi/kling-text-to-audio (default)
DurationAudio duration in seconds (default: 10)
Prompt OverrideCustom prompt (overrides flow context promptText)

How it works:

  1. Validates prompt text from flow context
  2. Builds payload with duration and prompt
  3. Submits to generate-image edge function with mediaType: 'audio'
  4. Polls for completion (2-minute timeout)
  5. Returns audio artifact

Use cases:

  • Voiceovers for videos
  • Sound effects generation
  • Music generation from text descriptions
  • Narration for content

Video Foley​

Generate foley sound effects for video

Automatically generates synchronized foley sound effects for a video using the Hunyuan Video Foley model.

PropertyDescription
Inputsvideo (video), ctx (text, optional prompt)
Outputsvideo (video -- with audio)

Configuration:

SettingDescription
Prompt OverrideCustom prompt for sound guidance (overrides flow context)
SeedReproducibility (-1 for random)

Model: wavespeed-ai/hunyuan-video-foley

How it works:

  • Analyzes the video content and generates appropriate foley sounds synchronized with the visuals
  • Prompt can guide the type of sounds to generate
  • Returns a video artifact with the generated audio track

Use cases:

  • Add environmental sounds to silent video
  • Generate footstep, impact, and action sounds
  • Create ambient audio matching video content

Kling Video To Audio​

Generate SFX and background music for video

Generates synchronized sound effects and background music for videos using Kling's specialized video-to-audio model. Supports separate prompts for SFX and BGM.

PropertyDescription
Inputsvideo (video), sfx (text, optional SFX prompt), bgm (text, optional BGM prompt)
Outputsvideo (video -- with audio)

Configuration:

SettingDescription
ASMR ModeEnhanced micro-detail audio for immersive experience

Model: kwaivgi/kling-video-to-audio

Payload parameters:

  • video -- Source video URL
  • sound_effect_prompt -- Text prompt for sound effects (from sfx input)
  • bgm_prompt -- Text prompt for background music (from bgm input)
  • asmr_mode -- Boolean for enhanced audio detail

Example:

Video: Nature scene
SFX: "Gentle rustling leaves, distant bird calls"
BGM: "Soft ambient piano with nature atmosphere"
ASMR Mode: enabled

Use cases:

  • Add atmospheric audio to AI-generated videos
  • Create immersive soundscapes with ASMR mode
  • Combine distinct SFX and BGM layers

MMAudio​

Multi-modal audio generation with fine control

Advanced audio generation using MMAudio v2 that analyzes video content to create synchronized audio with negative prompt support and configurable parameters.

PropertyDescription
Inputsvideo (video), ctx (text, prompt), neg (text, optional negative prompt)
Outputsvideo (video -- with audio)

Configuration:

SettingDefaultDescription
Duration8sAudio duration (1-30s)
Guidance Scale4.5Prompt adherence strength (1.0-10.0)
Inference Steps25Quality vs speed tradeoff (10-50)
Mask Away ClipfalseMask away video clip features

Model: wavespeed-ai/mmaudio-v2

Negative prompt support:

Connect a Text Prompt node to the neg input to specify what to avoid:

Prompt: "Gentle acoustic guitar with soft ambient sounds"
Negative: "No drums, no electronic sounds, no vocals"

How it works:

  1. Extracts prompt from flow context and negative prompt from neg input
  2. Builds payload with duration, guidance scale, inference steps, and mask settings
  3. Submits to generate-image edge function
  4. Polls for completion (2-minute timeout)
  5. Returns video artifact with generated audio

Parameter guidance:

  • Guidance scale -- Higher values (4.5-7.0) follow prompts more closely
  • Inference steps -- More steps (25-50) produce better quality but run slower
  • Mask away clip -- Enable to ignore video features for pure audio generation

Audio Workflow Examples​

Voiceover Pipeline​

Add narration to video:

Text Prompt --> Text to Audio --> Lip Sync
^
Upload Video

Automatic Video Audio​

Let AI generate appropriate audio:

Upload Video --> MMAudio --> Library
^
Text Prompt: "Natural ambient sounds
with subtle background music"

Kling Video with Layered Audio​

Add both SFX and BGM to Kling-generated video:

                                  Text Prompt (SFX) --v
Kling Video --> Kling Video To Audio --> Library
Text Prompt (BGM) --^

Foley Sound Pipeline​

Upload Video --> Video Foley --> Library

Tips for Audio Quality​

Text to Audio​

  1. Be descriptive -- "Warm male narrator with calm energy" works better than "male voice"
  2. Specify duration -- Set duration to match your video length

Video Foley​

  1. Clear video content -- Videos with distinct actions produce better foley
  2. Guide with prompt -- Use the prompt to specify the type of sounds desired

Kling Video To Audio​

  1. ASMR mode -- Enable for immersive, detailed audio experiences
  2. Separate prompts -- Use distinct SFX and BGM text inputs for independent control
  3. Best with Kling videos -- Optimized for Kling-generated video characteristics

MMAudio​

  1. Guidance scale -- Start at 4.5, increase for stronger prompt following
  2. Inference steps -- Use 25 for speed, 40+ for quality
  3. Negative prompts -- Effective for excluding unwanted audio elements
  4. Mask away clip -- Enable to generate audio independent of video features