Audio Models
DUTO provides multiple AI models for generating speech, sound effects, and synchronized audio for videos.
Text to Audio
Kling Text-to-Audio
Audio generation from text prompts
kwaivgi/kling-text-to-audio| Attribute | Value |
|---|---|
| Quality | Good |
| Speed | Medium |
| Duration | 5-30s (configurable) |
| Provider | Kwaivgi (via WaveSpeed) |
Features:
- Text-prompt-based audio generation
- Configurable duration (5s to 30s, in 5s steps)
- Music, sound effects, and ambient audio
Usage:
- Connect a Text Prompt node with a description of the desired audio
- Set duration using the slider on the node
- Audio is generated asynchronously and polled for completion
Video Foley
Hunyuan Video Foley
Automatic sound effects for video
wavespeed-ai/hunyuan-video-foley| Attribute | Value |
|---|---|
| Quality | Good |
| Speed | Medium |
| Provider | WaveSpeed AI |
Inputs:
- Video (required) - Source video to generate foley for
- Description (optional) - Text prompt to guide sound generation
Capabilities:
- Analyzes video content automatically
- Generates synchronized sound effects
- Supports prompt-guided audio generation
- Outputs video with added foley audio
Best for:
- Adding sound effects to silent videos
- Environmental ambience
- Action-synced sounds
MMAudio
Multi-Modal Audio Generation
mmaudio| Attribute | Value |
|---|---|
| Quality | High |
| Speed | Medium |
| Provider | MMAudio (via WaveSpeed) |
Inputs:
- Video (required) - Source video
- Description (optional) - Text prompt for audio guidance
Capabilities:
- Scene-appropriate audio generation
- Music and sound effects
- Ambient soundscapes
- Motion-to-sound mapping
Best for:
- Full audio design for video
- Scene-matched background audio
- Synchronized timing
Kling Video to Audio
Kling-Optimized Audio
kling/video-to-audio| Attribute | Value |
|---|---|
| Quality | High |
| Speed | Medium |
| Provider | Kwaivgi (via WaveSpeed) |
Inputs:
- Video (required) - Source video
Features:
- Optimized for video content analysis
- Style-appropriate audio generation
- Good motion synchronization
Lip Sync
Audio-Visual Synchronization
lip-syncAvailable Models:
| Model | Provider | Key Feature |
|---|---|---|
veed/lipsync | Veed | Default, reliable |
kwaivgi/kling-lipsync/audio-to-video | Kling | Good quality sync |
sync/lipsync-2 | Sync Labs | Sync mode options |
sync/lipsync-2-pro | Sync Labs | Pro quality |
wavespeed-ai/infinitetalk/video-to-video | InfiniteTalk | Prompt support, resolution options |
wavespeed-ai/infinitetalk-fast/video-to-video | InfiniteTalk | Fast generation |
Inputs:
- Video (required) - Source video with visible face
- Audio (required) - Audio to sync
- Prompt (optional, InfiniteTalk only) - Text guidance
Model-specific settings:
- Sync Labs: Sync mode (loop, bounce, cut off, silence, remap)
- InfiniteTalk: Resolution (480p, 720p), Seed for reproducibility
Use cases:
- Dubbing videos in different languages
- Voice replacement
- Voiceover synchronization
Audio Model Comparison
| Model | Purpose | Inputs | Quality | Speed |
|---|---|---|---|---|
| Kling Text-to-Audio | Generate from text | Text prompt | Good | Medium |
| Hunyuan Video Foley | Sound effects | Video + optional text | Good | Medium |
| MMAudio | Full audio design | Video + optional text | High | Medium |
| Kling Video-to-Audio | Video audio | Video | High | Medium |
| Lip Sync (6 models) | Sync audio to face | Video + audio | Good | Varies |
Workflow Patterns
Adding Sound Effects to Video
Video → Video Foley → Output (video with foley)
Adding Voice to Video
Text Prompt → Text to Audio → (audio)
Video → Lip Sync ← (audio) → Output
Complete Audio Design
Video → MMAudio → (audio + video)
Video → Video Foley → (video with effects)
Text-based Audio Creation
Text Prompt ("ocean waves, seagulls") → Text to Audio → Output (audio file)
Audio Quality Tips
Text to Audio
- Be descriptive about the desired audio style and content
- Set appropriate duration for your needs
- Use clear descriptions of instruments, sounds, or environments
Video Foley
- Higher quality source video produces better sound matching
- Use the optional description input to guide sound generation
- Clear visual actions get better foley effects
Lip Sync
- Ensure clear face visibility in the video
- Use high-quality audio input
- Audio and video durations should be reasonably matched
- Try InfiniteTalk for prompt-guided results
- Use Sync Labs models for sync mode control (loop, bounce, etc.)