Skip to main content

AI Models Overview

DUTO integrates multiple AI models for different content generation tasks, routing requests to the best provider automatically.

Model Categories​

DUTO uses AI models across four categories:

CategoryPurposeExamples
ImageGenerate and edit imagesNano Banana Pro, SeedDream v4.5/v5 Lite, Flux Kontext
VideoGenerate and edit videosVeo 3, Kling O1/v2.6 Pro, Hailuo 2.3, Wan 2.5, Sora, Waver, Seedance
AudioGenerate speech, music, soundTTS, MMAudio
LLMReasoning and textGemini 2.5 Flash/Pro, DeepSeek

Model Selection Philosophy​

DUTO provides multiple models because:

  1. Quality vs Speed - Some tasks need quality, others need speed
  2. Cost Efficiency - Budget models for testing, premium for finals
  3. Specialization - Different models excel at different tasks
  4. Audio Support - Some video models include native audio generation

AI Providers​

Primary Providers​

ProviderServices
WaveSpeedPrimary image and video generation (hosts most models)
OpenRouterLLM access (Gemini, DeepSeek)
TavilyWeb research and image search

Model Routing​

DUTO automatically routes generation requests to the correct provider based on the model path prefix:

Model PrefixProviderEdge Function
google/*WaveSpeedgenerate-image / poll-wavespeed-job
bytedance/*WaveSpeedgenerate-image / poll-wavespeed-job
kwaivgi/*WaveSpeedgenerate-image / poll-wavespeed-job
minimax/*WaveSpeedgenerate-image / poll-wavespeed-job
alibaba/*WaveSpeedgenerate-image / poll-wavespeed-job
openai/*WaveSpeedgenerate-image / poll-wavespeed-job
wavespeed-ai/*WaveSpeedgenerate-image / poll-wavespeed-job

All image and video generation requests go through WaveSpeed as the primary provider. Requests are asynchronous: a job is submitted via generate-image, then polled for completion via poll-wavespeed-job.

Choosing Models​

Image Generation​

NeedRecommended Model
Default / fastNano Banana Pro
Higher qualityNano Banana 2
Budget testingSeedDream v4.5
Editing / inpaintingSeedDream v4.5 Edit, SeedDream v5 Lite Edit
LoRA supportFlux Kontext

Video Generation​

NeedRecommended Model
Best qualityVeo 3
Quality + speedVeo 3 Fast
Cinematic / narrativeSora
Good quality with audioKling v2.6 Pro
Flexible duration (3-10s)Wan 2.5
Fast iterationsHailuo 2.3 Fast
Budget with audioKling v2.6 Pro (also used as budget default)

Intelligence​

NeedRecommended Model
Complex reasoningGemini 2.5 Pro
Fast responsesGemini 2.5 Flash
Technical analysisDeepSeek

Audio Support​

Some video models generate synchronized audio natively:

ModelNative Audio
Veo 3Yes
Veo 3 FastYes
Veo 3.1Yes
Kling v2.6 ProYes
SoraNo
Waver 1.0No
Hailuo 2.3No
Wan 2.5No
Kling O1No
SeedanceNo

Budget Mode Models​

Budget Mode uses cost-effective alternatives when the model is set to "auto":

CategoryStandard (auto)Budget (auto)
ImageNano Banana ProSeedDream v4.5
VideoVeo 3Kling v2.6 Pro
LLMGemini 2.5 ProGemini 2.5 Flash

Duration Constraints​

Each video model supports specific durations:

ModelSupported Durations (seconds)
Veo 3 / Veo 3 Fast4, 6, 8
Waver 1.05-10
Hailuo 2.3 (all tiers)6, 10
Wan 2.5 / Wan 2.5 Fast3-10
Kling O15, 10
Kling v2.6 Pro5, 10
Seedance5-10

DUTO automatically clamps requested durations to the nearest valid value for the selected model.

Model Updates​

AI models evolve regularly:

  • New models are added as they become available
  • Existing models may improve over time
  • Deprecated models are phased out gracefully
  • Users are notified of significant changes

Learn More​

Detailed information about each category: