AI Models Overview
DUTO integrates multiple AI models for different content generation tasks, routing requests to the best provider automatically.
Model Categories
DUTO uses AI models across four categories:
| Category | Purpose | Examples |
|---|---|---|
| Image | Generate and edit images | Nano Banana Pro, SeedDream v4.5/v5 Lite, Flux Kontext |
| Video | Generate and edit videos | Veo 3, Kling O1/v2.6 Pro, Hailuo 2.3, Wan 2.5, Sora, Waver, Seedance |
| Audio | Generate speech, music, sound | TTS, MMAudio |
| LLM | Reasoning and text | Gemini 2.5 Flash/Pro, DeepSeek |
Model Selection Philosophy
DUTO provides multiple models because:
- Quality vs Speed - Some tasks need quality, others need speed
- Cost Efficiency - Budget models for testing, premium for finals
- Specialization - Different models excel at different tasks
- Audio Support - Some video models include native audio generation
AI Providers
Primary Providers
| Provider | Services |
|---|---|
| WaveSpeed | Primary image and video generation (hosts most models) |
| OpenRouter | LLM access (Gemini, DeepSeek) |
| Tavily | Web research and image search |
Model Routing
DUTO automatically routes generation requests to the correct provider based on the model path prefix:
| Model Prefix | Provider | Edge Function |
|---|---|---|
google/* | WaveSpeed | generate-image / poll-wavespeed-job |
bytedance/* | WaveSpeed | generate-image / poll-wavespeed-job |
kwaivgi/* | WaveSpeed | generate-image / poll-wavespeed-job |
minimax/* | WaveSpeed | generate-image / poll-wavespeed-job |
alibaba/* | WaveSpeed | generate-image / poll-wavespeed-job |
openai/* | WaveSpeed | generate-image / poll-wavespeed-job |
wavespeed-ai/* | WaveSpeed | generate-image / poll-wavespeed-job |
All image and video generation requests go through WaveSpeed as the primary provider. Requests are asynchronous: a job is submitted via generate-image, then polled for completion via poll-wavespeed-job.
Choosing Models
Image Generation
| Need | Recommended Model |
|---|---|
| Default / fast | Nano Banana Pro |
| Higher quality | Nano Banana 2 |
| Budget testing | SeedDream v4.5 |
| Editing / inpainting | SeedDream v4.5 Edit, SeedDream v5 Lite Edit |
| LoRA support | Flux Kontext |
Video Generation
| Need | Recommended Model |
|---|---|
| Best quality | Veo 3 |
| Quality + speed | Veo 3 Fast |
| Cinematic / narrative | Sora |
| Good quality with audio | Kling v2.6 Pro |
| Flexible duration (3-10s) | Wan 2.5 |
| Fast iterations | Hailuo 2.3 Fast |
| Budget with audio | Kling v2.6 Pro (also used as budget default) |
Intelligence
| Need | Recommended Model |
|---|---|
| Complex reasoning | Gemini 2.5 Pro |
| Fast responses | Gemini 2.5 Flash |
| Technical analysis | DeepSeek |
Audio Support
Some video models generate synchronized audio natively:
| Model | Native Audio |
|---|---|
| Veo 3 | Yes |
| Veo 3 Fast | Yes |
| Veo 3.1 | Yes |
| Kling v2.6 Pro | Yes |
| Sora | No |
| Waver 1.0 | No |
| Hailuo 2.3 | No |
| Wan 2.5 | No |
| Kling O1 | No |
| Seedance | No |
Budget Mode Models
Budget Mode uses cost-effective alternatives when the model is set to "auto":
| Category | Standard (auto) | Budget (auto) |
|---|---|---|
| Image | Nano Banana Pro | SeedDream v4.5 |
| Video | Veo 3 | Kling v2.6 Pro |
| LLM | Gemini 2.5 Pro | Gemini 2.5 Flash |
Duration Constraints
Each video model supports specific durations:
| Model | Supported Durations (seconds) |
|---|---|
| Veo 3 / Veo 3 Fast | 4, 6, 8 |
| Waver 1.0 | 5-10 |
| Hailuo 2.3 (all tiers) | 6, 10 |
| Wan 2.5 / Wan 2.5 Fast | 3-10 |
| Kling O1 | 5, 10 |
| Kling v2.6 Pro | 5, 10 |
| Seedance | 5-10 |
DUTO automatically clamps requested durations to the nearest valid value for the selected model.
Model Updates
AI models evolve regularly:
- New models are added as they become available
- Existing models may improve over time
- Deprecated models are phased out gracefully
- Users are notified of significant changes
Learn More
Detailed information about each category:
- Image Models - All image generation models
- Video Models - All video generation models
- Audio Models - All audio generation models
- LLM Models - All language models