Skip to main content

LLM Models

DUTO uses large language models (LLMs) via OpenRouter for intelligence features like the Brain node, prompt expansion, and content analysis. The default model for all LLM operations is Gemini 2.5 Flash.

Available Models​

Gemini 2.5 Flash​

Fast, efficient reasoning

google/gemini-2.5-flash
AttributeValue
SpeedVery Fast
QualityGood
ContextLarge
CostLow
ProviderOpenRouter

Strengths:

  • Very fast responses
  • Cost efficient
  • Large context window
  • Vision capable
  • Good for simple tasks

Best for:

  • Quick analysis
  • Simple transformations
  • High-volume processing
  • Speed-critical tasks
  • Image analysis with text

Gemini 2.5 Pro​

High-quality reasoning

google/gemini-2.5-pro
AttributeValue
SpeedMedium
QualityExcellent
ContextVery Large
CostMedium
ProviderOpenRouter

Strengths:

  • Excellent reasoning
  • Complex task handling
  • Nuanced understanding
  • Creative capabilities
  • Vision capable

Best for:

  • Complex analysis
  • Creative writing
  • Detailed reasoning
  • Important decisions
  • Image-heavy analysis

DeepSeek​

Technical and analytical model

deepseek/deepseek-r1
AttributeValue
SpeedMedium
QualityHigh
ContextLarge
CostLow
ProviderOpenRouter

Strengths:

  • Strong technical reasoning
  • Good at structured tasks
  • Cost effective
  • Logical analysis
  • Good for code-like tasks

Best for:

  • Technical analysis
  • Data processing
  • Structured outputs
  • Logic-heavy tasks

Model Comparison​

ModelSpeedQualityCostBest For
Gemini Flash★★★★★★★★LowQuick tasks
Gemini Pro★★★★★★★★MediumQuality tasks
DeepSeek★★★★★★★★LowTechnical tasks

Usage in DUTO​

Brain Node Modes​

The Brain node uses LLMs differently based on mode:

ModeDefault ModelUse Case
CreativeGemini 2.5 FlashCreative writing, ideas
ResearchGemini 2.5 FlashAnalysis, synthesis (with tool calling)
HybridGemini 2.5 FlashBalanced tasks (with tool calling)
AnalysisGemini 2.5 FlashTechnical analysis

Research and Hybrid modes enable tool calling (web search, image search via Tavily), while Creative and Analysis modes use direct generation without tools.

Prompt Expander​

The Prompt Expander node uses LLMs to enhance prompts:

ModeDefault ModelPurpose
CinematicGemini FlashFilm-like descriptions
CommercialGemini FlashAdvertising style
AnimeGemini FlashAnime style
DocumentaryGemini FlashRealistic style

Content Analysis​

LLMs are used for:

  • Text classification
  • Sentiment analysis
  • Content summarization
  • Image description
  • Reference analysis

OpenRouter Integration​

DUTO uses OpenRouter as its LLM gateway, providing access to multiple model providers through a single API:

Edge Functions:

  • brain-reasoning - Brain node agentic reasoning (uses Vercel AI SDK)
  • openrouter-chat - Prompt expansion, analyzer, and general LLM tasks

Default Model: google/gemini-2.5-flash (used for both fast and creative tasks)

Available Models:

  • google/gemini-2.5-flash - Default, fast, cost-efficient
  • google/gemini-2.5-pro - Higher quality reasoning

Credit Costs​

ModelSimple QueryComplex Task
Gemini Flash1 credit2 credits
Gemini Pro2-3 credits4-5 credits
DeepSeek1-2 credits2-3 credits

Task Recommendations​

Simple Tasks​

Use Gemini Flash for:

  • Extracting data
  • Simple transformations
  • Quick categorization
  • Format conversion
  • Basic prompt expansion

Complex Tasks​

Use Gemini Pro for:

  • Creative writing
  • Complex analysis
  • Nuanced decisions
  • Multi-step reasoning
  • Image-heavy analysis
  • Story generation

Technical Tasks​

Use DeepSeek for:

  • Data parsing
  • Structured extraction
  • Technical analysis
  • Logic-heavy tasks
  • Code generation

Prompt Optimization​

For All Models​

  1. Be specific

    ✓ "Extract the product name and price from this description"
    ✗ "Get info from this"
  2. Provide context

    ✓ "This is a product description. Extract: name, price, category"
    ✗ "Parse this text"
  3. Specify format

    ✓ "Return as JSON: {name: string, price: number}"
    ✗ "Give me the data"

Gemini Flash Optimization​

Keep prompts concise:

Extract product name and price as JSON.
Input: [text]
Output: {"name": "...", "price": ...}

Gemini Pro Optimization​

Can handle complex prompts:

Analyze this product description and provide:
1. Product name
2. Key features (list)
3. Target audience
4. Suggested improvements

Consider market positioning and competitor analysis.

DeepSeek Optimization​

Structure requests clearly:

Task: Parse the following data
Format: JSON array
Fields: name, email, company
Rules:
- Normalize email to lowercase
- Extract company from domain if not stated

Advanced Usage​

Chain of Thought​

For complex reasoning:

Think through this step by step:
1. First, identify...
2. Then, analyze...
3. Finally, conclude...

Output Schemas​

Define exact output structure:

{
"analysis": {
"sentiment": "positive|negative|neutral",
"confidence": 0.0-1.0,
"key_points": ["string"]
}
}

Few-Shot Examples​

Provide examples for consistency:

Example input: "Blue cotton t-shirt, size M, $29.99"
Example output: {"product": "t-shirt", "color": "blue", "price": 29.99}

Now process: "Red wool sweater, size L, $49.99"

Troubleshooting​

Inconsistent Outputs​

Problem: Same input gives different outputs

Solutions:

  1. Be more specific in prompt
  2. Use output schema
  3. Provide examples
  4. Use Pro model for consistency

Wrong Format​

Problem: Output doesn't match expected format

Solutions:

  1. Specify format explicitly
  2. Provide JSON schema
  3. Include example output
  4. Parse and validate in workflow

Slow Responses​

Problem: LLM taking too long

Solutions:

  1. Use Flash model
  2. Simplify prompt
  3. Reduce output requirements
  4. Break into smaller tasks

Vision Capabilities​

Both Gemini models support vision input for:

  • Image analysis
  • Reference description
  • Visual Q&A
  • Content moderation
  • Storyboard context

When using image inputs with LLMs:

  • Provide clear context for the image
  • Specify what to extract/analyze
  • Use appropriate model (Pro for complex analysis)
  • Consider image size/quality