Aion 3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. Multiple specialized models collaborate on each response to produce stronger narrative structure and more compelling tension and conflict. It handles mature and darker themes with nuance and supports tool calling for richer interactive fiction.
Venice
inference provider · 210 models
Access 210 models served through Venice on AnonRouter's privacy-first gateway, including Aion 3.0, Aion 3.0 Mini, and Claude Fable 5. Venice serves privacy-first inference with no payload logging, and AnonRouter strips identity before requests ever reach it.
Models
210
Modalities
4
Text, Image, Audio, Embeddings
From (input)
$0.0125
per 1M tokens
Max context
2M
Private routes
98 / 210
not anonymous-only
Catalog by modality
210 routesVenice models210
Aion 3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. Multiple specialized models collaborate on each response to produce stronger narrative structure and more compelling tension and conflict at lower cost. It handles mature and darker themes with nuance and supports tool calling for richer interactive fiction.
Claude Fable 5 is Anthropic's most capable widely released model, designed for demanding reasoning and long-horizon agentic work. It features a 1M token context window, 128K max output tokens, always-on adaptive thinking, and strong multimodal capabilities.
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work such as long code refactors, front-end development, and finance and analysis tasks. It features a 1M token context window, 128K max output tokens, always-on adaptive thinking, and strong multimodal capabilities, and tends to be more concise in its plans and summaries.
Claude Opus 4.5 is Anthropic's frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and reasoning benchmarks, and improved robustness to prompt injection.
Claude Opus 4.6 is Anthropic's most capable reasoning model, building on Opus 4.5 with enhanced performance across complex software engineering, agentic workflows, and long-horizon tasks. It features a 1M token context window, improved multimodal capabilities, and stronger robustness to prompt injection.
Claude Opus 4.7 is Anthropic's most capable generally available model for complex reasoning and agentic coding. It features a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports long-horizon agentic work, complex multi-step coding, and memory-driven tasks where coherence over extended sessions matters. It features a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.
Claude Opus 4.8 (Fast) is a speed-optimized variant of Anthropic's most capable generally available Opus model, offering the same 1M token context window and strong performance across long-horizon agentic work and complex coding — with lower latency.
Claude Opus 5 is Anthropic's most capable model in the Opus family. It delivers major gains over Opus 4.8 in agentic coding, professional knowledge work, and long-horizon reasoning, with a 1M token context window, 128K max output tokens, adaptive thinking, and strong multimodal capabilities.
Claude Opus 5 (Fast) is a speed-optimized variant of Anthropic's most capable Opus model, offering the same 1M token context window and strong performance across agentic coding, professional knowledge work, and long-horizon reasoning — with lower latency.
Claude Sonnet 4.5 is Anthropic's balanced model offering strong performance on coding, reasoning, and general tasks with good speed and cost efficiency.
Claude Sonnet 4.6 is Anthropic's best combination of speed and intelligence, offering strong performance on coding, reasoning, and general tasks with excellent speed and cost efficiency. It features a 1M token context window and 64K max output tokens.
Claude Sonnet 5 is Anthropic's latest Sonnet model, substantially improving on Sonnet 4.6 in coding and agentic work and reaching near-Opus quality on many tasks. It features a 1M token context window, adaptive thinking, and strong document and vision understanding.
DeepSeek-V3.2 is an efficient large language model with DeepSeek Sparse Attention (DSA) for long contexts. It features strong reasoning and tool-use skills, achieving top results on the 2025 IMO and IOI.
DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.
DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.
DeepSeek V4 Pro is a 1.6T-parameter Mixture-of-Experts model with 49B active parameters and a 1M-token context window. Built for advanced reasoning, coding, and long-horizon agentic workflows with a hybrid attention system for efficient long-context processing.
DeepSeek V4 Pro is a 1.6T-parameter Mixture-of-Experts model with 49B active parameters and a 1M-token context window. Built for advanced reasoning, coding, and long-horizon agentic workflows with a hybrid attention system for efficient long-context processing.
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro level reasoning with substantially lower latency.
Gemini 3.1 Pro is the latest evolution of Google flagship frontier model with 1M context, advancing high-precision multimodal reasoning across text, image, and code.
Gemini 3.5 Flash is a high speed, high value thinking model with 1M context, designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro level reasoning with substantially lower latency.
Gemini 3.5 Flash-Lite is the fastest, most cost-efficient Gemini 3.5 model with 1M context, ideal for everyday questions, summarization, and lightweight coding tasks.
Gemini 3.6 Flash is a high speed, high value thinking model with 1M context, designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro level reasoning with substantially lower latency.
Gemini 3.7 Flash is Google's most capable Flash model, built for complex coding, agentic workflows, and reliable multi-step execution, with 1M context and tunable thinking.
Gemini 3.8 Flash is Google's most capable Flash model, built for complex coding, agentic workflows, and reliable multi-step execution, with 1M context and tunable thinking.
Gemma 4 26B A4B Uncensored running in a Trusted Execution Environment (TEE). An uncensored variant of Google's Gemma 4 MoE model with 25.2B total / 3.8B active parameters, supporting multimodal input across text and images, with hardware attestation evidence available for independent verification.
Gemma 4 31B is a dense model from Google DeepMind with 31B parameters, delivering frontier-level reasoning performance. It handles text, image, and video input, supports 256K context, function calling, and configurable thinking modes.
Gemma 4 Uncensored is an uncensored variant of Google Gemma 4 26B, a Mixture-of-Experts model with 26B total parameters and only 4B active per token. Fine-tuned for uncensored chat without content filtering, it supports 256K context, coding, and general-purpose conversation.
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to Gemma 2.
Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.
Mercury 2 is a diffusion-based reasoning LLM from Inception, delivering over 1,000 tokens per second — 5x faster than leading speed-optimized models — with strong reasoning, tool use, and structured output capabilities.
Hermes 3 405B is a frontier level, full parameter finetune of the Llama-3.1 405B foundation model, focused on aligning LLMs to the user, with powerful steering capabilities and control given to the end user.
Llama 3.2 3B is a text model.
Llama 3.3 70B is a text model.
MiniMax-M2.5 is a state-of-the-art large language model optimized for coding, agentic workflows, and modern application development with enhanced reasoning capabilities.
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity with advanced agentic capabilities through multi-agent collaboration.
MiniMax-M3 preview is a 1.4T-parameter frontier model from MiniMax for coding, agentic workflows, and complex reasoning, served at fp8 with a 512K context window.
Mistral Small 3.2 is a 24B parameter model optimized for efficiency and performance. Ideal for general-purpose tasks with balanced speed and capability.
Mistral Small 4 unifies instruction following, reasoning, coding, and vision in a single 119B MoE model with 256K context and configurable reasoning effort.
Kimi K2.5 is Moonshot AIs most advanced open reasoning model, featuring trillion-parameter Mixture-of-Experts architecture with 32B active parameters and 256K context windows.
Kimi K2.6 is an open-source, native multimodal agentic model from Moonshot AI with 1T total parameters and 32B active parameters. It excels at long-horizon coding, coding-driven design, agent swarm orchestration, and proactive autonomous execution with 256K context windows.
Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built on Kimi K2.6, with 1T total parameters and 32B active parameters. It always operates in thinking mode, supports text and image input, and targets long-horizon software engineering, agentic task decomposition, and multi-turn coding workflows with 256K context.
Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.
Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.
NVIDIA Nemotron 3 Nano 30B is a compact and efficient language model from NVIDIA, optimized for fast inference while maintaining strong performance across diverse tasks.
NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.
OpenAI's multimodal flagship model with vision capabilities, strong reasoning, and broad knowledge. Popular for its balanced performance across tasks. Version: 2024-11-20.
OpenAI's cost-efficient small model that delivers GPT-4 level intelligence at a fraction of the cost. Ideal for high-volume applications requiring strong reasoning. Version: 2024-07-18.
GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context performance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly to simple queries while spending more depth on complex tasks.
GPT-5.2 Codex is OpenAI specialized coding model built on GPT-5.2, optimized for advanced software development, code generation, and technical problem-solving.
GPT-5.3 Codex is OpenAI specialized coding model built on GPT-5.3, optimized for advanced software development, code generation, and technical problem-solving.
GPT-5.4 is the latest frontier model in the GPT-5 series with a 1M+ context window, offering improved agentic and long context performance. It uses adaptive reasoning to dynamically allocate computation across tasks.
GPT-5.4 Mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use.
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning for complex, high-stakes tasks. It provides a 1M+ token context window (922K input, 128K output) and supports text and image inputs.
GPT-5.5 is the latest frontier model in the GPT-5 series with a 1M+ context window, offering improved agentic and long context performance. It uses adaptive reasoning to dynamically allocate computation across tasks.
GPT-5.5 Pro is OpenAI's most advanced model, building on GPT-5.5's unified architecture with enhanced reasoning for complex, high-stakes tasks. It provides a 1M+ token context window (922K input, 128K output) and supports text and image inputs.
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks where capability and cost need to be balanced.
GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex tasks.
GPT-6 Astra is OpenAI's most capable model, built for the hardest end-to-end work. It is suited for complex reasoning, coding, computer use, research, and document creation, with a 1.05M token context window (922K input, 128K output) and support for text and image inputs.
GPT-6 Astra with pro reasoning mode for difficult tasks that benefit from more model work. Supports text and image inputs with a 1.05M token context window. Pro mode can use more tokens and take longer than standard Astra.
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation
Qwen 2.5 7B Instruct running in a Trusted Execution Environment (TEE). A compact model with strong coding, math, and multilingual capabilities supporting 29+ languages, with hardware attestation evidence available for independent verification.
Built for in-depth research and handling long, complex documents. Ideal for technical work, multimodal input, and high-precision tasks.
Built for in-depth research and handling long, complex documents. Ideal for technical work, multimodal input, and high-precision tasks.
Turbo variant of Qwen3 Coder 480B, optimized for faster inference on code tasks.
Optimized for speed and efficiency.
Qwen 3.5 35B A3B is a highly efficient MoE model with 35B total parameters and only 3B active parameters. It surpasses the larger Qwen3-235B-A22B while being 6.7x smaller, excelling at reasoning, coding, and general knowledge tasks.
Qwen 3.5 is Alibaba flagship reasoning model featuring a 397B parameter Mixture-of-Experts architecture with 17B active parameters. It excels at complex reasoning, coding, and general knowledge tasks.
A 9B dense model with 262K native context window (extendable to 1M). Features Gated DeltaNet hybrid attention architecture for efficient long-context processing. Supports 201 languages, thinking/reasoning mode, and function calling.
The Qwen 3.6 27B native vision-language dense model builds upon the 3.5-27B version, with key improvements in agentic coding capabilities and enhanced STEM reasoning and inference skills. In the vision modality, it demonstrates significant advances in spatial intelligence, object localization, and detection, while video understanding, document OCR, and visual agent capabilities continue to improve steadily.
Qwen 3.6 35B A3B is a fast mixture-of-experts model with 35B total parameters and ~3B active per token. Strong at agentic coding, STEM reasoning, and tool use, with a native 256K context window.
Qwen 3.6 Plus Uncensored is Alibaba's latest flagship reasoning model with exceptional performance across coding, reasoning, and general knowledge tasks. Features mixed reasoning, function calling, and multimodal input support.
Qwen 3.7 Max is the largest model in the Qwen 3.7 series, with deep thinking, function calling, prompt caching, and multimodal input support for images and video. It excels at programming, office and productivity tasks, and long-running autonomous agent workflows.
Qwen 3.7 Plus is Alibaba's latest flagship reasoning model with exceptional performance across coding, reasoning, and general knowledge tasks. Features mixed reasoning, function calling, and multimodal input support.
Qwen 3.8 2.4T is Alibaba's open-weight 2.4-trillion-parameter MoE model (95B active), with major gains in software engineering, research, and long-horizon agentic tasks. It is text-only, requires thinking mode, and supports a 262K-token context window.
Qwen 3.8 27B is a native vision-language dense model with 27B parameters. It improves coding, professional work, research, and long-horizon agentic tasks, with flexible thinking control and image and video understanding. It supports a native 262K-token context window.
Qwen 3.8 Max is Alibaba's flagship 2.4-trillion-parameter MoE model, with major gains over Qwen 3.7 Max in software engineering and office-productivity workflows and strong long-horizon, multi-agent performance. It accepts both text and vision-language input (images and video), operates in thinking mode only, and supports a 1M-token context window.
Qwen3-VL 235B vision-language model with MoE architecture. The most powerful VL model in the Qwen series with superior visual perception, OCR, and multimodal reasoning.
Inkling is a general-purpose multimodal model from Thinking Machines Lab that accepts text, image, and audio inputs and generates text. It is a 66-layer sparse MoE (975B total / 41B active) with hybrid local/global attention, 512K context, and variable thinking effort — suited for chat, coding, tool use, and agentic workflows. Video input is not supported on Venice.
Seed 2.1 Turbo (Dola-Seed-2.1) is ByteDance’s next-generation multimodal model for the coding and agent era, with engineering-grade code delivery, long-horizon agent execution, and upgraded GUI and video understanding. Supports text, image, and video inputs with a 256K context window.
Optimized for creative roleplay scenarios with maximum freedom. Designed for immersive storytelling, character interactions, and open-ended creative writing.
Venice Uncensored 1.2 is designed for maximum creative freedom and authentic interaction. Built for open-ended exploration, roleplay, and unfiltered dialogue with improved capabilities over 1.1.
Grok 4.20 is xAI's latest multimodal reasoning model with strong tool use, structured output support, and a 2M-token context window.
Grok 4.20 Multi-Agent is a variant of xAI Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information across complex tasks.
Grok 4.3 is xAI's most intelligent and fastest reasoning model with function calling, structured outputs, and a 1M-token context window. Suited for agentic workflows, instruction-following tasks, and applications requiring high factual accuracy.
Grok 4.5 is xAI's intelligent coding model for agentic software engineering and workflow tasks, with function calling, structured outputs, and a 500K-token context window.
Grok 4.6 is xAI's multimodal chat and reasoning model with function calling, structured outputs, adjustable reasoning effort (low/medium/high/xhigh), and a 500K-token context window.
xAI's fast coding model trained specifically for agentic coding, currently in early access.
GLM-4.6 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.
GLM-4.7 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.
GLM-4.7-Flash-Heretic is an uncensored experimental variant of GLM-4.7-Flash, optimized for creative freedom and unfiltered dialogue with fast inference speed.
GLM-5 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple languages. Supports large context windows for processing extensive text and detailed analysis.
GLM-5 Turbo is a fast inference model from Z.ai tuned for strong performance in agent-driven environments and production coding workflows.
GLM-5.1 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple languages. Supports large context windows for processing extensive text and detailed analysis with fast inference speed.
GLM 5.2 running in a Trusted Execution Environment (TEE). Z.AI's flagship model for long-horizon tasks with enhanced reasoning and project-level engineering context, with hardware attestation evidence available for independent verification.
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.
GLM-5.3 Flash is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.
GLM-5V-Turbo is Z.ai's first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks with image, video, and text inputs.
DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks.
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities.
MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. Built upon the MiMo-V2-Flash backbone and extended with dedicated vision and audio encoders, it delivers robust performance across multimodal perception, long-context reasoning, and agentic workflows.
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.
GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks.
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.
GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks.
Qwen 3.6 35B A3B FP8 running in a Trusted Execution Environment (TEE). A fast mixture-of-experts model with ~3B active parameters per token. Hardware attestation evidence is available for independent verification of enclave identity and configuration.
GLM-4.7-Flash is a fast inference variant of GLM-4.7, optimized for speed while maintaining strong reasoning capabilities. Ideal for applications requiring quick responses with good quality.
Wan 2.7 Pro is a text-to-image model. High-fidelity results with strong prompt adherence and fast turnaround.
Wan 2.7 is a text-to-image model. High-fidelity results with strong prompt adherence and fast turnaround.
Z-Image Turbo is a text-to-image model. High-fidelity results with strong prompt adherence and fast turnaround.
Embedding model for semantic search and retrieval.
Embedding model for semantic search and retrieval.
Flux 2 Max is a text-to-image model. Photorealistic, high-detail output with excellent typography and layout.
Flux 2 Pro is a text-to-image model. Photorealistic, high-detail output with excellent typography and layout.
Seedream V4.5 is a text-to-image model. Vivid color, sharp detail, and coherent multi-subject scenes.
Seedream V5 Lite is a text-to-image model. Vivid color, sharp detail, and coherent multi-subject scenes.
ElevenLabs Turbo v2.5 is a text-to-speech model. Natural, expressive speech with clear articulation.
Gemini 3.1 Flash TTS is a text-to-speech model. Natural, expressive speech with clear articulation.
Embedding model for semantic search and retrieval.
Nano Banana 2 Lite is a text-to-image model. Creative, high-quality results with strong instruction following and clean edits.
Ideogram V4 is a text-to-image model. Best-in-class in-image text and typography.
Krea v2 Large is a text-to-image model. Aesthetic-focused, with a distinctive and stylized look.
Krea v2 Medium is a text-to-image model. Aesthetic-focused, with a distinctive and stylized look.
Luma Uni-1 is a text-to-image model. Cinematic lighting and composition with photoreal detail.
Luma Uni-1 Max is a text-to-image model. Cinematic lighting and composition with photoreal detail.
Embedding model for semantic search and retrieval.
Clone your voice from a short recording and generate natural speech in it across 30+ languages.
Embedding model for semantic search and retrieval.
GPT Image 1.5 is a text-to-image model. Instruction-following generation with broad world knowledge and reliable text rendering.
Embedding model for semantic search and retrieval.
Embedding model for semantic search and retrieval.
Qwen 3 TTS 0.6B is a text-to-speech model. Natural, expressive speech with clear articulation.
Qwen 3 TTS 1.7B is a text-to-speech model. Natural, expressive speech with clear articulation.
Qwen Image is a text-to-image model. Flexible across a wide range of prompts and edits.
Qwen Image 2 is a text-to-image model. Flexible across a wide range of prompts and edits.
Qwen Image 2 Pro is a text-to-image model. Flexible across a wide range of prompts and edits.
Qwen Image 3 is a text-to-image model. Flexible across a wide range of prompts and edits.
Embedding model for semantic search and retrieval.
Embedding model for semantic search and retrieval.
Recraft V4 is a text-to-image model. Brand- and vector-friendly, with precise control over style.
Recraft V4 Pro is a text-to-image model. Brand- and vector-friendly, with precise control over style.
Lustify SDXL is a text-to-image model. Open-weight, with fine-grained control over style and composition.
Venice SD35 is a text-to-image model. Open-weight, with fine-grained control over style and composition.
Hunyuan Image 3.0 is a text-to-image model. Smooth motion with strong scene and subject consistency.
Anime (WAI) is a text-to-image model.
Background Remover is a text-to-image model.
Chatterbox HD (Resemble AI) is a text-to-speech model. Natural, expressive speech with clear articulation.
Chroma is a text-to-image model.
Gradium TTS is a text-to-speech model. Natural, expressive speech with clear articulation.
ImagineArt 1.5 Pro is a text-to-image model.
Inworld TTS-1.5 Max is a text-to-speech model. Natural, expressive speech with clear articulation.
Kokoro Text to Speech is a text-to-speech model. Natural, expressive speech with clear articulation.
Lustify v7 is a text-to-image model.
Lustify v8 is a text-to-image model.
Muse Image is a text-to-image model.
Orpheus TTS is a text-to-speech model. Natural, expressive speech with clear articulation.
xAI TTS v1 is a text-to-speech model. Natural, expressive speech with clear articulation.
Feature-rich song generation with optional lyrics and detailed musical controls.
Wan 2.7 Pro Edit is an image editing and inpainting model. High-fidelity results with strong prompt adherence and fast turnaround.
Flux 2 Max is an image editing and inpainting model. Photorealistic, high-detail output with excellent typography and layout.
Seedream V4.5 is an image editing and inpainting model. Vivid color, sharp detail, and coherent multi-subject scenes.
Seedream V5 Lite is an image editing and inpainting model. Vivid color, sharp detail, and coherent multi-subject scenes.
Seedream V5 Pro is a text-to-image model. Vivid color, sharp detail, and coherent multi-subject scenes.
Seedream V5 Pro is an image editing and inpainting model. Vivid color, sharp detail, and coherent multi-subject scenes.
Multilingual text-to-speech using ElevenLabs. Supports 29 languages with high-quality natural-sounding voices, configurable speed, and accent accuracy.
High-quality instrumental music generation with configurable duration. Best for polished, production-ready tracks across a wide range of genres.
ElevenLabs Scribe V2 is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.
Generate high-quality sound effects from text descriptions using ElevenLabs. Ideal for films, games, and digital content with configurable duration.
Generate natural text-to-speech audio using ElevenLabs Eleven-v3. High-quality voices with stability control and automatic text normalization.
Google's Lyria 3 Pro generates full-length, structured songs up to 3 minutes long from a single text prompt. Supports vocals, lyrics, and multi-language generation across genres.
Nano Banana 2 is a text-to-image model. Creative, high-quality results with strong instruction following and clean edits.
Nano Banana 2 is an image editing and inpainting model. Creative, high-quality results with strong instruction following and clean edits.
Nano Banana 2 Lite is an image editing and inpainting model. Creative, high-quality results with strong instruction following and clean edits.
Nano Banana Pro is a text-to-image model. Creative, high-quality results with strong instruction following and clean edits.
Nano Banana Pro is an image editing and inpainting model. Creative, high-quality results with strong instruction following and clean edits.
Krea 2 Turbo is a text-to-image model. Aesthetic-focused, with a distinctive and stylized look.
Luma Uni-1 is an image editing and inpainting model. Cinematic lighting and composition with photoreal detail.
Luma Uni-1 Max is an image editing and inpainting model. Cinematic lighting and composition with photoreal detail.
Full song generation with vocals and lyrics. Provide your own lyrics with verse/chorus structure for complete songs with singing.
Advanced song generation with vocals, lyrics optimizer, and instrumental mode. Supports structure tags and up to 3500 character lyrics.
Latest MiniMax song generation with vocals, instrumental mode, and support for rich structure tags in lyrics.
Parakeet ASR is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.
GPT Image 1.5 is an image editing and inpainting model. Instruction-following generation with broad world knowledge and reliable text rendering.
GPT Image 2 is a text-to-image model. Instruction-following generation with broad world knowledge and reliable text rendering.
GPT Image 2 is an image editing and inpainting model. Instruction-following generation with broad world knowledge and reliable text rendering.
Whisper Large V3 is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.
Wizper (Whisper v3) is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.
Qwen Edit Uncensored is an image editing and inpainting model. Flexible across a wide range of prompts and edits.
Qwen Image 2 is an image editing and inpainting model. Flexible across a wide range of prompts and edits.
Qwen Image 2 Pro is an image editing and inpainting model. Flexible across a wide range of prompts and edits.
Qwen Image 3 Edit is an image editing and inpainting model. Flexible across a wide range of prompts and edits.
Qwen Image 3 Pro is a text-to-image model. Flexible across a wide range of prompts and edits.
Qwen Image 3 Pro Edit is an image editing and inpainting model. Flexible across a wide range of prompts and edits.
Fast, lightweight audio generation for sound effects, ambient textures, and short musical clips. Flexible duration from 5 seconds to over 3 minutes.
FireRed Edit is an image editing and inpainting model.
Generate synchronized audio and sound effects from text prompts with MMAudio V2.
Muse Image is an image editing and inpainting model.
Generate expressive multilingual speech and audio from a text prompt with BytePlus Seed Audio 1.0 (20 languages, timestamp length control).
Generate licensed, commercial-use-safe music with precise control over style, mood, instrumentation, and duration.
Generate licensed, commercial-use-safe sound effects with precise control over type, texture, intensity, and duration.
Upscaler is an image upscaling model.
Grok Imagine is a text-to-image model.
Grok Imagine 2.0 is a text-to-image model.
Grok Imagine 2.0 is an image editing and inpainting model.
Grok Imagine is an image editing and inpainting model.
Grok Imagine High Quality is an image editing and inpainting model.
Grok Imagine High Quality (SOTA) is a text-to-image model.
xAI Speech to Text v1 is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.