Parakeet ASR is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.
Audio and Speech Models
collections/audio · 32 models
Audio routes covering speech synthesis, transcription, music, and sound design. Billing is per character or per minute depending on the model, and the unit is shown on the card.
Models
32
Labs
10
model creators
From (input)
Varies
per 1M tokens
Max context
Varies
Private routes
8 / 32
up to Private
In this collection32
Ordered by strongest privacy.
Whisper Large V3 is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.
Wizper (Whisper v3) is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.
Qwen 3 TTS 0.6B is a text-to-speech model. Natural, expressive speech with clear articulation.
Qwen 3 TTS 1.7B is a text-to-speech model. Natural, expressive speech with clear articulation.
Chatterbox HD (Resemble AI) is a text-to-speech model. Natural, expressive speech with clear articulation.
Kokoro Text to Speech is a text-to-speech model. Natural, expressive speech with clear articulation.
Orpheus TTS is a text-to-speech model. Natural, expressive speech with clear articulation.
Feature-rich song generation with optional lyrics and detailed musical controls.
Multilingual text-to-speech using ElevenLabs. Supports 29 languages with high-quality natural-sounding voices, configurable speed, and accent accuracy.
High-quality instrumental music generation with configurable duration. Best for polished, production-ready tracks across a wide range of genres.
High-quality realistic music generation with optional vocals and configurable duration. Best for polished, production-ready tracks across a wide range of genres.
Latest ElevenLabs music generation with higher-quality realistic tracks, optional vocals, and configurable duration. Best for polished, production-ready songs across a wide range of genres.
ElevenLabs Scribe V2 is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.
Generate high-quality sound effects from text descriptions using ElevenLabs. Ideal for films, games, and digital content with configurable duration.
Generate natural text-to-speech audio using ElevenLabs Eleven-v3. High-quality voices with stability control and automatic text normalization.
ElevenLabs Turbo v2.5 is a text-to-speech model. Natural, expressive speech with clear articulation.
Gemini 3.1 Flash TTS is a text-to-speech model. Natural, expressive speech with clear articulation.
Google's Lyria 3 Pro generates full-length, structured songs up to 3 minutes long from a single text prompt. Supports vocals, lyrics, and multi-language generation across genres.
Full song generation with vocals and lyrics. Provide your own lyrics with verse/chorus structure for complete songs with singing.
Advanced song generation with vocals, lyrics optimizer, and instrumental mode. Supports structure tags and up to 3500 character lyrics.
Latest MiniMax song generation with vocals, instrumental mode, and support for rich structure tags in lyrics.
Clone your voice from a short recording and generate natural speech in it across 30+ languages.
Fast, lightweight audio generation for sound effects, ambient textures, and short musical clips. Flexible duration from 5 seconds to over 3 minutes.
Gradium TTS is a text-to-speech model. Natural, expressive speech with clear articulation.
Inworld TTS-1.5 Max is a text-to-speech model. Natural, expressive speech with clear articulation.
Generate synchronized audio and sound effects from text prompts with MMAudio V2.
Generate expressive multilingual speech and audio from a text prompt with BytePlus Seed Audio 1.0 (20 languages, timestamp length control).
Generate licensed, commercial-use-safe music with precise control over style, mood, instrumentation, and duration.
Generate licensed, commercial-use-safe sound effects with precise control over type, texture, intensity, and duration.
xAI Speech to Text v1 is a speech-to-text transcription model. Accurate transcription across languages, accents, and noisy audio.
xAI TTS v1 is a text-to-speech model. Natural, expressive speech with clear articulation.
This list is rebuilt from the live catalog rather than stored as a snapshot, so it tracks pricing, context windows, and privacy tiers as providers change them. Ordering is yours to pick, and there is no popularity option: prompts are never retained, and the usage metadata kept for billing is not turned into a public ranking.
Explore more collections
Strongest guarantee in this collection:Private