30 similar Audio Models you might want to consider.
Resemble AI is a Audio Models model developed by Resemble AI. This page compares it with 30 other Audio Models models in the same category — 15 open source and 0 offering a free tier.
| Model | Developer | Release date | Context window | Max output tokens | Modalities (input → output) | Open source / licence | Free tier | API available |
|---|---|---|---|---|---|---|---|---|
| Resemble AI | Resemble AI | Jun 1, 2024 | — | — | — | No | — | — |
| AudioCraft | Meta | Mar 1, 2025 | — | — | — | Yes | — | — |
| ElevenLabs v3 | ElevenLabs | Sep 1, 2025 | — | — | — | No | — | — |
| Suno v4 | Suno | Nov 1, 2025 | — | — | — | No | — | — |
| Whisper v4 | OpenAI | Jul 1, 2025 | — | — | — | Yes | — | — |
| Amphion | Open Source Community | — | — | — | — | Yes | — | — |
| AudioGen 2 | Meta | — | — | — | — | Yes | — | — |
| Bark 2 | Suno | May 1, 2025 | — | — | — | Yes | — | — |
| Coqui XTTS v2 | Coqui | Jan 1, 2024 | — | — | — | Yes | — | — |
| Deepgram Nova-3 | Deepgram | — | — | — | — | No | — | — |
| Dia | Nari Labs | Feb 1, 2025 | — | — | — | Yes | — | — |
| F5-TTS | SWivid | Apr 1, 2025 | — | — | — | Yes | — | — |
| Google SoundStream | Jan 1, 2024 | — | — | — | No | — | — | |
| Hume AI | Hume | — | — | — | — | No | — | — |
| Lyria 2 | Google DeepMind | Aug 1, 2025 | — | — | — | No | — | — |
| Mars5-TTS | CAMB.AI | — | — | — | — | No | — | — |
| MusicGen | Meta | Mar 1, 2025 | — | — | — | Yes | — | — |
| NotebookLM Audio | Apr 1, 2025 | — | — | — | No | — | — | |
| Parler-TTS | Hugging Face | — | — | — | — | Yes | — | — |
| PlayHT 3 | PlayHT | Jan 1, 2025 | — | — | — | No | — | — |
| Riffusion v2 | Riffusion | Jun 1, 2025 | — | — | — | Yes | — | — |
| Sesame CSM | Sesame | Mar 1, 2025 | — | — | — | Yes | — | — |
| Sesame CSM 2 | Sesame AI | Dec 1, 2025 | — | — | — | Yes | — | — |
| SoundHound AI | SoundHound | — | — | — | — | No | — | — |
| Speechify AI | Speechify | Jun 1, 2024 | — | — | — | No | — | — |
| Stable Audio 3 | Stability AI | May 1, 2025 | — | — | — | Yes | — | — |
| Suno v4.5 | Suno | Jan 1, 2026 | — | — | — | No | — | — |
| Tortoise TTS | James Betker | Dec 1, 2024 | — | — | — | Yes | — | — |
| Udio v2 | Udio | Oct 1, 2025 | — | — | — | No | — | — |
| Voicebox 2 | Meta | Jun 1, 2025 | — | — | — | No | — | — |
| WavCraft | Microsoft Research | — | — | — | — | No | — | — |
Meta
Open-source audio generation suite including MusicGen and AudioGen.
ElevenLabs
Industry-leading voice synthesis with natural-sounding speech and voice cloning.
Suno
AI music creation platform generating full songs with vocals and lyrics.
OpenAI
State-of-the-art speech recognition supporting 100+ languages.
Open Source Community
Open-source audio toolkit for speech, music, and sound generation research, providing unified framework for audio synthesis tasks.
Meta
Second-generation audio generation model from Meta that creates realistic sound effects and environmental audio from text descriptions.
Suno
Generative audio model capable of speech, music, and sound effects.
Coqui
Open-source multilingual text-to-speech model with voice cloning from just 6 seconds of audio.
Deepgram
Third-generation speech-to-text AI model offering industry-leading transcription accuracy, speed, and cost efficiency with support for 40+ languages.
Nari Labs
Open-source dialogue audio generation model for creating realistic multi-speaker conversations.
SWivid
Fast and fluent text-to-speech with natural flow and minimal artifacts.
Neural audio codec for efficient high-quality audio compression and generation at low bitrates.
Hume
Empathic AI voice model that understands and generates speech with emotional intelligence, detecting tone, sentiment, and vocal expressions in real-time.
Google DeepMind
Google's music generation model powering YouTube's Dream Track.
CAMB.AI
Advanced text-to-speech model capable of generating highly natural speech with prosody control and voice cloning from short audio samples.
Meta
Text-to-music generation model producing high-quality musical compositions.
Podcast-style audio generation from documents, creating engaging conversational summaries.
Hugging Face
Open-source text-to-speech model that generates natural-sounding speech with controllable speaker characteristics described in natural language.
PlayHT
Advanced text-to-speech model generating ultra-realistic voices with natural prosody and emotional range.
Riffusion
Real-time music generation using spectral diffusion techniques.
Sesame
Conversational speech model with emotional expression and natural turn-taking capabilities.
Sesame AI
Next-gen conversational speech model with ultra-realistic voice synthesis and emotional expression.
SoundHound
Voice AI platform specializing in conversational intelligence with real-time speech recognition, natural language understanding, and voice commerce solutions.
Speechify
AI-powered text-to-speech platform with natural-sounding voices for accessibility and content consumption.
Stability AI
Open-source audio and music generation from text prompts.
Suno
Suno's latest music generation model with improved lyrics, instrumentation, and genre diversity.
James Betker
High-quality multi-voice TTS known for natural prosody.
Udio
Competing music generation platform with emphasis on audio quality.
Meta
Versatile speech generation with editing and style transfer.
Microsoft Research
AI audio editing and generation system from Microsoft that uses LLMs to understand and execute complex audio manipulation tasks from text instructions.