30 similar Audio Models you might want to consider.
WavCraft is a Audio Models model developed by Microsoft Research. This page compares it with 30 other Audio Models models in the same category — 15 open source and 0 offering a free tier.
| Model | Developer | Release date | Context window | Max output tokens | Modalities (input → output) | Open source / licence | Free tier | API available |
|---|---|---|---|---|---|---|---|---|
| WavCraft | Microsoft Research | — | — | — | — | No | — | — |
| AudioCraft | Meta | Mar 1, 2025 | — | — | — | Yes | — | — |
| ElevenLabs v3 | ElevenLabs | Sep 1, 2025 | — | — | — | No | — | — |
| Suno v4 | Suno | Nov 1, 2025 | — | — | — | No | — | — |
| Whisper v4 | OpenAI | Jul 1, 2025 | — | — | — | Yes | — | — |
| Amphion | Open Source Community | — | — | — | — | Yes | — | — |
| AudioGen 2 | Meta | — | — | — | — | Yes | — | — |
| Bark 2 | Suno | May 1, 2025 | — | — | — | Yes | — | — |
| Coqui XTTS v2 | Coqui | Jan 1, 2024 | — | — | — | Yes | — | — |
| Deepgram Nova-3 | Deepgram | — | — | — | — | No | — | — |
| Dia | Nari Labs | Feb 1, 2025 | — | — | — | Yes | — | — |
| F5-TTS | SWivid | Apr 1, 2025 | — | — | — | Yes | — | — |
| Google SoundStream | Jan 1, 2024 | — | — | — | No | — | — | |
| Hume AI | Hume | — | — | — | — | No | — | — |
| Lyria 2 | Google DeepMind | Aug 1, 2025 | — | — | — | No | — | — |
| Mars5-TTS | CAMB.AI | — | — | — | — | No | — | — |
| MusicGen | Meta | Mar 1, 2025 | — | — | — | Yes | — | — |
| NotebookLM Audio | Apr 1, 2025 | — | — | — | No | — | — | |
| Parler-TTS | Hugging Face | — | — | — | — | Yes | — | — |
| PlayHT 3 | PlayHT | Jan 1, 2025 | — | — | — | No | — | — |
| Resemble AI | Resemble AI | Jun 1, 2024 | — | — | — | No | — | — |
| Riffusion v2 | Riffusion | Jun 1, 2025 | — | — | — | Yes | — | — |
| Sesame CSM | Sesame | Mar 1, 2025 | — | — | — | Yes | — | — |
| Sesame CSM 2 | Sesame AI | Dec 1, 2025 | — | — | — | Yes | — | — |
| SoundHound AI | SoundHound | — | — | — | — | No | — | — |
| Speechify AI | Speechify | Jun 1, 2024 | — | — | — | No | — | — |
| Stable Audio 3 | Stability AI | May 1, 2025 | — | — | — | Yes | — | — |
| Suno v4.5 | Suno | Jan 1, 2026 | — | — | — | No | — | — |
| Tortoise TTS | James Betker | Dec 1, 2024 | — | — | — | Yes | — | — |
| Udio v2 | Udio | Oct 1, 2025 | — | — | — | No | — | — |
| Voicebox 2 | Meta | Jun 1, 2025 | — | — | — | No | — | — |
Meta
Open-source audio generation suite including MusicGen and AudioGen.
ElevenLabs
Industry-leading voice synthesis with natural-sounding speech and voice cloning.
Suno
AI music creation platform generating full songs with vocals and lyrics.
OpenAI
State-of-the-art speech recognition supporting 100+ languages.
Open Source Community
Open-source audio toolkit for speech, music, and sound generation research, providing unified framework for audio synthesis tasks.
Meta
Second-generation audio generation model from Meta that creates realistic sound effects and environmental audio from text descriptions.
Suno
Generative audio model capable of speech, music, and sound effects.
Coqui
Open-source multilingual text-to-speech model with voice cloning from just 6 seconds of audio.
Deepgram
Third-generation speech-to-text AI model offering industry-leading transcription accuracy, speed, and cost efficiency with support for 40+ languages.
Nari Labs
Open-source dialogue audio generation model for creating realistic multi-speaker conversations.
SWivid
Fast and fluent text-to-speech with natural flow and minimal artifacts.
Neural audio codec for efficient high-quality audio compression and generation at low bitrates.
Hume
Empathic AI voice model that understands and generates speech with emotional intelligence, detecting tone, sentiment, and vocal expressions in real-time.
Google DeepMind
Google's music generation model powering YouTube's Dream Track.
CAMB.AI
Advanced text-to-speech model capable of generating highly natural speech with prosody control and voice cloning from short audio samples.
Meta
Text-to-music generation model producing high-quality musical compositions.
Podcast-style audio generation from documents, creating engaging conversational summaries.
Hugging Face
Open-source text-to-speech model that generates natural-sounding speech with controllable speaker characteristics described in natural language.
PlayHT
Advanced text-to-speech model generating ultra-realistic voices with natural prosody and emotional range.
Resemble AI
Enterprise voice AI platform for creating custom synthetic voices with emotional control and real-time generation.
Riffusion
Real-time music generation using spectral diffusion techniques.
Sesame
Conversational speech model with emotional expression and natural turn-taking capabilities.
Sesame AI
Next-gen conversational speech model with ultra-realistic voice synthesis and emotional expression.
SoundHound
Voice AI platform specializing in conversational intelligence with real-time speech recognition, natural language understanding, and voice commerce solutions.
Speechify
AI-powered text-to-speech platform with natural-sounding voices for accessibility and content consumption.
Stability AI
Open-source audio and music generation from text prompts.
Suno
Suno's latest music generation model with improved lyrics, instrumentation, and genre diversity.
James Betker
High-quality multi-voice TTS known for natural prosody.
Udio
Competing music generation platform with emphasis on audio quality.
Meta
Versatile speech generation with editing and style transfer.