- What is the core architecture of Voicebox 2?
- Voicebox 2 uses a non-autoregressive flow-matching approach that allows for parallel generation of speech. This design makes it faster and less prone to cumulative errors compared to traditional sequential models.
- What tasks can Voicebox 2 perform?
- The model is designed for zero-shot text-to-speech, style transfer, and speech editing. It can handle tasks such as noise removal, content editing, and style modification with high fidelity.
- How was Voicebox 2 trained?
- It was trained on over 50,000 hours of speech and transcripts from publicly available audiobooks. The training utilized self-supervised learning techniques and a novel flow-matching objective to learn rich speech representations.
- Is Voicebox 2 available for public use?
- No, Voicebox 2 is currently a research preview and is not publicly released. This proprietary status limits accessibility for researchers and developers outside of Meta.
- What are the main ethical concerns associated with Voicebox 2?
- The model's powerful capabilities raise significant ethical concerns regarding the potential for misuse in creating deepfakes or impersonating individuals. Additionally, it may inherit biases from its training data.