Static

ElevenLabs’ new v4 speech model supports more expression control and 90 languages

First reported by TechCrunch ·

The signal ●○○○ Compiled by AI from TechCrunch, the single source so far
Why you might care

Your synthesized speech can now sound more nuanced, with greater emotional range and more accurate language representation.

What happened

ElevenLabs has released two new speech models, v4 and v4 Turbo, enhancing expression control and expanding language support to over 90 languages. The v4 models utilize a new architecture enabling faster voice cloning, requiring only 10 seconds of audio, and better handling of voice identity over longer text segments. These models maintain context to adjust expressions and support sequential control through expanded inline tags. The previous v3 model supported 70 languages, with significant quality improvements noted in Japanese, Brazilian Portuguese, Mandarin, and Cantonese. The new v4 generation is optimized for voice agents due to lower latency, allowing for more fluid, real-time conversational audio generation. This development occurs amidst increased competition from startups and major tech companies in the expressive speech model market. ElevenLabs recently raised $500 million at an $11 billion valuation and is reportedly seeking further funding.

What it means

The advancements in ElevenLabs' v4 model, particularly its improved expression control and expanded language support to 90 languages, signal a maturing market for generative AI voice technology. The ability to clone voices with just 10 seconds of audio and maintain context for nuanced expression suggests these models are moving beyond simple text-to-speech to sophisticated content generation tools. This increased sophistication directly impacts creators, developers of voice agents, and anyone requiring highly realistic and emotionally resonant synthetic voices, potentially lowering the barrier to entry for high-quality voiceover production.

This evolution in speech synthesis, coupled with lower latency for voice agents, indicates a push towards more seamless human-computer interaction. The competitive landscape, with other startups and tech giants also improving their voice models, suggests an acceleration in AI capabilities. For businesses relying on voice interfaces or synthetic media, the improved quality and control offered by ElevenLabs and its competitors mean more realistic and engaging user experiences, potentially transforming customer service, entertainment, and accessibility.

AI-written summary. May contain errors.

Tech