ElevenLabs Upgrades App with v4 Model, Natural Speech, and Hebrew Support
ElevenLabs has updated its text-to-speech app with the advanced v4 model, offering natural intonation, emotion control tags, and ultra-low latency for developers.

ElevenLabs has rolled out a major upgrade to its text-to-speech app and platform, introducing its advanced v4 model and Turbo version. The update significantly enhances the naturalness of generated speech, moving far beyond monotonous robotic voices.
Advanced v4 Architecture and Emotional Control
The transition to the v4 architecture changes how text-to-speech systems handle generated audio. The new model analyzes the full context of sentences and text before and during reading, dynamically adjusting intonation, speech tempo, and pauses. This makes listening to long articles, boring reports, or digital books smooth and natural.
One of the most interesting features is the ability to precisely control emotion and expression using built-in text tags. Users can insert specific commands directly into the text, such as [whisper], [laugh], [sigh], or [sad]. The system can process several tags in sequence, allowing the virtual voice to shift smoothly from a professional tone to a dramatic whisper.
High Performance for Developers
For developers building voice agents using the company's API, the v4 Turbo version delivers a major performance leap with exceptionally low latency of just around 100 milliseconds. This enables smooth, real-time conversations with bots powered by large language models, generating audio the moment the model outputs the first words.
Additional updates in the v4 release include support for 90 languages—including Hebrew, which was added about a year ago—and the ability to quickly clone any voice by providing voice samples.





