ElevenLabs releases Eleven v4 and Turbo

The speech models target expressive audio and quicker voice responses. A synthesis latency figure does not measure a complete conversation.

ElevenLabs, a synthetic audio developer, announced Eleven v4 and Eleven v4 Turbo on September 28, adding new text-to-speech models for produced audio and responsive voice applications.

The company turns written words into spoken audio. Its new models aim to improve how speech sounds and how quickly it begins. Developers need to assess both qualities in the conversations they intend to support.

Why it matters: A voice assistant can give the right answer and still be difficult to use if it speaks unnaturally or responds too slowly. The launch provides new options for testing those problems, while leaving accuracy, user consent and complete application performance as separate questions.

ElevenLabs’ own announcement describes Turbo as a low-latency variant and reports approximately 100 milliseconds of median inference latency. A median is the middle observation in a set of measurements. It should not be read as a maximum delay, nor as the time every caller will wait for an entire answer.

That distinction changes the evaluation. A proposed test should start when a person finishes speaking and end when the application begins useful speech. It should also measure unusually slow responses. Measuring only the synthesis stage would leave other stages of the conversation outside the result.

The company says the models better interpret tone, pacing and context, with improvements across more than 90 languages. Those are developer claims. They give evaluators concrete dimensions to inspect, but do not establish uniform quality across languages, accents, speakers or recording conditions.

A useful listening test would therefore separate intelligibility from performance style. Reviewers could first check whether names, numbers and key instructions are spoken correctly, then judge whether the delivery suits the situation. A more dramatic voice is not necessarily preferable for a support call or a factual announcement.

The same separation applies to dialogue. A conversation may sound fluid while assigning a line to the wrong speaker or delivering an instruction with the wrong emphasis. Testing should use scripts with known intended meanings, so that an attractive sample cannot conceal a content error.

Test the voice inside the complete application

ElevenLabs lists the release across its creative products, agent products and developer interface. That makes the launch relevant to both prerecorded output and interactive applications. The acceptance criteria for those uses should differ: an editor can review a recorded clip before publication, while a live caller hears the first version immediately.

For recorded work, an appropriate comparison would include revision effort. If a producer must regenerate an entire passage to correct one word, the cost of a usable clip may differ from the cost of the first generation. Keep the script, selected voice and review standard consistent when comparing alternatives.

For an interactive assistant, consider a caller interrupting midway through an answer. A practical test would ask whether the application stops, understands the correction and resumes with the right information. These are proposed application tests; the reported inference figure does not tell us how the complete system performs them.

Voice cloning raises another distinct acceptance question: whether the person whose voice is represented has authorized the intended use. Higher similarity is a technical objective, not evidence of permission. A production review should keep those two decisions separate and document the approved speaker and use case.

The release also warrants caution about broad rankings. A preference score in a listening comparison depends on the samples, languages and alternatives selected. A leaderboard position is useful only alongside the evaluation record, including which competing systems were tested and how listeners were selected.

The next useful evidence is a reproducible comparison of finished audio and full conversations, including slower cases and correction handling. Teams can then decide whether the new voices improve the experience their listeners actually receive, instead of treating a synthesis benchmark as a complete service guarantee.

Verification

Claim groupTierPrimary evidence
September 28 release and product accessVERIFIEDElevenLabs launch post
Turbo latency, expressive controls, language coverage and cloning claimsPARTIALLY VERIFIED — company-reported; not independently tested hereElevenLabs’ company announcement
Whole-conversation, listening and consent distinctionsAnalysis and proposed evaluation criteria; no performance outcome assertedInference from the launch scope above

Glossary candidates: text-to-speech — converting written text into audio; median latency — middle observed delay; voice cloning — generating speech with characteristics of a reference voice.

Cold-reader sentence: ElevenLabs released expressive v4 speech models and a Turbo variant, but its synthesis speed claim does not establish full voice-assistant response time.