Cartesia announced Sonic 3.6 on August 27, 2026, expanding the languages, locales and voice options available for text-to-speech. A broader catalog matters for multilingual products, but adoption should depend on pronunciation, interruption behavior and total conversational delay in the intended deployment.
Coverage grows beyond a single language label
The official announcement lists 44 languages, 61 locales and expanded Indic support, alongside a larger voice catalog. It also discusses locale handling for mixed-language speech and formatting-sensitive content. These are provider descriptions of the release, not independent measurements from Nerova.
Language support is a starting point. A product may need regional pronunciation, names, abbreviations and code switching within one sentence. Build a listening evaluation with speakers from the target audience and include the actual terms users will hear.
Measure the complete conversation path
The provider's latency claims concern speech synthesis under its measured conditions. A live assistant also spends time capturing audio, recognizing speech, choosing an answer and transporting output. A fast synthesis engine cannot remove delay elsewhere in that chain.
Measure time from the user's completed utterance to useful audible response, along with interruption handling. Test weak connections and competing sessions. Report percentile behavior rather than only the fastest response, because delayed turns can damage a conversation even when the average looks acceptable.
Plans and concurrency affect rollout
The provider's plan page documents commercial plan features and operating limits. Those current terms should be checked before deployment; they are not a historical price guarantee for every August customer.
Model volume using the number of simultaneous conversations and generated audio, not only request count. A system that buffers long answers may consume more resources while worsening the user experience. Keep answers concise when the task benefits from rapid turn taking, and observe queue behavior under load.
Voice selection is a product responsibility
Evaluate cloned or recognizable voices only with appropriate permission for the intended use. Keep a record of what voice was selected and make the product's identity clear to listeners. Technical voice fidelity does not establish permission to impersonate a person.
Sonic 3.6 is most useful where expanded coverage removes a concrete language or locale gap. Compare it with the current voice path on the same scripts and conditions, then roll out by locale with a way to report pronunciation failures. Catalog size alone does not establish a better conversation.