As Competition in Real-Time Voice AI Grows, Soniox Releases 'TTS v2' With Multilingual Speech and Voice Cloning

Voice interfaces have moved from experimental features to core components of many applications in recent years. 

Developers building conversational systems now have to balance cost, latency, language coverage, pronunciation accuracy, and performance at scale, particularly when voice is generated during live interactions.

Soniox is entering that market with 'TTS v2,' available through its tts-rt-v2 model. 

The system converts text into speech and supports more than 60 languages, voice cloning from short samples, expressive controls through audio tags such as whispering and laughter, multilingual speech within a single utterance, and streaming generation.

The model also provides character-level timestamps, which can be used to synchronize generated speech with text and manage interruptions in conversational applications.

Soniox lists pricing at approximately $0.70 per hour of generated speech, with the final cost calculated from input text and output audio token usage. The company, founded in 2020 by Klemen Simonic and Ambroz Bizjak and based in Foster City, California, previously focused primarily on speech-to-text before expanding further into speech generation.

Initial reactions from users testing the release have focused on the quality of cloned voices and its ability to handle mixed-language speech. 

Some testers have specifically highlighted English-Japanese utterances, although broader independent testing will be needed to establish how consistently the model performs across languages, voices, and workloads.

The wider text-to-speech market already includes several established providers. 

ElevenLabs is widely used for natural-sounding and expressive speech, while its Flash models are designed for low-latency applications. OpenAI's tts-1 is priced at $15 per million characters and offers a smaller selection of voices without publicly available high-fidelity voice cloning. Cartesia's Sonic models focus heavily on streaming latency, with reported figures frequently below 100 milliseconds in certain configurations.

Cloud platforms such as Google Cloud and Microsoft Azure offer extensive language and voice selections, with Azure supporting more than 100 locales in some parts of its speech platform. 

Their approaches to custom voice cloning and related capabilities vary by product and account level.

Against these systems, Soniox's TTS v2 combines several capabilities that are often distributed across different products. 

These include voice cloning, expressive speech controls, multilingual output, structured-text pronunciation, and character-level timing information. 

The model is also designed to handle elements such as email addresses, phone numbers, and identifiers directly.

Those features could be relevant to applications such as voice agents, customer-support systems, education tools, and interactive services, although their practical value will depend on how reliably they work under real-world conditions. 

Character-level timestamps, for example, can help applications determine where generated speech corresponds to text, but their usefulness ultimately depends on the accuracy of the timestamps and the surrounding application logic.

Soniox also provides real-time speech-to-text services, which the company prices at roughly $0.10 to $0.12 per hour depending on the configuration, with features such as translation and diarization available. 

This means developers can potentially obtain speech recognition and speech generation from the same provider, although using a single provider is not necessarily an advantage for every application.

The company also states that it does not sell customer content or personal information and does not use customer content to train, evaluate, benchmark, or improve its models. 

According to Soniox, real-time API requests are processed transiently, while audio, transcripts, prompts, and other customer content are not stored unless storage is explicitly requested, required for a particular feature, or otherwise agreed upon.

Operational logging is limited to technical metadata used for purposes such as reliability, billing, debugging, security, abuse prevention, and service operation. Soniox says these logs do not contain raw audio, transcripts, prompts, model inputs, or model outputs. 

Customer content is encrypted during transmission and when stored by supported services, while regional processing and storage are available for projects configured for data residency.

As with any newly released speech model, comparisons based solely on company specifications or early user reactions should be treated cautiously. 

Large-scale independent testing is still needed to establish how TTS v2 performs against competing systems in areas such as latency, pronunciation, voice cloning, emotional expression, multilingual speech, and reliability.

For developers evaluating the model, the more useful comparison will ultimately come from testing it against their own workloads. Time-to-first-audio, sustained streaming latency, cloning quality, pronunciation of names and structured information, language switching, and total cost under expected traffic are likely to matter more than headline specifications alone.

TTS v2 therefore enters an already competitive market rather than fundamentally changing it. Its combination of features gives developers another system to evaluate, but whether it offers a meaningful advantage over established alternatives will depend on independent testing and the requirements of each application.

Published