Blog
Why one voice engine can't read all 11 Indian languages
Paithra's sample calls run on two different text-to-speech engines, not one: ElevenLabs Flash v2.5 for Hindi, English and Bengali, and Cartesia Sonic for the other eight languages. That split exists because Cartesia is the only engine of the two that reads all eleven, but ElevenLabs is the stronger voice where it does cover a language.
The split, as it actually runs
Every sample call on the homepage is generated by one of these two engines, chosen per language, not per account:
- ElevenLabs Flash v2.5: Hindi, Hindi and English, Bengali, English.
- Cartesia Sonic: Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia.
Why not run everything on one engine
ElevenLabs does not yet cover Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi or Odia at the quality bar a live sales call needs. Cartesia Sonic is the one model of the two that reads all eleven — which is why it is the fallback for the eight ElevenLabs does not reach — but on Hindi, English and Bengali specifically, ElevenLabs is the stronger-sounding voice, so that is what runs there instead.
Why this is a platform decision, not a limitation
Every component of the call — speech-to-text, the language model, and text-to-speech — is swappable per tenant without a deploy. The two-engine split for voice synthesis is the same principle applied to language coverage: pick the best available engine for each language rather than accept the lowest common denominator of a single vendor. Deepgram nova-3 covers most of the speech-to-text side the same way, with ElevenLabs Scribe filling in for Malayalam and Odia, which Deepgram does not yet transcribe.