OpenAI Realtime — direct voice-to-voice dialogue with context handling and tool calling. Suitable when natural responses and the ability to interrupt the assistant are important. CRM integrations are performed through separately authorized actions.
ElevenLabs Flash — fast text-to-speech synthesis. The provider reports approximately 75 ms model latency; for languages other than English, we consider Flash v2.5. This describes synthesis performance, not a guaranteed pause in a completed call.
Cartesia Sonic — streaming speech synthesis. For Sonic 3.5, the provider reports approximately 90 ms to the first audio byte. We test the specific voice, language, and speed on the phone line using the client's scenarios.
Yandex SpeechKit — another synthesis option. API v3 accepts text in chunks within a session, allowing the response to be voiced as it is generated. The dialogue model, knowledge base, and telephony are connected separately.
We begin by comparing two options: a direct voice model and the chain “speech recognition → language model → streaming synthesis.” We choose based on quality in your dialogues, total latency, cost per completed interaction, and data-processing terms. Service capabilities were checked against the documentation on September 17, 2026; regional availability and connection terms are verified before the project.