What is the Voice Live API?
The Voice Live API is a solution for low-latency, high-quality speech-to-speech interactions for voice agents. It targets development teams building scalable, voice-driven experiences without manually orchestrating multiple components. Speech recognition, generative AI, and text-to-speech functionality are integrated into a single unified interface.
The service is fully managed, so you handle neither backend orchestration nor component integration. Developers provide audio input and receive audio output, avatar visuals, and action triggers, all with minimal latency. You also don’t deploy or manage generative AI models yourself, since the API handles the underlying infrastructure.
The API is designed for compatibility with the Azure OpenAI Realtime API. Supported real-time events mostly match those of the Azure OpenAI Realtime API. Features unique to the Voice Live API are optional and additive: you can add Azure Speech capabilities such as noise suppression, echo cancellation, and advanced end-of-turn detection without changing your existing architecture. The API is supported through WebSocket events, which suits server-to-server integration.
Core Features
- Broad locale coverage: over 140 locales for speech to text and over 600 standard voices across 150+ locales for text to speech
- Customizable input and output through phrase list, custom speech, and custom voice
- Flexible model choice from GPT-Realtime through GPT-4.1 and the GPT-5 family to Phi models
- Conversational features: noise suppression, echo cancellation, robust interruption detection, and advanced end-of-turn detection
- Avatar integration with standard or customizable avatars synchronized with audio output
- Function calling for external actions, tool use, and grounded responses using the VoiceRAG pattern
- Data zone models that keep prompts and responses within the data zone of the resource region
- Bring Your Own Model (BYOM) for models that are not pre-deployed
Typical Use Cases
Contact centers
Interactive voice bots handle customer requests, navigate product catalogs, and enable self-service without building an orchestrator of your own.
Automotive assistants
Hands-free in-car voice assistants execute commands, support navigation, and answer general inquiries.
Education
Voice-enabled learning companions and virtual tutors support interactive training and teaching.
Public services
Voice agents assist citizens with administrative queries and public service information.
Human resources
Voice-enabled tools support HR processes such as employee support, career development, and training.
Benefits
- One interface instead of orchestrating speech recognition, LLM, and speech synthesis
- Lower latency perceived by end users and lower engineering cost
- Conversational quality through noise suppression, echo cancellation, and end-of-turn detection
- Freedom of model choice along intelligence, speed, and cost
- Brand-aligned output through custom voice and avatars
- Compatibility with the Azure OpenAI Realtime API eases migration of existing applications
Integration with innFactory
As a Microsoft Solutions Partner, innFactory supports the implementation of voice agents with the Voice Live API: model selection along your latency, quality, and cost requirements, WebSocket integration, connecting tools through function calling, and building grounded responses using the VoiceRAG pattern.
We describe how we take AI applications on Azure to production in our articles CompanyGPT cloud stack on Azure and CompanyGPT with Microsoft Foundry, agents, and Bedrock. Contact us for a no-obligation consultation.
Typical Use Cases
Technical Specifications
Frequently Asked Questions
What does the Voice Live API do?
The Voice Live API enables low-latency speech-to-speech interactions for voice agents. Speech recognition, generative AI, and text to speech are combined into a single interface, so you don't orchestrate individual components. The service is fully managed: you provide audio input and receive audio output, avatar visuals, and action triggers.
Which models are available?
Microsoft lists gpt-realtime, gpt-realtime-1.5, gpt-realtime-mini, GPT-4o and GPT-4o mini, GPT-4.1 in several sizes, the GPT-5 family, and azure-realtime, among others. The models phi4-mm-realtime and phi4-mini are marked as preview in the model table. All natively supported models are fully managed, so no deployment of your own is required.
How is the Voice Live API billed?
Pricing is tiered by the generative AI model used into Voice Live pro, basic, and lite. You don't select a tier; you choose a model and the corresponding pricing applies. Custom speech, custom voice, and custom avatar are charged separately for training and hosting. Microsoft states that pricing for the Voice Live API takes effect on July 1, 2025.
How many languages and voices are supported?
Microsoft states over 140 locales for speech to text and over 600 standard voices across more than 150 locales for text to speech.
Does data stay in the EU?
For the models gpt-realtime-datazone and gpt-realtime-1.5-datazone, Microsoft states data zone standard processing applies: prompts and responses stay within the data zone associated with the resource region. For authoritative region information, see the Azure Speech service regions list.
Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Azure (official documentation). This page does not represent an offer by Azure.
