Skip to main content
Cloud / Azure / Products / Voice Live API in Foundry Tools - Voice Agents

Voice Live API in Foundry Tools - Voice Agents

Voice Live API: low-latency speech-to-speech for voice agents with 140+ locales, avatars, noise suppression, and model choice from GPT-Realtime to Phi.

ai-machine-learning
Pricing Model Tiered by the generative model used into Voice Live pro, basic, and lite; billed by tokens
Availability Regions per the Azure Speech service regions list in the official documentation
Data Sovereignty Data zone variants (gpt-realtime-datazone, gpt-realtime-1.5-datazone) keep prompts and responses within the data zone of the resource region
Reliability per provider / see official documentation SLA

What is the Voice Live API?

The Voice Live API is a solution for low-latency, high-quality speech-to-speech interactions for voice agents. It targets development teams building scalable, voice-driven experiences without manually orchestrating multiple components. Speech recognition, generative AI, and text-to-speech functionality are integrated into a single unified interface.

The service is fully managed, so you handle neither backend orchestration nor component integration. Developers provide audio input and receive audio output, avatar visuals, and action triggers, all with minimal latency. You also don’t deploy or manage generative AI models yourself, since the API handles the underlying infrastructure.

The API is designed for compatibility with the Azure OpenAI Realtime API. Supported real-time events mostly match those of the Azure OpenAI Realtime API. Features unique to the Voice Live API are optional and additive: you can add Azure Speech capabilities such as noise suppression, echo cancellation, and advanced end-of-turn detection without changing your existing architecture. The API is supported through WebSocket events, which suits server-to-server integration.

Core Features

  • Broad locale coverage: over 140 locales for speech to text and over 600 standard voices across 150+ locales for text to speech
  • Customizable input and output through phrase list, custom speech, and custom voice
  • Flexible model choice from GPT-Realtime through GPT-4.1 and the GPT-5 family to Phi models
  • Conversational features: noise suppression, echo cancellation, robust interruption detection, and advanced end-of-turn detection
  • Avatar integration with standard or customizable avatars synchronized with audio output
  • Function calling for external actions, tool use, and grounded responses using the VoiceRAG pattern
  • Data zone models that keep prompts and responses within the data zone of the resource region
  • Bring Your Own Model (BYOM) for models that are not pre-deployed

Typical Use Cases

Contact centers
Interactive voice bots handle customer requests, navigate product catalogs, and enable self-service without building an orchestrator of your own.

Automotive assistants
Hands-free in-car voice assistants execute commands, support navigation, and answer general inquiries.

Education
Voice-enabled learning companions and virtual tutors support interactive training and teaching.

Public services
Voice agents assist citizens with administrative queries and public service information.

Human resources
Voice-enabled tools support HR processes such as employee support, career development, and training.

Benefits

  • One interface instead of orchestrating speech recognition, LLM, and speech synthesis
  • Lower latency perceived by end users and lower engineering cost
  • Conversational quality through noise suppression, echo cancellation, and end-of-turn detection
  • Freedom of model choice along intelligence, speed, and cost
  • Brand-aligned output through custom voice and avatars
  • Compatibility with the Azure OpenAI Realtime API eases migration of existing applications

Integration with innFactory

As a Microsoft Solutions Partner, innFactory supports the implementation of voice agents with the Voice Live API: model selection along your latency, quality, and cost requirements, WebSocket integration, connecting tools through function calling, and building grounded responses using the VoiceRAG pattern.

We describe how we take AI applications on Azure to production in our articles CompanyGPT cloud stack on Azure and CompanyGPT with Microsoft Foundry, agents, and Bedrock. Contact us for a no-obligation consultation.

Typical Use Cases

Voice bots for contact centers, product catalog navigation, and self-service
Hands-free in-car voice assistants
Voice-enabled learning companions and virtual tutors
Public service voice agents for citizen and administrative queries
Voice-enabled tools for HR processes

Technical Specifications

0th Speech to speech through a unified interface: speech recognition, generative AI, and text to speech combined
1st Over 140 locales for speech to text and over 600 standard voices across 150+ locales for text to speech
2nd Model choice including gpt-realtime, gpt-realtime-1.5, gpt-realtime-mini, GPT-4o, GPT-4.1, the GPT-5 family, plus phi4-mm-realtime and phi4-mini (both marked preview)
3rd Data zone variants gpt-realtime-datazone and gpt-realtime-1.5-datazone
4th Conversational features: noise suppression, echo cancellation, robust interruption detection, and advanced end-of-turn detection
5th Avatar integration with standard or customizable avatars synchronized with audio output
6th Function calling for external actions, tool use, and grounded responses using the VoiceRAG pattern
7th Compatible with the Azure OpenAI Realtime API; communication over WebSocket events
8th Customization through phrase list, custom speech, custom voice, and Bring Your Own Model (BYOM)

Frequently Asked Questions

What does the Voice Live API do?

The Voice Live API enables low-latency speech-to-speech interactions for voice agents. Speech recognition, generative AI, and text to speech are combined into a single interface, so you don't orchestrate individual components. The service is fully managed: you provide audio input and receive audio output, avatar visuals, and action triggers.

Which models are available?

Microsoft lists gpt-realtime, gpt-realtime-1.5, gpt-realtime-mini, GPT-4o and GPT-4o mini, GPT-4.1 in several sizes, the GPT-5 family, and azure-realtime, among others. The models phi4-mm-realtime and phi4-mini are marked as preview in the model table. All natively supported models are fully managed, so no deployment of your own is required.

How is the Voice Live API billed?

Pricing is tiered by the generative AI model used into Voice Live pro, basic, and lite. You don't select a tier; you choose a model and the corresponding pricing applies. Custom speech, custom voice, and custom avatar are charged separately for training and hosting. Microsoft states that pricing for the Voice Live API takes effect on July 1, 2025.

How many languages and voices are supported?

Microsoft states over 140 locales for speech to text and over 600 standard voices across more than 150 locales for text to speech.

Does data stay in the EU?

For the models gpt-realtime-datazone and gpt-realtime-1.5-datazone, Microsoft states data zone standard processing applies: prompts and responses stay within the data zone associated with the resource region. For authoritative region information, see the Azure Speech service regions list.

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Azure (official documentation). This page does not represent an offer by Azure.

Microsoft Solutions Partner

innFactory is a Microsoft Solutions Partner. We provide expert consulting, implementation, and managed services for Azure.

Microsoft Solutions Partner Microsoft Data & AI

Ready to start with Voice Live API in Foundry Tools - Voice Agents?

Our certified Azure experts help you with architecture, integration, and optimization.

Schedule Consultation