What is Amazon Polly?
Amazon Polly is an AWS cloud service that converts text into natural-sounding speech. The service is suitable for applications, accessibility features, and automated content creation, and supports numerous languages.
Polly offers four voice types: classic standard voices, neural TTS voices with more natural emphasis through deep learning, long-form voices for consistent narration quality on longer texts, and generative voices, which currently provide the most expressive and human-like speech quality. The simple API enables quick integration into your own applications.
Core Features
- Four voice types: Standard, neural, long-form, and generative voices for different quality and use-case requirements
- Many languages: Including German, English, French, Spanish, and more
- SSML support: Fine control over pronunciation, pauses, and emphasis
- Speech Marks: Timing information for lip-sync and text highlighting
- Lexicons: Custom pronunciation dictionaries
- Caching allowed: Generated speech output can be cached and replayed at no additional cost
Typical Use Cases
Voice Assistants: Speech output for chatbots, IVR systems, and smart home devices. Neural and generative voices enable more natural conversations.
Accessibility: Reading web content, documents, and apps aloud for visually impaired users as an audio alternative to text.
Content Creation: Audio versions of articles, e-learning content, and podcasts. Automated production saves time and cost compared to manual voice-over.
Benefits
- Multiple voice types for different quality and budget requirements
- Usage-based billing per character
- Simple REST API for quick integration
- Support for German voices
Integration with innFactory
As an AWS Reseller, innFactory supports you with Amazon Polly: we help you choose the right voice type, integrate it into your applications, optimize speech quality with SSML, and combine it with other AWS services like Lex and Connect.
Typical Use Cases
Frequently Asked Questions
What is Amazon Polly?
Amazon Polly is an AWS text-to-speech service that converts text into natural-sounding speech. It offers voices in many languages across four technology tiers: standard, neural, long-form, and generative voices.
What is the difference between the voice types?
Standard voices use classic speech synthesis, neural TTS sounds more natural thanks to deep learning, long-form voices are optimized for long, consistent narration, and generative voices currently deliver the most expressive, human-like speech quality.
What does Amazon Polly cost?
Polly is billed on a pay-as-you-go basis per synthesized character, with pricing varying by voice type (standard, neural, long-form, generative). Generated audio can be cached and replayed at no additional cost. Current quotas and prices are listed on the official pricing page.
How can I customize pronunciation?
SSML tags enable control over pauses, emphasis, pronunciation, and speaking rate. Lexicons store custom pronunciation dictionaries.
Which output formats are supported?
Amazon Polly delivers audio in formats such as MP3, OGG Vorbis, and PCM, as well as JSON with Speech Marks for timing information, for example for lip-sync or text highlighting.
Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of AWS (official documentation). This page does not represent an offer by AWS.