Skip to main content
Cloud / Google Cloud / Products / Speech-to-Text - Speech Recognition with Chirp Models

Speech-to-Text - Speech Recognition with Chirp Models

Speech-to-Text converts spoken language to text using Chirp models and support for numerous languages. EU regions available.

AI/ML
Pricing Model Pay-per-use (per audio minute or second, depending on model)
Availability Global with EU regions
Data Sovereignty EU regions available
Reliability SLA as published by the provider SLA

Speech-to-Text converts spoken language to text. The current API generation (v2) uses the Chirp model family, offering high recognition accuracy and support for real-time streaming.

What is Google Cloud Speech-to-Text?

Speech-to-Text is a fully managed AI service for automatic speech recognition (ASR). The service converts audio to text and supports a wide range of languages and variants. With the Speech-to-Text API v2, Google overhauled the model architecture: instead of the former Standard, Enhanced, Phone Call, Video, and Medical models, the current offering consists of the Chirp models (Chirp 2, Chirp 3) plus a specialized telephony model for phone audio. Chirp is based on large, universal speech models and provides automatic language detection, diarization, and multilingual processing.

The service offers various processing modes: synchronous recognition for short audio clips, asynchronous processing for longer files, and real-time streaming for live audio. Streaming recognition delivers low-latency results, ideal for voice assistants, live subtitles, or voice commands. Batch processing is suited for transcribing large audio archives. The previous API version (v1) remains available, but v2 with the current Chirp models is recommended for new projects.

Core Features

  • Chirp model family: Current generation of large speech models for high accuracy and multilingual recognition
  • Telephony model: Specialized for phone audio (typically 8 kHz), suited for call centers and IVR systems
  • Automatic punctuation and diarization: Punctuation recognition and distinguishing multiple speakers
  • Real-time streaming and batch processing: Flexible processing depending on the use case
  • Word-level timestamps: For precise synchronization, e.g. for subtitles

Typical Use Cases

Call Center Transcription

A customer service center transcribes calls using the telephony model, specifically optimized for phone audio. Transcripts are automatically analyzed for quality assurance, sentiment analysis, and compliance checking.

Meeting Transcription

A company automatically transcribes internal meetings. Diarization identifies individual speakers, and transcripts are archived and made searchable.

Voice Assistants and Chatbots

A platform integrates voice commands via streaming recognition that processes user speech in real time. Customers can search for products, place orders, and ask questions by voice.

Accessibility and Subtitles

A media company creates automatic subtitles for videos. Word-level timestamps enable precise subtitle synchronization, including for live events.

Benefits

  • Current Chirp models with high recognition accuracy based on large speech models
  • Flexible processing: streaming and batch
  • Specialized models for telephony use cases
  • Integration with the Google Cloud AI ecosystem

Integration with innFactory

As a certified Google Cloud Partner, innFactory supports you with Speech-to-Text: API integration, migration to the current API v2 with Chirp models, streaming implementation, and recognition accuracy optimization.

Contact us for a consultation on Speech-to-Text and Google Cloud AI.

Available Tiers & Options

Telephony

Strengths
  • Optimized for phone audio (8 kHz)
  • Suited for call centers and IVR
Considerations
  • Specialized for telephony use cases

Typical Use Cases

Call center transcription
Voice commands
Meeting transcription
Accessibility
Voice search

Technical Specifications

API Speech-to-Text API v2 (v1 still available), RESTful API, gRPC, client libraries
Features Automatic punctuation, speaker diarization, word-level timestamps
Integration Native Google Cloud integration
Models Chirp 2, Chirp 3, telephony (successors to the former Standard/Enhanced/Phone Call/Medical models)
Security Encryption at rest and in transit
Streaming Real-time streaming and batch processing

Frequently Asked Questions

What is Google Cloud Speech-to-Text?

Speech-to-Text is an AI service that converts spoken language to text. The current Speech-to-Text API v2 uses the Chirp model family, based on large speech models, and offers automatic punctuation, speaker recognition, and processing of both recorded files and real-time audio.

What models are available today?

The current model generation is called Chirp (Chirp 2 and Chirp 3) and offers high accuracy along with multilingual recognition. A specialized telephony model is available for phone audio. The former Standard, Enhanced, Phone Call, Video, and Medical model names were superseded by these new models as part of the API v2 evolution.

What languages are supported?

Speech-to-Text supports a large number of languages and variants; the exact, current list is maintained in the official supported-languages documentation since coverage keeps expanding.

Does Speech-to-Text support real-time streaming?

Yes, Speech-to-Text offers real-time streaming transcription with low latency. Audio is processed continuously and results are returned progressively, ideal for live subtitles, voice assistants, and voice commands.

Can I customize recognition for specialized vocabulary?

Yes, Speech-to-Text provides mechanisms to adapt recognition to technical terms, product names, or industry-specific terminology, improving accuracy for specialized applications.

How is Speech-to-Text billed?

Billing is usage-based according to processed audio duration, with rates varying by the model used. Current details and any free-tier allowances are available on the official pricing page.

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Google Cloud (official documentation). This page does not represent an offer by Google Cloud.

Google Cloud Partner

innFactory is a certified Google Cloud Partner. We provide expert consulting, implementation, and managed services.

Google Cloud Partner

Ready to start with Speech-to-Text - Speech Recognition with Chirp Models?

Our certified Google Cloud experts help you with architecture, integration, and optimization.

Schedule Consultation