Skip to main content
Cloud / Google Cloud / Products / Gemini Models: Foundation Models via the Gemini API

Gemini Models: Foundation Models via the Gemini API

Gemini models: Google's foundation models with long context and multimodality via Gemini API and Agent Platform, in EU regions too.

AI/ML
Pricing Model Pay-per-use (tokens), plus batch, context caching, and provisioned throughput
Availability Global and EU endpoints (e.g. europe-west3 Frankfurt, europe-west4 Netherlands)
Data Sovereignty EU regions and EU data residency endpoints available
Reliability SLA as published by the provider SLA

The Gemini models are Google’s foundation model family for developers and enterprises that want to integrate language and multimodal models into their own applications and products. Access is available through the Gemini API or through the Gemini Enterprise Agent Platform (the former Vertex AI service), which adds enterprise governance, region control, and integration with the rest of the Google Cloud stack. This is a different access path than Gemini in Google Workspace, which addresses end-user features like Gmail summaries or Docs assistants.

The Gemini Model Family

The model family covers a range of requirements: powerful Pro variants for complex reasoning, code generation, and analytical questions, balanced Flash variants for high throughput, and especially cost-efficient Flash-Lite variants for latency-sensitive, high-volume workloads. Google continuously evolves the model family; older generations are deprecated after notice, so existing integrations should be migrated to current versions regularly.

The Gemini Enterprise Agent Platform model documentation currently lists, among others:

  • Pro: Gemini 3.1 Pro, Gemini 3 Pro Image, Gemini 2.5 Pro
  • Flash: Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini Omni 1.1 Flash, Gemini Omni Flash, Gemini 3.1 Flash Image, Gemini 3 Flash, Gemini 2.5 Flash
  • Flash-Lite: Gemini 3.5 Flash-Lite, Gemini 3.1 Flash-Lite, Gemini 3.1 Flash-Lite Image, Gemini 2.5 Flash-Lite
  • Specialized models: Gemini 3.5 Transcribe, Gemini Embedding 2, Gemini Robotics ER 2
  • Generative media: Veo 3.1, Veo 3, and Veo 2 for video, Lyria 3 and Lyria 2 for music

Partner and open models such as Claude, Llama, Mistral, DeepSeek, and Qwen are additionally available through Model Garden. Since Google adjusts the list continuously, the official model documentation is authoritative.

All current models are natively multimodal and process text, image, audio, video, and PDF within the same context window. The most capable variants offer a context window of up to 1 million tokens, so very long documents, large codebases, or extensive transcripts can be processed in a single request. Many models also support a controllable thinking or reasoning mode to balance response quality, latency, and cost.

Core Features

  • Current model family: Multiple generations and size classes (Pro, Flash, Flash-Lite), all natively multimodal via a unified API.
  • Long context and thinking mode: Up to 1 million tokens of context on the most capable models, plus a controllable reasoning mode to balance quality, latency, and cost.
  • Grounding and tools: Grounding with Google Search for current web information, grounding on your own data, function calling, and structured JSON output for production-ready integrations.
  • Customization: Supervised fine-tuning for selected models to adapt Gemini to your own data and tasks.
  • Cost optimization: Batch processing for asynchronous jobs, context caching for recurring long contexts, and provisioned throughput for predictable, reserved capacity.
  • Enterprise governance: EU endpoints and EU data residency, encryption in transit and at rest, and the commitment that customer data accessed through enterprise channels is not used to train the models.

Typical Use Cases

Text and code generation: Applications produce content, summaries, or source code and use the large context window to incorporate extensive input context.

Multimodal analysis: Models process text, image, audio, video, and PDF together, for example to analyze documents, extract structured data, or describe media.

Grounded assistants: Through grounding with Google Search or your own data sources, assistants deliver current and verifiable answers and reduce hallucinations on time-sensitive topics.

Domain-specific models: Through fine-tuning, enterprises adapt Gemini to their own terminology, formats, and tasks and operate the models in EU regions.

Benefits

  • A unified API for a current, natively multimodal model family ranging from cost-efficient (Flash-Lite) to powerful (Pro).
  • A long context window of up to 1 million tokens and a controllable thinking mode for demanding reasoning tasks.
  • EU endpoints and EU data residency, plus the commitment that customer data accessed through enterprise channels is not used to train the models.
  • Cost control through batch, context caching, and provisioned throughput, plus tight integration with the rest of the Google Cloud stack.

Integration with innFactory

As a certified Google Cloud Partner, innFactory supports the integration of Gemini models into your applications: API integration, prompt engineering, grounding and fine-tuning projects, model selection and migration, and architecture consulting for production-ready, EU-compliant deployments.

Contact us for technical consulting on the Gemini models.

Typical Use Cases

Text and code generation
Multimodal analysis (text, image, video, audio, PDF)
Grounding with Google Search and your own data
Fine-tuning on custom data

Frequently Asked Questions

What are the Gemini models and how do I access them?

The Gemini models are Google's foundation model family for text, image, audio, video, and code. Programmatic access is available through the Gemini API and through the Gemini Enterprise Agent Platform (formerly Vertex AI), which adds enterprise governance, region control, and a Model Garden with third-party models. This is a different access path than Gemini in Google Workspace, which addresses end-user features like Gmail summaries.

Which Gemini models are currently available?

The Gemini Enterprise Agent Platform model documentation currently lists Gemini 3.1 Pro and Gemini 3 Pro Image, the Flash models Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash, and Gemini Omni 1.1 Flash, the Flash-Lite models Gemini 3.5 Flash-Lite and Gemini 3.1 Flash-Lite, and the specialized models Gemini 3.5 Transcribe, Gemini Embedding 2, and Gemini Robotics ER 2, among others. Older generations such as Gemini 2.5 Pro and Gemini 2.5 Flash remain available for now and are deprecated after notice. The official model documentation lists the current models and retirement dates.

Is my data used to train the Gemini models?

Under the Google Cloud privacy commitment, customer data accessed through the enterprise channels is not used by default to train the foundation models: not prompts, not responses, and not adapter model training data. The foundation models remain frozen and only process input to produce the requested output.

Are Gemini models available in the EU?

Yes. Models can run via EU endpoints, including regions such as europe-west3 (Frankfurt) and europe-west4 (Netherlands). For strict requirements, EU data residency endpoints keep processing and storage within the EU geography. Model availability differs by region, so it is worth checking the regional availability matrix in advance.

How are the Gemini models billed?

Billing is token-based (pay-per-use), separated into input and output tokens and depending on model and modality. For cost optimization there is batch processing for asynchronous jobs, context caching for recurring long contexts, and provisioned throughput for predictable, reserved capacity. The official pricing page lists current prices.

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Google Cloud (official documentation). This page does not represent an offer by Google Cloud.

Google Cloud Partner

innFactory is a certified Google Cloud Partner. We provide expert consulting, implementation, and managed services.

Google Cloud Partner

Similar Products from Other Clouds

Other cloud providers offer comparable services in this category. As a multi-cloud partner, we help you choose the right solution.

AWS

Amazon Nova Forge: Build Your Own Frontier Models on Nova

Amazon Nova Forge provides access to Nova checkpoints across all training phases and blends your proprietary data with …

Pricing Subscription; pricing information is …
SLA Per provider / see official documentation
Compare →
AWS

Amazon Nova Multimodal Embeddings: Unified Vectors

Amazon Nova Multimodal Embeddings creates vectors for text, image, video, and audio in a single model — for agentic RAG …

Pricing Usage-based through Amazon Bedrock; per …
SLA SLA per provider (see the official Amazon Bedrock SLA page)
Compare →
AWS

Amazon SageMaker Ground Truth: Label Training Data

Amazon SageMaker Ground Truth builds labeled training datasets using human workforces. Closed to new customers since …

Pricing Rates per provider (see official pricing …
SLA SLA per provider (see official SLA page)
Compare →
AWS

AWS Context: Knowledge Graph for AI Agents

AWS Context maps the relationships across your data into a knowledge graph and provides governed agentic search for AI …

Pricing No official pricing page published yet
SLA Per provider / see official documentation
Compare →
AWS

AWS Deep Learning AMIs - Preconfigured ML Images

AWS Deep Learning AMIs (DLAMI) are preconfigured EC2 images with NVIDIA CUDA, cuDNN and recent deep learning frameworks …

Pricing No charge for the AMIs; you pay for the …
SLA As stated by the provider; see official documentation
Compare →
AWS

AWS Deep Learning Containers - Prebuilt ML Images

AWS Deep Learning Containers are prepackaged, fully tested Docker images for PyTorch, TensorFlow and Apache MXNet on …

Pricing You pay for the AWS services you use; as …
SLA As stated by the provider; see official documentation
Compare →

99 comparable products found across other clouds.

Ready to start with Gemini Models: Foundation Models via the Gemini API?

Our certified Google Cloud experts help you with architecture, integration, and optimization.

Schedule Consultation