What is Dataproc Metastore?
Dataproc Metastore is a fully managed, highly available Hive Metastore service from Google Cloud. The service acts as a central metadata repository for data lake workloads, storing table definitions, schemas, and partition information that various compute engines can access.
Without a managed metastore, clusters must run their own metadata databases, which are lost when the cluster is deleted. Dataproc Metastore decouples metadata from compute, enabling ephemeral clusters without data loss. It connects to both the Managed Service for Apache Spark (formerly Dataproc, including its serverless option) and self-managed clusters.
Core Features
- Managed Hive Metastore: Fully managed service without infrastructure management
- Multi-engine access: Shared metadata for Spark, Presto, Hive, and other engines
- High availability: Automatic replication and failover, with multi-zone options depending on service tier
- IAM integration: Fine-grained access control on metadata
Typical Use Cases
Data Lake Architecture
In data lake architectures on Cloud Storage, Dataproc Metastore serves as the central schema repository. Different teams and tools access the same table definitions.
Ephemeral Cluster Workflows
Data engineering teams create clusters for individual jobs via the Managed Service for Apache Spark and delete them afterwards. The central metastore preserves table definitions independently of the cluster lifecycle.
Benefits
- No metastore infrastructure to manage
- Metadata survives the cluster lifecycle
- Consistent schema definitions across teams and tools
- Integration with BigQuery for lakehouse architectures
Integration with innFactory
As a certified Google Cloud partner, innFactory supports you with Dataproc Metastore: data lake architecture, metadata management, and lakehouse strategies.
Typical Use Cases
Frequently Asked Questions
What is Dataproc Metastore?
Dataproc Metastore is a fully managed Hive Metastore service from Google Cloud. It stores and manages metadata for data lake workloads so that Spark, Presto, and Hive can access shared table definitions.
Why do I need a central metastore?
Without a central metastore, each cluster must manage its own metadata. A central metastore allows multiple clusters and services to access the same table definitions, improving consistency and reusability.
Which tools work with Dataproc Metastore?
Dataproc Metastore is compatible with Apache Spark, Presto, Apache Hive, the Managed Service for Apache Spark (formerly Dataproc, including its serverless option), and other tools that use the Hive Metastore interface.
What service tiers are available for Dataproc Metastore?
Google Cloud offers, among others, Standard, Enterprise, and Enterprise Plus tiers, which differ in scalability, fault tolerance, and multi-zone high availability. Google Cloud's current pricing page is authoritative for exact prices per tier.
How much does Dataproc Metastore cost?
Billing is hourly per metastore service, with the exact price depending on the chosen service tier. Google Cloud's current pricing page should be consulted for a binding calculation.
Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Google Cloud (official documentation). This page does not represent an offer by Google Cloud.
