What is Cloud Data Fusion?
Cloud Data Fusion is Google’s fully managed, cloud-native data integration service for quickly building and managing data pipelines. The service is based on the open-source project CDAP and enables visual ETL development via drag-and-drop, allowing data analysts to build pipelines without programming.
Core Features
- Visual pipeline designer: Drag-and-drop interface for ETL workflows
- Plugin Hub: Pre-built plugins for sources, transformations, aggregations, and sinks, extensible with custom plugins
- Data lineage: Tracking of data flows across pipelines
- Pipeline templates: Reusable templates for common integration patterns
- Dataproc integration: On-demand cluster provisioning for pipeline execution
Common Use Cases
Data Warehouse Loading
Load data from operational systems, SaaS applications, and files into BigQuery. Transformations happen visually, without requiring SQL knowledge in most cases.
Hybrid Integration
Cloud Data Fusion connects on-premises databases with cloud data lakes, for example via secure connections over VPN or Interconnect.
Data Migration
During cloud migrations, Data Fusion supports the initial data export and ongoing synchronization until cutover.
Benefits
- Little to no programming required for many pipelines
- Visual debugging and monitoring
- On-demand infrastructure with no permanently running clusters
- Enterprise security with CMEK and VPC-SC in the Enterprise edition
Integration with innFactory
As a certified Google Cloud Partner, innFactory supports you with Cloud Data Fusion: pipeline design, custom plugin development, migration from existing ETL tools, and performance optimization.
Available Tiers & Options
Developer
- Low-cost entry point
- For test and development environments
- Not intended for production
Basic
- Lower cost
- Simple pipelines
- Limited features
Enterprise
- Full feature set
- Advanced security
- Customer-managed encryption
- Higher cost
Typical Use Cases
Technical Specifications
Frequently Asked Questions
What is Cloud Data Fusion?
Cloud Data Fusion is a fully managed, cloud-native data integration service from Google Cloud for quickly building and managing data pipelines. It is based on the open-source project CDAP.
Which data sources does Cloud Data Fusion support?
Through its built-in Hub, Cloud Data Fusion offers a wide range of pre-built plugins for sources, transformations, and sinks, including databases, SaaS applications, cloud storage, and on-premises systems.
What's the difference between the editions?
The Developer edition is suited for testing and development, Basic covers simple batch pipelines, and Enterprise additionally offers advanced security, Customer-Managed Encryption Keys, VPC-SC support, and streaming pipelines.
How does Cloud Data Fusion scale?
Cloud Data Fusion uses Dataproc clusters for pipeline execution, which are provisioned on demand and shut down again once processing is complete.
Can I develop custom plugins?
Yes, Cloud Data Fusion supports custom plugins. The Hub also lets you add pre-built and community-provided plugins.
Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Google Cloud (official documentation). This page does not represent an offer by Google Cloud.
