Skip to main content
Cloud / Google Cloud / Products / Cloud Run worker pools - Pull-Based Background Work

Cloud Run worker pools - Pull-Based Background Work

Cloud Run worker pools: serverless resource for non-HTTP, pull-based workloads from queues like Pub/Sub and Kafka, with GPU support for AI/ML jobs.

Serverless
Pricing Model Pay-per-use, resource-based (vCPU plus memory over instance lifetime; GPU per second incl. idle)
Availability Multiple regions incl. EU (e.g. europe-west1 Belgium, europe-west4 Netherlands)
Data Sovereignty EU regions available (europe-west1, europe-west4)
Reliability Cloud Run SLA, tiered by configuration (e.g. different values for GPU instances); see official SLA page SLA

What is Cloud Run worker pools?

Cloud Run worker pools is a Cloud Run resource type for non-HTTP, pull-based background workloads. Instances are long-lived and continuously pull work from sources such as Pub/Sub pull subscriptions, Kafka and Redis task queues, and self-hosted GitHub Actions runners. Unlike Cloud Run services, worker pools have no load-balanced endpoint and no URL, and they do not scale in response to incoming HTTP requests.

Worker pools solve the problem that request-driven serverless models fit poorly with continuously running processing. Teams that consume messages from queues, run distributed AI/ML jobs, or operate CI/CD runners previously needed either self-managed VMs or Kubernetes. With Cloud Run worker pools, this background work runs on the serverless Cloud Run platform, with resource-based billing and GPU support, and without operating your own cluster.

Core features

  • Pull-based background processing: Long-lived instances continuously pull work from queues such as Pub/Sub, Kafka and Redis, without a load-balanced endpoint or URL.
  • GPU support for AI/ML: NVIDIA L4 (24 GB VRAM) and NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM), with a limit of one GPU per instance, for distributed inference and batch jobs.
  • Manual scaling and large instances: Instance count is configured manually; instances scale up to 44 vCPU and 176 GB RAM, with up to 10 containers (one main container plus up to nine sidecars).
  • Full Cloud Run integration: Environment variables, secrets, health checks, VPC egress and ingress, NFS and Cloud Storage volumes, and immutable revisions per deployment.

Typical use cases

Queue consumers: Worker pools continuously process messages from Pub/Sub pull subscriptions, Kafka topics or Redis task queues, replacing self-managed workers on VMs or in Kubernetes.

Distributed AI/ML jobs: With GPU support, inference and batch processing for AI models run serverlessly, without provisioning and operating a GPU cluster.

Self-hosted CI/CD runners: Worker pools operate self-hosted GitHub Actions runners that continuously wait for new jobs and scale as needed.

Benefits

  • Serverless model for continuously running background work without your own VMs or Kubernetes clusters
  • Resource-based billing, which Google states is around 40 percent cheaper than request-driven services or jobs for long-running work
  • GPU support for AI/ML workloads directly on the Cloud Run platform
  • EU regions available (europe-west1, europe-west4) for privacy-compliant processing

Integration with innFactory

As a certified Google Cloud Partner, innFactory supports you with the adoption and operation of this service.

Typical Use Cases

Continuously processing Pub/Sub pull subscriptions
Consumers for Kafka and Redis task queues
Distributed AI/ML inference and batch jobs with GPU
Self-hosted GitHub Actions runners

Frequently Asked Questions

What is Cloud Run worker pools?

Cloud Run worker pools is a Cloud Run resource type for non-HTTP, pull-based background workloads. Instances are long-lived and continuously pull work from queues such as Pub/Sub pull subscriptions, Kafka or Redis. Unlike Cloud Run services, worker pools have no load-balanced endpoint and no URL, and they do not scale in response to incoming requests.

When should I use Cloud Run worker pools?

Worker pools are a fit when you run continuous consumers for message queues (Pub/Sub, Kafka, Redis), execute distributed AI/ML inference or batch jobs with GPU, or operate self-hosted CI/CD runners such as GitHub Actions runners. For request-driven HTTP endpoints, Cloud Run services are the right choice.

How much does Cloud Run worker pools cost?

Billing is resource-based and pay-per-use. vCPU and memory are billed for the full lifetime of the instance, and GPU is billed per second including idle uptime. Regional Tier 1 and Tier 2 rates apply. For long-running background work, Google states this billing is around 40 percent cheaper than request-driven services or jobs. Current pricing is listed on the official Cloud Run pricing page.

Which GPUs and limits apply to worker pools?

GPU support is generally available with NVIDIA L4 (24 GB VRAM) and NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM), with a limit of one GPU per instance. GPU worker pools cannot be autoscaled. Instances can be configured with up to 44 vCPU and 176 GB RAM, with up to 10 containers per instance (one main container plus up to nine sidecars).

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of Google Cloud (official documentation). This page does not represent an offer by Google Cloud.

Google Cloud Partner

innFactory is a certified Google Cloud Partner. We provide expert consulting, implementation, and managed services.

Google Cloud Partner

Similar Products from Other Clouds

Other cloud providers offer comparable services in this category. As a multi-cloud partner, we help you choose the right solution.

AWS

Amazon EC2 - Virtual Servers

Amazon EC2 provides scalable virtual servers in the cloud with over 1000 instance types. GDPR-compliant in EU regions.

Pricing Pay-per-use (On-Demand), plus Reserved …
SLA 99.99% Monthly Uptime Percentage at the region level, 99.5% at the instance level (see official EC2 SLA page)
Compare →
AWS

Amazon EC2 Auto Scaling - Automatic Capacity Adjustment

Amazon EC2 Auto Scaling automatically adjusts EC2 capacity to demand. Optimal performance at minimum cost.

Pricing No additional charge for Auto Scaling …
SLA Covered by the Amazon EC2 SLA (see official SLA page)
Compare →
AWS

Amazon Lightsail - Simple Cloud Hosting

Amazon Lightsail offers virtual servers, containers, and databases with fixed monthly pricing for simple workloads.

Pricing Fixed monthly pricing
SLA SLA as published by the provider
Compare →
AWS

Amazon Linux 2023 - Optimized Linux Distribution for AWS

Amazon Linux 2023 is an AWS-optimized Linux distribution with long-term support, security updates, and seamless AWS …

Pricing Free (included in the price of the EC2 …
SLA No dedicated SLA; covered by the SLA of the underlying compute service (e.g. EC2)
Compare →
AWS

AWS App Runner - Container Hosting Without Infrastructure

AWS App Runner is a managed service for container-based web apps with auto-scaling; closed to new customers since April …

Pricing Pay for vCPU and memory usage, plus …
SLA SLA as published by the provider
Compare →
AWS

AWS Batch - Batch Computing in the Cloud

AWS Batch runs batch jobs automatically on optimal compute infrastructure. Serverless or with EC2/Spot instances.

Pricing No charge for Batch, pay for resources
SLA N/A
Compare →

49 comparable products found across other clouds.

Ready to start with Cloud Run worker pools - Pull-Based Background Work?

Our certified Google Cloud experts help you with architecture, integration, and optimization.

Schedule Consultation