Which GCP Services Need a VPC? Networking Guide




















 

What Needs a VPC in Google Cloud—and What Doesn’t?

When designing a Google Cloud architecture, a common question is:

Which services need a VPC network, and which can operate without one?

Dataflow asks you to configure a network or subnetwork. BigQuery asks you to select a project. Cloud Run can deploy without any VPC configuration. Cloud Storage lets you create a bucket without selecting a network.

The explanation lies in where the service runs its resources and how applications connect to it.

Some services deploy compute resources into your VPC. Others expose Google-managed APIs without placing their underlying infrastructure in your network. A third category can operate independently but connect to your VPC when an application needs private resources.

Cloud Run belongs in that third category. It does not require a VPC for a basic deployment, but it supports private connectivity.

The Core Distinction: Resource Placement vs. Access Path

Two separate questions determine the networking design:

  1. Does this service place resources in my VPC?
  2. Does my application need a private network path to reach this service or its dependencies?

These questions can have different answers.

For example, a Dataflow worker VM runs in a VPC subnet. A Cloud Storage bucket does not. Yet that worker can access the bucket through private connectivity to Google APIs.

A private access path does not mean the destination service resides in your subnet.

Quick Comparison

Service Requires a customer VPC for ordinary use? Relationship to your VPC
Dataflow Yes, for its worker VMs Workers use a VPC network and subnet; a default network can be selected implicitly.
BigQuery No Google manages query processing and storage; clients can use private API access.
Cloud Run services and jobs No VPC connectivity is optional and supports access to private resources and controlled outbound routing.
Cloud Storage No Buckets are accessed through service endpoints; private API connectivity can be configured.

These differences reflect the services’ execution and access models.

1. Dataflow: Managed Processing with Workers in Your VPC

Dataflow manages Apache Beam pipelines, including worker provisioning and scaling. However, its worker VMs still require network connectivity.

When submitting a job, you can specify a network, a subnetwork, or both. Specifying the subnet identifies its parent VPC automatically.

There is a useful correction to the statement that Dataflow always requires you to explicitly pick a network: if you omit both settings, Dataflow attempts to use an auto mode VPC named default. If that network is absent, you must specify another network or subnet.

Why the network matters

The workers need to communicate with pipeline dependencies, which might include:

  • Cloud Storage for input files, staging files, and temporary data.
  • BigQuery for analytics output.
  • Pub/Sub for streaming messages.
  • Private databases or application endpoints.
  • External systems accessed by pipeline code.

Google manages the workers, but your network configuration still determines their connectivity. Routes, firewall rules, DNS, and access to Google APIs can affect whether a pipeline succeeds.

Workers without public IP addresses

Dataflow workers can operate without external IP addresses. In that configuration, enable Private Google Access on the subnet so the workers can reach required Google APIs.

If pipeline code also needs public internet destinations—for example, an external API—provide an appropriate outbound path, such as Cloud NAT. Private Google Access provides access to Google APIs; it does not provide general internet access.

Subnet sizing also matters: autoscaling requires available IP addresses for additional workers.

Dataflow takeaway: A managed service can still depend on compute resources deployed in your VPC.

2. BigQuery: A Project and Dataset, Without a Customer Subnet

BigQuery separates analytics from infrastructure management. You work with projects, datasets, tables, and query jobs while Google manages the underlying storage and processing.

For ordinary BigQuery use, you do not select a VPC, assign a subnet, or provision query-worker VMs. You can run queries through the console, command-line tools, client libraries, or APIs.

Why a project is required

A project provides the administrative context for resources, permissions, API usage, and billing. It is not a network boundary.

Creating a BigQuery dataset in a project that also contains a VPC does not place the dataset inside that VPC.

Likewise, running a query from a VM in a private subnet does not relocate BigQuery into the subnet. The VM is a client accessing a managed service.

BigQuery can still use private connectivity and perimeter controls

An enterprise can configure private access to Google APIs and use VPC Service Controls to restrict access and supported data movement across a service perimeter.

These controls protect how BigQuery is accessed and how its data moves. They do not turn BigQuery into a subnet-hosted database.

BigQuery takeaway: The service does not require your VPC, but the clients accessing it may need a carefully designed network path.

3. Cloud Run: VPC Optional, Private Connectivity Supported

A Cloud Run service can deploy and serve requests without selecting a VPC network.

However, the explanation “Cloud Run does not need a network because it cannot use a private IP” is inaccurate.

Cloud Run supports VPC integration, including private outbound connectivity and private inbound access patterns.

Outbound connectivity: Cloud Run to private resources

Suppose a Cloud Run application needs to reach a private database or an internal application endpoint.

You can configure:

  • Direct VPC egress, which connects outbound traffic to a VPC without a connector.
  • Serverless VPC Access, which provides connectivity through a connector.

The selected configuration lets the application reach internal addresses through the VPC. You can route only private destination traffic through the network or configure all outbound traffic to use it.

Inbound connectivity: Private clients to Cloud Run

Inbound access is a separate design decision.

A Cloud Run service can be reached through private access patterns such as:

  • Private Google Access.
  • A Private Service Connect endpoint with appropriate DNS configuration.
  • An internal Application Load Balancer.

These options can provide private access, including access through an internal IP address. Direct VPC egress itself does not create an inbound endpoint for a Cloud Run service.

Cloud Run worker pools have a different networking model: Google documents direct VPC connectivity with private IP addresses and support for both ingress and egress. Avoid applying every networking assumption about services and jobs to worker pools.

Cloud Run takeaway: A VPC is optional for deployment. It becomes relevant when application dependencies, private access requirements, or outbound traffic policies require it.

4. Cloud Storage: Buckets Are Not Placed in Your VPC

Creating a Cloud Storage bucket does not require a VPC or subnet.

Cloud Storage is accessed through service endpoints rather than through a network interface assigned to each bucket. Calling it “not network based” can be misleading: applications still use network connections to read and write objects. The distinction is that the bucket is not deployed inside your customer VPC.

A publicly reachable endpoint does not mean public data

A bucket can be accessed through a publicly reachable API endpoint while its objects remain restricted by access permissions.

Cloud Storage also provides public access prevention to block public access granted through permissions for allUsers and allAuthenticatedUsers. This protects against accidental public exposure; it does not require moving the bucket into a subnet.

Private access to Cloud Storage

Applications in a VPC can use private connectivity to Google APIs. Private Service Connect can provide an internal endpoint IP for accessing supported Google services, including Cloud Storage.

The endpoint belongs to the access architecture. The bucket remains a Google-managed storage resource.

Cloud Storage takeaway: No VPC is needed to create or use a bucket, but private connectivity may be part of the architecture used to access it.

VPC and VPC Service Controls Solve Different Problems

The similar names often cause confusion.

Control Primary purpose
VPC network Provides network connectivity, subnets, routes, and network traffic controls.
IAM Determines which identities can perform actions on resources.
Private Google Access / Private Service Connect Provides connectivity to services through configured private access paths.
VPC Service Controls Adds service perimeter controls that help restrict access and reduce data exfiltration through supported services.

VPC Service Controls complements IAM. It can protect services such as BigQuery and Cloud Storage even though those services do not reside in your subnet.

It also does not replace general network egress controls: a service perimeter does not block arbitrary third-party internet APIs.

Putting the Services Together

Consider an application that processes files and exposes analytics:

  1. Files arrive in Cloud Storage.
  2. Dataflow workers process them.
  3. Results are written to BigQuery.
  4. A Cloud Run application queries the results.

In this architecture, Dataflow’s worker VMs require a VPC. The bucket and BigQuery dataset do not.

Cloud Run may operate without VPC integration if its dependencies and policies allow standard managed-service connectivity. If it also needs a private database or controlled outbound routing, VPC integration becomes part of its design.

The useful architectural question is therefore:

“Where does this workload execute, what must it reach, and which controls must govern that access?”

That question explains why Dataflow needs worker networking, why BigQuery and Cloud Storage need no customer subnet, and why Cloud Run can operate either independently or with private VPC connectivity.

More GCP Services: Which Need a VPC—and Which Don’t?

The same distinction extends beyond Dataflow, BigQuery, Cloud Run, and Cloud Storage.

Services that execute workloads on VM instances in your network require a VPC. Services accessed through Google-managed APIs generally do not require you to provide one.

However, “requires a VPC” does not always mean “requires creating a new VPC.” A service may use an existing network, a Shared VPC, or the default network.

A. Services That Run VM-Based Workloads in a VPC

Service What runs in the VPC? Example workload
Compute Engine VM instances with network interfaces attached to VPC subnets. Hosting an application server, database, or security appliance.
Google Kubernetes Engine—GKE Worker nodes that execute container workloads. Running microservices on a Kubernetes cluster.
Dataflow Worker VMs that execute pipeline processing. Transforming streaming events before writing them to BigQuery.
Dataproc clusters / Managed Service for Apache Spark clusters Cluster VMs running Spark and other supported processing components. Running distributed Spark processing against a data lake.
Batch VM instances provisioned to execute batch tasks. Running simulations, rendering jobs, or parallel file processing.
App Engine flexible environment Compute Engine VM instances running the application. Hosting a web application that needs a flexible runtime environment.

Google’s VPC documentation identifies Compute Engine, GKE, and App Engine flexible instances as VM-based resources using VPC networking. Dataflow, Spark clusters, and Batch also use networks for their worker or execution VMs.

Compute Engine: The Direct Example

A Compute Engine VM uses a network interface attached to a VPC subnet. The subnet supplies its internal addressing, while routes and firewall rules govern connectivity.

For example, an application VM in 10.10.1.0/24 can communicate with a database VM in 10.10.2.0/24, subject to the relevant network and application controls.

An external IP address is optional. A VM without an external IP still uses its VPC network.

GKE: Containers Run on Nodes with VPC Networking

GKE adds Kubernetes orchestration, but the containers still execute on worker nodes connected to a VPC.

For a VPC-native cluster, network planning includes addressing for nodes, Pods, and Services. This makes subnet and IP range planning part of the cluster architecture.

GKE Autopilot reduces node administration, but it does not remove the cluster’s dependence on VPC networking.

Example: An internal microservices platform needs sufficient address space for its nodes and Pods, plus connectivity to private application dependencies.

Dataproc Clusters: Distributed Processing on Networked VMs

Dataproc’s VM-based clusters use a network to connect their processing nodes and reach data sources.

Example: A Spark job reads files from Cloud Storage, performs distributed transformations across worker VMs, and writes results to BigQuery.

The cluster VMs require networking even though the source bucket and destination dataset are not deployed in that network. Google’s current documentation describes this cluster offering under Managed Service for Apache Spark.

Batch: Temporary Compute Still Requires Networking

Batch provisions VMs to run tasks and manages their execution lifecycle.

Those VMs need a VPC and subnet. If you do not specify networking options, Batch uses the default network and the subnet appropriate for the VMs’ location.

Example: A financial simulation runs hundreds of parallel tasks, reads input from Cloud Storage, and uploads the resulting files. The execution VMs use a VPC even if they exist only for the duration of the job.

App Engine Flexible: Managed Hosting on Compute Engine

App Engine flexible runs applications on Compute Engine VMs and uses VPC networking. It can also use a Shared VPC configuration.

This distinction is specific to the flexible environment; do not assume that every App Engine deployment has the same infrastructure model.

B. Services That Do Not Require Your VPC for Ordinary Use

For the following services, you create service resources and access them through managed endpoints. You do not provision their underlying servers into your subnet.

Service What you create or use Example
Pub/Sub Topics and subscriptions. Distributing order events to multiple consumers.
Firestore in Native mode A document database, collections, and documents. Storing application profiles and preferences.
Spanner Managed database instances and databases. Running a distributed transactional database.
Secret Manager Secrets and secret versions. Supplying credentials to an application.
Cloud KMS Key rings, keys, and key versions. Managing encryption keys and cryptographic operations.
Artifact Registry Repositories for container images and other artifacts. Storing images deployed to Cloud Run or GKE.

These services expose managed functionality without requiring customers to deploy the service’s servers into a VPC subnet. Application connectivity and security requirements can still introduce private networking into the overall architecture.

Pub/Sub: Messaging Without Customer-Managed Brokers

Creating a Pub/Sub topic does not require deploying broker VMs or selecting a subnet.

Publishers and subscribers communicate with the managed messaging service through its APIs.

Example: A Cloud Run application publishes an order event, and a Dataflow pipeline consumes it. Pub/Sub does not require your VPC, while Dataflow’s processing workers do.

The networking requirements of a subscriber are separate from those of the messaging service.

Firestore: A Managed Document Database

Firestore in Native mode provides a managed document database accessed through SDKs and APIs.

Example: A web application stores user preferences in Firestore without deploying a database VM or selecting a database subnet.

The application may run on a VM, GKE, or Cloud Run. Its hosting environment determines its own networking needs; Firestore does not inherit them.

Spanner: Provisioned Database Capacity Without Subnet Placement

Spanner demonstrates that provisioning database capacity does not necessarily mean provisioning VMs into your VPC.

You configure a managed database instance and its capacity, while Google operates the underlying infrastructure.

Example: An application uses Spanner for distributed transactions without assigning the database a customer subnet. Private connectivity can be part of the access design, but it is separate from database placement.

Secret Manager and Cloud KMS: Security Services Accessed Through APIs

Neither service requires you to deploy its infrastructure into a VPC.

Secret Manager stores and serves secret values. Cloud KMS manages cryptographic keys and operations.

Example: A Cloud Run application retrieves a credential from Secret Manager, while another workload uses Cloud KMS for encryption. These API interactions do not, by themselves, require a customer VPC.

Artifact Registry: The Repository and Its Consumers Have Different Requirements

Artifact Registry stores software artifacts without requiring a repository subnet.

Example: A container image stored in Artifact Registry can be deployed to Cloud Run without VPC integration, or to a GKE cluster that requires VPC networking.

The repository’s networking model is separate from that of the platform running the image.

An Additional Distinction: Managed Services with Private Network Endpoints

Some services need private connectivity without placing their underlying VMs directly in your customer subnet.

Cloud SQL with private IP is a useful example. Its private connectivity can use private services access or Private Service Connect. With private services access, the database runs in a service producer network connected to your VPC.

That differs from a Compute Engine database VM that you deploy directly into your subnet.

A useful classification therefore has three categories:

Category Architectural meaning Examples
VM-based execution in your VPC Workload compute uses your VPC networking. Compute Engine, GKE, Dataflow, Dataproc clusters, Batch, App Engine flexible.
Managed service without a required customer VPC Service resources are accessed through managed endpoints. BigQuery, Cloud Storage, Pub/Sub, Firestore, Spanner, Secret Manager, Cloud KMS, Artifact Registry.
Optional or configuration-dependent private connectivity The chosen deployment or access pattern introduces a VPC dependency. Cloud Run with VPC integration; Cloud SQL with private connectivity.

For each service, examine both its execution infrastructure and its access requirements. That keeps VM placement, private endpoints, and managed API access clear when designing the architecture.