GCP Governance, Billing and VMware Migration Guide

















A Practical Guide to GCP Governance, Billing, VMware Migration, and Disaster Recovery

A successful Google Cloud deployment requires more than moving workloads into the cloud. Organizations must also establish the right administrative roles, billing controls, network segmentation, migration strategy, backup architecture, and disaster-recovery model.

This guide covers the key areas that should be reviewed when assessing a Google Cloud environment, particularly one that includes Google Cloud VMware Engine (GCVE), Compute Engine, BigQuery billing exports, and GKE Enterprise.

1. Understanding GCP Administrative and Billing Roles

One of the most confusing aspects of Google Cloud is that resource administration and billing administration are separate responsibilities.

Having broad access to projects, folders, or the organization does not automatically provide full access to Cloud Billing accounts.

Owner

The basic Owner role provides extensive control over resources within the scope where the role is assigned. A project Owner, for example, can manage most resources and IAM policies in that project.

However, ownership of a project does not automatically make the user an administrator of the Cloud Billing account linked to that project.

Organization Administrator

The Organization Administrator role provides extensive control over the Google Cloud resource hierarchy. This can include managing:

  • Organization-level IAM policies
  • Folders and projects
  • Organization policies
  • Resource hierarchy permissions
  • Delegated administrative access

However, the Organization Administrator role does not, by itself, provide complete administrative access to every Cloud Billing account.

Billing Account Administrator

The Billing Account Administrator role—roles/billing.admin—is specifically designed to manage Cloud Billing accounts.

Depending on where the role is granted, a Billing Account Administrator can:

  • View billing accounts
  • Review costs and pricing
  • Configure billing exports
  • Create and manage budgets
  • Link or unlink projects
  • Manage billing-related IAM permissions
  • Update certain billing and payment settings

The user who creates a Cloud Billing account is normally assigned the Billing Account Administrator role for that account by default. Other administrators must be granted access separately. Google Cloud explains the separation between resource permissions, billing permissions, and Google Payments permissions in its billing access documentation.

Viewing Contracted or Custom Pricing

The billing.accounts.getPricing permission is included in roles such as:

  • Billing Account Administrator
  • Billing Account Viewer

The Billing Account Viewer role is appropriate for finance, audit, and FinOps personnel who need to review billing data and pricing without changing billing-account settings.

It is also important to distinguish Cloud Billing permissions from Google Payments profile permissions. A user may be able to analyze cloud costs but still be unable to modify payment methods or manage every aspect of the associated payments profile.

2. Recommended Resource Hierarchy

A well-designed Google Cloud environment normally follows this hierarchy:

Organization
└── Folders
    └── Projects
        └── Cloud resources

Folders provide a useful policy and administrative boundary. They can represent:

  • Production and nonproduction environments
  • Business units
  • Geographic regions
  • Regulated and nonregulated workloads
  • Shared infrastructure
  • Security and networking services

A typical structure might include separate folders for:

  • Production
  • Development and testing
  • Shared networking
  • Security and logging
  • Data and analytics
  • Sandbox workloads

IAM policies and organization constraints can then be applied at the appropriate level and inherited by the resources below it.

The Organization Policy Service should be used to enforce enterprise guardrails such as:

  • Permitted deployment regions
  • Restrictions on external IP addresses
  • Domain-restricted sharing
  • Allowed services
  • Encryption requirements
  • Service-account controls
  • Restrictions on public resource access

A service catalog can also help provide approved infrastructure patterns, templates, and applications to internal teams.

3. Budgets and Cost Alerts

Budgets should be created for billing accounts, folders, projects, or other meaningful cost boundaries.

A budget does not normally stop spending by itself. Its primary purpose is to monitor cost and trigger notifications when actual or forecasted spending reaches defined thresholds.

Common thresholds include:

  • 50% of budget
  • 80% of budget
  • 90% of budget
  • 100% of budget
  • Forecasted overspend

Alert recipients should include more than the billing-account administrators. Depending on the organization, notifications may need to reach:

  • Finance and FinOps teams
  • Project owners
  • Application owners
  • Engineering managers
  • Cloud operations
  • Security teams

Pub/Sub notifications can also be used to trigger automated workflows. For example, an organization could open a service ticket, send a collaboration alert, or initiate a review when forecasted spending exceeds a threshold.

4. Exporting Cloud Billing Data to BigQuery

Exporting billing data to BigQuery provides much more flexibility than relying exclusively on the Cloud Billing console. It enables organizations to build custom reports, identify cost anomalies, allocate shared costs, analyze service consumption, and create FinOps dashboards.

Google Cloud supports several billing-data exports, including:

  • FOCUS usage-cost export
  • Standard usage-cost export
  • Detailed usage-cost export
  • Pricing export
  • Committed-use-discount metadata export

Data Availability and Retroactive Exports

Dataset location affects how much historical usage-cost data is initially exported.

When FOCUS, standard, or detailed usage-cost export is enabled for the first time:

  • A dataset in the US or EU multi-region can receive data retroactively from the beginning of the previous month.
  • A dataset in a supported single region receives data beginning on the date the export is enabled.
  • Pricing-export data is not retroactive.

The initial multi-region backfill can take several days. If an export is disabled or redirected to another dataset, earlier records are not automatically copied into the new dataset. Google documents the current availability behavior for each export type here.

Encryption

Current Google Cloud documentation supports using customer-managed encryption keys with Cloud Billing export datasets, provided CMEK is configured correctly at the dataset or table level.

This is an important change from older guidance that treated CMEK-protected billing datasets as unsupported.

GKE Cost Allocation

Detailed billing exports include resource-level information for Compute Engine and several other services.

For meaningful GKE cluster, namespace, and workload-level cost analysis, however, GKE cost allocation must also be enabled. Without it, the billing export may not contain the detail needed to allocate Kubernetes costs accurately.

Tags and Labels

Resource-level tags can be included in standard and detailed billing exports for supported resources. Changes can take approximately an hour to propagate.

Labels should still be used consistently for operational and cost reporting. Common labels include:

  • environment
  • application
  • business_unit
  • cost_center
  • owner
  • data_classification

A mandatory labeling and tagging standard is one of the most effective foundations for cloud cost allocation.

5. BigQuery Access for Billing Analysts

Billing analysts generally need two different types of BigQuery access:

  • Permission to read the billing dataset
  • Permission to run query jobs

A common least-privilege model is:

  • roles/bigquery.dataViewer on the billing dataset
  • roles/bigquery.jobUser on the project from which queries are executed

The broader roles/bigquery.user role may be appropriate when analysts also need to create datasets or perform additional BigQuery operations, but it should not be assigned automatically if the narrower Job User role is sufficient.

6. Example Billing Query by Service

The following query summarizes consumption, gross cost, credits, and net cost by service and SKU:

SELECT
  service.description AS service,
  sku.description AS sku,
  usage.pricing_unit,
  ROUND(SUM(usage.amount_in_pricing_units), 3) AS usage_quantity,
  ROUND(SUM(cost), 2) AS gross_cost,
  ROUND(
    SUM(
      IFNULL(
        (
          SELECT SUM(credit.amount)
          FROM UNNEST(credits) AS credit
        ),
        0
      )
    ),
    2
  ) AS credits,
  ROUND(
    SUM(cost) +
    SUM(
      IFNULL(
        (
          SELECT SUM(credit.amount)
          FROM UNNEST(credits) AS credit
        ),
        0
      )
    ),
    2
  ) AS net_cost
FROM
  `PROJECT_ID.DATASET_ID.gcp_billing_export_resource_v1_BILLING_ACCOUNT_ID`
WHERE
  invoice.month = 'YYYYMM'
GROUP BY
  service,
  sku,
  usage.pricing_unit
ORDER BY
  net_cost DESC;

Replace the project, dataset, table, and invoice-month values with those from your environment.

This query reports cost and consumption by service. If the goal is to analyze published or contracted rates independently of consumption, query the separate pricing-export table.

7. Network Architecture Assessment

Before migrating workloads, document the current and target network architecture.

Important questions include:

  • Are production and nonproduction environments properly separated?
  • Are regulated workloads isolated from general-purpose workloads?
  • Is Shared VPC being used?
  • Which projects own the shared network?
  • How are application teams granted access to subnets?
  • How does traffic flow between on-premises systems and Google Cloud?
  • Is connectivity provided through Cloud VPN, Dedicated Interconnect, or Partner Interconnect?
  • Where are firewalls and security inspection services located?
  • How are DNS, routing, and IP address management handled?
  • Will VMware workloads retain their existing IP addresses?
  • Is east-west traffic inspected?
  • Are overlapping CIDR ranges present?

The network design must also consider whether workloads are moving to GCVE or directly to Compute Engine. These are different target environments with different connectivity, migration, and operational requirements.

8. Understanding the Workloads Before Selecting Services

The nature of the application determines which Google Cloud services should be used.

Discovery should identify whether workloads are:

  • Web applications
  • Transactional systems
  • Batch-processing systems
  • Streaming applications
  • Data warehouses
  • Kubernetes-based applications
  • Commercial off-the-shelf systems
  • Legacy applications with VMware dependencies
  • Latency-sensitive applications
  • Systems with strict licensing restrictions

Additional discovery questions should cover:

  • Operating systems and versions
  • Application dependencies
  • Database platforms
  • Storage requirements
  • Recovery-point and recovery-time objectives
  • Peak CPU and memory utilization
  • Network throughput
  • Compliance requirements
  • Maintenance windows
  • Licensing portability

A virtual machine should not be moved simply because it can be moved. Some systems belong in GCVE, some can move to Compute Engine, and others may benefit from containerization or modernization.

9. GCVE Versus Compute Engine

Google Cloud VMware Engine and Compute Engine solve different migration problems.

Google Cloud VMware Engine

GCVE is generally appropriate when an organization wants to retain VMware technologies and operating practices while moving the underlying infrastructure into Google Cloud.

It can reduce the amount of immediate application refactoring by preserving familiar VMware components and operational models.

Compute Engine

Compute Engine is a native Google Cloud virtual-machine platform. Moving a VMware workload to Compute Engine normally involves converting or importing the VM image and adapting the workload to Google Cloud networking, storage, IAM, monitoring, and backup services.

A VMware-to-GCVE migration is therefore different from a VMware-to-Compute Engine migration. The distinction also affects backup and disaster-recovery design.

10. Importing a VMDK into Compute Engine

A VMDK can be imported into Compute Engine to test whether a VMware workload is compatible with native Google Cloud virtual machines.

The process generally involves:

  1. Confirming that the operating system and image format are supported.
  2. Consolidating split VMDK files when necessary.
  3. Uploading the image to Cloud Storage.
  4. Enabling the required Compute Engine and VM Migration services.
  5. Assigning the necessary permissions to the migration service accounts.
  6. Importing the image.
  7. Creating a test VM.
  8. Validating boot behavior, drivers, networking, licensing, and application functionality.

Older procedures used gcloud compute images import. The current Google Cloud workflow uses the VM Migration image-import command:

gcloud compute migration image-imports create IMAGE_NAME \
  --source-file=gs://BUCKET_NAME/PATH/IMAGE.vmdk \
  --location=REGION_ID \
  --target-project=projects/HOST_PROJECT_ID/locations/global/targetProjects/TARGET_PROJECT

Current image-import documentation supports .vmdk and .tar.gz source files. Review the current prerequisites and syntax before running an import.

A successful image import does not guarantee that the complete application will work correctly. Always test:

  • Boot configuration
  • Network interfaces
  • VMware-specific drivers and tools
  • Static IP dependencies
  • Application services
  • Database connectivity
  • Authentication
  • Monitoring agents
  • Backup agents
  • Licensing

11. vSphere Access and Privilege Elevation in GCVE

GCVE provides access to VMware management components, including vCenter, but Google retains control of certain infrastructure-level operations.

Organizations should document:

  • Who will administer vCenter?
  • Which tasks require elevated privileges?
  • How will temporary elevation be requested and approved?
  • How will administrative activity be logged?
  • Which third-party tools require additional vCenter permissions?
  • Which operations remain the responsibility of Google?

Temporary privilege elevation may be needed when installing or managing tools that require administrative vCenter access. It should be treated as a controlled operational process, with documented approval and audit requirements.

12. Migration with VMware HCX

VMware HCX can support workload mobility between an on-premises VMware environment and GCVE.

A typical migration design may include:

  • HCX Connector in the source VMware environment
  • HCX Cloud components in GCVE
  • Service mesh configuration
  • Network extension appliances
  • Extended VLANs
  • Mobility Optimized Networking
  • Replication-assisted migration
  • Bulk migration
  • HCX vMotion
  • Planned switchover
  • Reverse-migration testing

Before production migration, perform:

  1. A small pilot migration.
  2. Application validation in GCVE.
  3. Network and DNS validation.
  4. Performance testing.
  5. Backup and restore testing.
  6. A controlled failback or reverse-migration test.
  7. A documented production cutover.

Extending a Layer 2 network can simplify migration by allowing a VM to retain its IP address. However, it also extends the failure domain and can complicate routing. It should generally be treated as a transition mechanism rather than a permanent substitute for sound cloud network design.

13. Backup Is Not the Same as Migration or Disaster Recovery

Migration, backup, high availability, and disaster recovery are related but separate capabilities.

  • Migration moves a workload to another environment.
  • Backup creates recoverable copies of data or systems.
  • High availability reduces disruption from local component failures.
  • Disaster recovery restores service after a major site or regional failure.

GCVE supports integration with several third-party data-protection and disaster-recovery platforms, including products from Veeam, NetApp, Dell, Cohesity, and Zerto. Product selection should be based on workload requirements, recovery objectives, licensing, and the organization’s existing operational model.

For each workload, define:

  • Recovery Point Objective
  • Recovery Time Objective
  • Backup frequency
  • Retention period
  • Immutable-copy requirements
  • Off-site or cross-region copy requirements
  • Application-consistent backup requirements
  • Restore-testing schedule
  • Ransomware recovery procedures

A backup solution should never be considered complete until restore operations have been tested.

14. Site Recovery Manager Planning

VMware Site Recovery Manager can orchestrate recovery between a protected VMware environment and a recovery environment.

An SRM engagement may include:

  • Deploying SRM appliances at the protected and recovery sites
  • Configuring replication
  • Creating protection groups
  • Mapping networks
  • Remapping IP addresses
  • Creating recovery plans
  • Testing recovery plans
  • Documenting failover and failback
  • Conducting knowledge-transfer sessions

The number of protected VMs, protection groups, and recovery plans should be treated as implementation scope—not as universal product limits—unless current licensing and product documentation explicitly define those limits.

Key design questions include:

  • How many VMs require protection?
  • Which applications must recover together?
  • What is the required startup sequence?
  • Are IP addresses retained or remapped?
  • How will DNS be updated?
  • Where will replicated data be stored?
  • How frequently will recovery plans be tested?
  • Is recovery occurring within one region, across regions, or between on-premises infrastructure and GCVE?

15. Stretched Private Clouds

A stretched private cloud distributes a GCVE cluster across two data zones within a region, with a witness in a third zone.

This architecture is intended to improve availability during a zonal failure.

A stretched cluster requires:

  • Equal numbers of data nodes in both data zones
  • A minimum configuration of six data nodes, arranged as 3+3
  • Nonoverlapping management and HCX network ranges
  • Sufficient network capacity and appropriate workload placement
  • Storage policies aligned with availability requirements

Google manages the witness node. A properly designed stretched cluster can survive the loss of a single zone, but it does not eliminate the need for backups or a broader disaster-recovery strategy. Google documents the current stretched-cluster architecture and node requirements here.

16. GKE Enterprise and Hybrid Kubernetes

Organizations that previously referred to “Anthos” should now evaluate the relevant capabilities under GKE Enterprise.

On-premises or hybrid Kubernetes may be justified by:

  • Data-residency requirements
  • Regulatory constraints
  • Very low-latency dependencies
  • Factory, branch, or edge workloads
  • Applications that cannot yet move to a public-cloud region

Potential benefits include:

  • Centralized policy management
  • Consistent cluster governance
  • Multi-cluster visibility
  • Fleet management
  • Service-mesh capabilities
  • Centralized configuration
  • Multi-cluster ingress

However, hybrid Kubernetes should not be adopted solely to create a “single pane of glass.” The operational complexity, skills, connectivity, lifecycle management, and support model must justify the architecture.

17. Final Assessment Checklist

Before beginning a GCP or GCVE migration, confirm the following:

Governance

  • Is the organization, folder, and project hierarchy documented?
  • Are production and nonproduction workloads separated?
  • Are organization policies and IAM guardrails defined?
  • Are administrative duties separated?

Billing

  • Are Billing Account Administrator and Viewer roles assigned correctly?
  • Can finance teams view custom pricing?
  • Are budgets and alerts configured?
  • Is billing data exported to BigQuery?
  • Are labels, tags, and cost centers consistently applied?

Networking

  • Is Shared VPC required?
  • Are CIDR ranges nonoverlapping?
  • Is hybrid connectivity sized appropriately?
  • Is DNS resolution defined in both directions?
  • Are firewall and inspection paths documented?

Migration

  • Is each workload mapped to GCVE, Compute Engine, GKE, or a managed service?
  • Have application dependencies been discovered?
  • Has a pilot migration been completed?
  • Has reverse migration or failback been tested?

Backup and Recovery

  • Are RPO and RTO requirements documented?
  • Are backups stored outside the primary failure domain?
  • Are recovery plans documented and tested?
  • Is ransomware recovery included?
  • Is cross-region recovery required?

Conclusion

Google Cloud administration, billing, VMware migration, and disaster recovery should be designed as one coordinated architecture—not as separate technical projects.

The most important lessons are:

  • Resource ownership does not automatically provide billing administration.
  • Billing data should be exported early because historical backfill is limited.
  • Network and application discovery must precede migration-tool selection.
  • GCVE and Compute Engine represent different destination architectures.
  • HCX moves workloads, while backup products and SRM address different recovery requirements.
  • A stretched private cloud improves zonal availability but does not replace backup or disaster recovery.
  • Every migration should include testing, rollback planning, cost controls, and recovery validation.

When governance, billing, networking, migration, and recovery are designed together, the result is a cloud environment that is easier to operate, more secure, and financially accountable.