
Google Cloud: A Practical Guide to Cloud Services (Part 1)
Google Cloud Fundamentals & Compute & Virtual Machines
Google Cloud is Google’s cloud computing platform, offering services to build, deploy, secure, and manage applications and data. It brings together computing, storage, networking, databases, analytics, artificial intelligence, and developer tools to support everything from small websites to large enterprise systems.
Imagine building an online application: it needs computing power to run, databases to store transactions, networking to connect services, security to protect customer information, and monitoring to track performance. Google Cloud provides services for each of these needs, helping teams design systems that can grow with demand.
In this documentation series, we’ll explore Google Cloud services one by one through simple explanations, practical examples, diagrams, hands-on steps, and interview point of view. You’ll learn what each service does, when to use it, and how it fits into a real-world cloud architecture.
Here are the main Google Cloud services, grouped for documentation series:
| Category | Services to cover |
| Cloud basics | Regions and Zones, Projects, Resource Hierarchy, Billing |
| Compute | Compute Engine, Instance Templates, Managed Instance Groups, Cloud Run, App Engine |
| Containers | Google Kubernetes Engine (GKE), Artifact Registry |
| Networking | Virtual Private Cloud (VPC), Cloud Load Balancing, Cloud DNS, Cloud NAT, Cloud VPN, Cloud Interconnect |
| Storage | Cloud Storage, Persistent Disk, Hyperdisk, Filestore |
| Databases | Cloud SQL, AlloyDB for PostgreSQL, Spanner, Firestore, Bigtable, Memorystore |
| Security and identity | Identity and Access Management (IAM), Secret Manager, Cloud Key Management Service (Cloud KMS), Security Command Center, Cloud Armor, Identity-Aware Proxy (IAP) |
| Monitoring | Cloud Monitoring, Cloud Logging, Cloud Trace |
| DevOps and automation | Cloud Build, Cloud Deploy, Infrastructure Manager, Cloud Scheduler |
| Messaging and integration | Pub/Sub, Cloud Tasks, Workflows, Eventarc, Apigee |
| Data and analytics | BigQuery, Dataflow, Dataproc, Cloud Composer |
| AI and machine learning | Vertex AI, Document AI, Vision AI, Speech-to-Text, Text-to-Speech |
For your first article, start with “Introduction to Google Cloud: Projects, Regions, and Zones.” Then move to Compute Engine → VPC → Cloud Storage → IAM → GKE → CI/CD → Monitoring.
Google Cloud advertises 150+ products. Its catalog changes as services are introduced, renamed, or retired. Below is a broad documentation index covering services, product families, and important features; it should not be treated as a permanently exhaustive inventory. For the live list, use the
| Category | Services and documentation topics |
| Compute | Compute Engine, Batch, Cloud TPU, GPU workloads, Sole-tenant Nodes, Confidential VM, Shielded VM, VM Manager, Instance Templates, Managed Instance Groups, Autoscaling |
| Containers and application hosting | Google Kubernetes Engine (GKE), GKE Autopilot, GKE Standard, Cloud Run, Cloud Run functions, App Engine |
| Hybrid and multicloud | Google Distributed Cloud, GKE Enterprise, Fleet Management, Connect Gateway, Cloud Service Mesh, Config Sync, Policy Controller |
These categories cover virtual machines, containers, serverless applications, and workloads across cloud and on-premises environments.
| Category | Services and documentation topics |
| Core networking | Virtual Private Cloud (VPC), Shared VPC, VPC Network Peering, Private Service Connect, Private Google Access, Cloud NAT, Network Service Tiers |
| Hybrid connectivity | Cloud VPN, Cloud Interconnect, Cloud Router, Network Connectivity Center |
| Load balancing and DNS | Cloud Load Balancing, Cloud DNS, Cloud Domains, Service Extensions |
| Content delivery | Cloud CDN, Media CDN, CDN Interconnect |
| Network protection and troubleshooting | Cloud Next Generation Firewall, Cloud Armor, Cloud IDS, Secure Web Proxy, Network Intelligence Center, Connectivity Tests, Network Topology, Network Analyzer, VPC Flow Logs |
Google’s networking documentation groups these around connectivity, scalability, security, and network observability.
| Category | Services and documentation topics |
| Object, block, and file storage | Cloud Storage, Persistent Disk, Hyperdisk, Local SSD, Filestore, Google Cloud NetApp Volumes, Parallelstore |
| Backup and recovery | Backup and DR Service, Backup for GKE |
| Data transfer | Storage Transfer Service, Transfer Appliance |
Treat storage classes, snapshots, replication, retention policies, and lifecycle rules as separate articles within the relevant storage service.
| Category | Services and documentation topics |
| Relational databases | Cloud SQL for MySQL, Cloud SQL for PostgreSQL, Cloud SQL for SQL Server, AlloyDB for PostgreSQL, AlloyDB Omni, Spanner, Spanner Omni |
| Document and wide-column databases | Firestore in Native mode, Firestore with MongoDB compatibility, Datastore, Bigtable |
| In-memory databases | Memorystore for Redis, Memorystore for Redis Cluster, Memorystore for Valkey |
| Database migration and management | Database Migration Service, Datastream, Database Center |
These include managed cloud databases and downloadable editions such as AlloyDB Omni and Spanner Omni.
| Category | Services and documentation topics |
| Analytics and business intelligence | BigQuery, BigQuery ML, BigQuery Data Transfer Service, BigQuery Sharing, Looker, Data Studio |
| Data processing | Dataflow, Dataproc, Managed Service for Apache Spark |
| Data integration and orchestration | Cloud Data Fusion, Dataform, Datastream, Managed Service for Apache Airflow, Orchestration Pipelines |
| Data governance and lakehouse | Knowledge Catalog, BigQuery Unified Governance, Lakehouse, Sensitive Data Protection |
For your articles, also cover batch versus streaming processing, data pipelines, metadata, access controls, and cost optimization.
| Category | Services and documentation topics |
| CI/CD and software supply chain | Cloud Build, Cloud Deploy, Artifact Registry, Artifact Analysis, Binary Authorization |
| Developer tools | Cloud Shell, Google Cloud CLI, Cloud Code, Cloud Workstations, Gemini Code Assist, Developer Connect, Developer Device Platform |
| Application organization and design | App Hub, Application Design Center |
| API management | Apigee, Apigee Hybrid, API Gateway, Cloud Endpoints, API Keys, Service Infrastructure |
| Messaging and event processing | Pub/Sub, Cloud Tasks, Cloud Scheduler, Eventarc Standard, Eventarc Advanced |
| Workflow and application integration | Workflows, Application Integration, Integration Connectors |
These provide a strong foundation for your DevOps documentation, particularly build, scan, publish, deploy, and event-driven automation workflows.
| Category | Services and documentation topics |
| Observability | Cloud Monitoring, Cloud Logging, Cloud Trace, Error Reporting, Cloud Profiler, Google Cloud Managed Service for Prometheus |
| Operational monitoring topics | Cloud Audit Logs, Uptime Checks, Synthetic Monitoring, Alerting Policies, Dashboards, Ops Agent, Personalized Service Health |
| Identity and access | Identity and Access Management (IAM), Cloud Identity, Identity Platform, Identity-Aware Proxy (IAP), Access Context Manager |
| Resource administration | Resource Manager, Organization Policy Service, Cloud Asset Inventory, Service Usage, Cloud Quotas, Essential Contacts |
| Access topics | Service Accounts, Workload Identity Federation, Workforce Identity Federation, IAM Recommender |
Some entries are features or administration topics within larger services, which makes them useful individual chapters in your documentation.
| Category | Services and documentation topics |
| Secrets and encryption | Secret Manager, Cloud Key Management Service (Cloud KMS), Cloud HSM, Cloud External Key Manager |
| Certificates | Certificate Manager, Certificate Authority Service |
| Security posture and operations | Security Command Center, Google Security Operations, Google Threat Intelligence, Advisory Notifications |
| Data security and governance | Sensitive Data Protection, VPC Service Controls, Access Approval, Access Transparency, Assured Workloads |
| Application and AI protection | Binary Authorization, Model Armor, Web Risk, Fraud Defense, Cloud Armor, Secure Web Proxy |
Google’s security portfolio covers access control, threat detection, application protection, encryption, and governance.
| Category | Services and documentation topics |
| AI platforms | Gemini Enterprise Agent Platform, Gemini Enterprise |
| Agent development | Agent Development Kit (ADK), Agent Gateway, Agent Registry, Agent Runtime, Agent Studio, Google Cloud MCP Servers |
| Search and grounding | Agent Search, Enterprise Knowledge Graph, RAG Engine, Vector Search |
| Notebooks and model development | Workbench, Colab Enterprise, Model Garden, Managed Training, Experiments, Deep Learning Containers, Deep Learning VM Images |
| MLOps and model serving | Pipelines, Inference, Feature Store, Model Registry, Model Monitoring, Vertex AI TensorBoard |
| Language and speech | Cloud Natural Language API, Cloud Translation, Speech-to-Text, Text-to-Speech |
| Vision and documents | Cloud Vision, Video Intelligence API, Document AI |
| Conversational and customer AI | Dialogflow CX, Dialogflow ES, Agent Assist, Conversational AI Platform, Customer Experience Insights, Contact Center as a Service |
| Specialized AI | AI Commerce Search, Anti Money Laundering AI, Talent Solution |
Use the current official names when publishing AI articles: this part of the catalog changes particularly quickly.
| Category | Services and documentation topics |
| Migration and modernization | Migration Center, Migrate to Virtual Machines, Migrate to Containers, Mainframe Connector, Mainframe Assessment Tool, Dual Run |
| Media services | Live Stream API, Transcoder API, Video Stitcher API |
| Healthcare | Cloud Healthcare API |
| Infrastructure automation | Infrastructure Manager, Terraform on Google Cloud |
| Billing and cost management topics | Cloud Billing, Billing Budgets and Alerts, Billing Export, FinOps Hub, Committed Use Discounts, Recommender |
| Related application platforms | Firebase, AppSheet |
Migration, industry APIs, infrastructure automation, and cost management are additional areas to include in your series.
Google Cloud Fundamentals — practical and interview notes
These fundamentals answer four questions: Where does an application run? Who owns it? Who can access it? Who pays for it?
| Concept | Main purpose | Example |
|---|---|---|
| Region | Geographic deployment location | Mumbai: asia-south1 |
| Zone | Deployment area within a region | asia-south1-a |
| Organization | Company-level ownership and governance | A bank’s Google Cloud organization |
| Folder | Group projects for administration and policies | Production, Nonproduction |
| Project | Organize service resources and configuration | A banking application’s production project |
| Resource hierarchy | Structure ownership and policy inheritance | Organization → Folder → Project → Resources |
| Google Cloud Console | Manage Google Cloud through a browser | Inspect VMs, IAM, logs, and billing |
| Billing account | Track charges and determine who pays | One account funding several projects |
These concepts form Google Cloud’s location, administration, and billing foundations.
1. Regions
A region is an independent geographic area containing zones. Region selection affects latency, service availability, resilience, and data-location requirements.
Examples in India include:
- Mumbai:
asia-south1 - Delhi:
asia-south2
Check availability for the specific service and machine type before choosing a location.
Practical example
For a banking application serving users in India, evaluate Mumbai as the primary deployment region and Delhi as a recovery location. This is an example design: validate latency, replication, recovery targets, service availability, and organizational requirements before implementing it.
Interview questions
Q: How do you choose a region?
Consider user location, latency, data-location requirements, service availability, cost, and disaster recovery needs.
Q: Does deploying in one region provide protection against a regional outage?
It does not automatically provide regional disaster recovery. You need an appropriate strategy for data replication or backups, application recovery, and traffic failover.
2. Zones
A zone is a deployment area inside a region and should be treated as a failure domain. Compute Engine VM instances are examples of zonal resources.
The Mumbai region includes zones such as:
asia-south1-aasia-south1-basia-south1-c
A zone is a logical infrastructure boundary; avoid assuming that every zone corresponds to exactly one physical data center.
Practical example
Place application instances across multiple zones and configure load balancing and health checks. If one zone becomes unavailable, healthy instances in other zones can continue serving traffic—provided dependencies and remaining capacity support that operation.
Interview questions
Q: What is the difference between a region and a zone?
A region is a geographic area; a zone is a deployment area within that region.
Q: Does creating a second VM in the same zone provide zone-level resilience?
No. Both VMs remain exposed to the same zonal failure. Distribute workloads and design their dependencies appropriately.
3. Organizations
An organization resource represents a company and sits at the top of its Google Cloud resource hierarchy. It gives the company centralized ownership and governance over its cloud resources.
Organization resources are associated with Google Workspace or Cloud Identity accounts.
Practical example
A bank places its application projects under its organization. Projects remain company resources when individual employees leave. Administrators delegate responsibilities to groups rather than relying on one employee’s account.
Interview questions
Q: Is an organization mandatory for a personal Google Cloud learning project?
No. You can have a project without an organization. However, an organization is required to use folders.
Q: Does Organization Administrator automatically include every administrative permission?
No. For example, creating folders or projects requires additional permissions. Managing organization IAM and operating every service are separate responsibilities.
4. Folders
A folder groups projects and other folders under an organization. Folders provide useful points for delegated administration and policy application.
Common grouping approaches include:
- Business units: Banking, Insurance, Analytics.
- Environments: Production, Nonproduction.
- Teams or applications: Platform, Payments, Customer Services.
Folders are optional, but they require an organization.
Practical example
Create separate Production and Nonproduction folders. Give developers appropriate access to Nonproduction, while restricting Production administration to approved operations groups.
Folder separation alone does not configure network isolation; network connectivity and access must also be designed.
Interview questions
Q: Can a folder contain another folder?
Yes. Folders can contain projects, nested folders, or both.
Q: What happens when you grant an IAM role on a folder?
The role is inherited by its descendant projects and resources, subject to applicable access controls. Choose the grant carefully because its scope can be broad.
5. Projects
A project is a fundamental organizing unit for Google Cloud service resources. Projects provide a context for API enablement, access management, usage, and billing.
For example, you might use separate projects for development, UAT, and production.
Project identifiers
| Identifier | Purpose | Can it change? |
|---|---|---|
| Project name | Human-readable display name | Yes |
| Project ID | Globally unique identifier used in many commands and APIs | No, after creation |
| Project number | Automatically generated unique numeric identifier | No |
A deleted project’s ID cannot be reused.
Practical example
Use illustrative project IDs such as banking-dev-4821, banking-uat-4821, and banking-prod-4821. Actual IDs must be globally available.
Separate environments so that deployment permissions and resource usage can be managed independently.
Interview questions
Q: Is a project restricted to one region?
No. A project can contain resources in multiple regions and zones. Each resource’s location and scope depend on its service.
Q: Does creating a project automatically enable every API?
No. Enable the APIs required for the services you intend to use; project creation and service configuration are separate steps.
6. Resource hierarchy
The typical enterprise hierarchy is:
Organization → Folders → Projects → Service resources
This structure controls ownership and provides attachment points for IAM and organization policies.
Practical example

Regions and zones describe where appropriate resources run; they are separate from this administrative hierarchy.
Interview questions
Q: Can a project have two parent folders simultaneously?
No. A project has one parent at a time.
Q: Can you remove an inherited IAM allow grant by deleting a project-level binding?
No. IAM allow grants accumulate through the hierarchy. Removing a local binding does not remove a grant inherited from a parent.
Q: Can a deny policy block an action even when a role allows it?
Yes. An applicable IAM deny rule can block supported permissions despite an allow grant. Organization Policy separately governs permitted resource configurations.
7. Google Cloud Console
The Google Cloud Console is Google Cloud’s browser-based graphical interface:
You can use it to manage projects and resources. Google Cloud CLI provides command-line access, and Cloud Shell offers a browser-accessible shell with tools preinstalled.
Practical example
Before investigating a production issue:
- Confirm the selected project.
- Open the relevant service.
- Check the resource’s region or zone.
- Inspect its configuration, logs, and metrics.
- Confirm the target again before making changes.
Interview questions
Q: Does Console access automatically grant permission to manage resources?
No. Signing in identifies the user; IAM determines which resource operations they can perform.
Q: When would you use Console versus CLI or Terraform?
Use Console for exploration and visual inspection. Use CLI for commands and scripts, and Terraform for repeatable infrastructure configuration.
8. Billing accounts
A Cloud Billing account tracks costs for linked projects and determines who pays. It is associated with a Google payments profile.
A project can have one linked billing account at a time, while one billing account can fund multiple projects.
Practical example
Link development, UAT, and production projects to a company billing account. Track spending by project and configure separate budgets so each environment’s costs remain visible.
Interview questions
Q: Is the billing account the IAM parent of its linked projects?
No. Payment linkage is separate from administrative ownership. Projects do not inherit resource permissions from their linked billing account.
Q: Does a budget automatically stop spending?
An alerts-only budget does not cap spending. It triggers notifications at configured thresholds. Google also documents separate spend-cap budgets for supported services; check eligibility and behavior before relying on them.
Practical starter exercise
For a learning environment:
- Create a project with a unique ID.
- Link an appropriate active billing account.
- Configure a budget and alert thresholds.
- Enable the API required for your chosen service.
- Select a suitable region and, where applicable, a zone.
- Deploy a small application.
- Inspect its permissions, logs, and costs.
- Delete unneeded resources after practice.
For a company environment, first place the project under the approved organization and folder and follow its inherited policies.
Compute & Virtual Machines
Google Cloud Compute provides the processing power to run applications, virtual machines, batch jobs, and AI workloads.
Your list includes standalone services and features within Compute Engine. Understanding that distinction will help both your documentation and interview preparation.
| Topic | What it means | Practical example |
|---|---|---|
| Compute Engine | Virtual machines with configurable CPU, memory, disks, networking, and operating systems. | Host a Java banking application on Linux VMs. |
| Batch | A managed service that schedules and executes jobs and provisions the required compute resources. | Process thousands of end-of-day transaction files. |
| Cloud TPU | Google’s specialized accelerators for supported machine-learning workloads. | Train a large machine-learning model. |
| GPU workloads | Applications running on GPU-equipped compute instances. | Accelerate AI inference or video processing. |
| Sole-tenant Nodes | Physical Compute Engine servers dedicated to VMs from your projects. | Meet a requirement for dedicated hardware. |
| Confidential VM | VMs that use hardware-backed protection for data while it is being processed. | Process sensitive customer information. |
| Shielded VM | VM security features that help protect and verify the boot process. | Detect unexpected changes to a VM’s boot environment. |
| VM Manager | Tools for operating-system inventory, patching, and configuration management. | Manage security patches across a Linux VM fleet. |
| Instance Templates | Reusable VM configuration definitions. | Define a standard application-server configuration. |
| Managed Instance Groups | Groups of VMs managed together, with capabilities such as autohealing and rolling updates. | Run multiple application servers across zones. |
| Autoscaling | Automatically adjusts the number of VMs in a managed instance group. | Add application servers when customer traffic increases. |
The examples are illustrative designs; the service capabilities come from Google’s compute documentation.
Google Cloud Compute — Practical Guide with Diagrams and Interview Answers
Google Cloud compute services provide the CPU, memory, and accelerators needed to run applications and process data. Some services run continuously, such as application VMs; others execute jobs or accelerate machine-learning workloads.
The examples below are illustrative real-world scenarios, not claims about a particular organization’s production environment.
Understand the compute family first

Compute Engine, Batch, and Cloud TPU are services. Instance templates, MIGs, autoscaling, and several VM security capabilities belong to the broader Compute Engine ecosystem.
1. Compute Engine — Run your application on virtual machines
Compute Engine lets you create Linux or Windows VMs with selected CPU, memory, disks, and networking.
Example: A loan-processing application requires Java, custom OS packages, and a background service. A VM provides the operating-system control needed to install these components.

How the example works
- The customer submits a loan request.
- The load balancer sends it to a healthy application instance.
- The application validates the request and accesses the database.
- Monitoring records latency, errors, and resource utilization.
Google manages the underlying infrastructure; your team manages the guest OS and application. Multiple VMs help only when the surrounding architecture supports failover.
Interview answer: “Compute Engine is Google Cloud’s VM service. I use it when workloads need control over the operating system, installed software, and machine configuration.”
2. Batch — Process large jobs without maintaining permanent workers
Batch schedules and executes scripts or containers, provisioning resources for their tasks.
Example: Generate monthly statements for a large customer base. Divide the input into partitions so independent tasks can run in parallel.

Practical design
- Give each output a deterministic identifier, such as customer ID plus statement month.
- Record task completion so retries do not produce duplicate statements.
- Keep outputs in durable storage.
- Inspect failed-task logs before rerunning work.
Batch manages execution infrastructure; your application handles business correctness. Its underlying compute, storage, and other resources incur charges.
Interview answer: “Batch suits finite processing jobs, especially parallel workloads. It manages provisioning and task execution instead of requiring us to maintain a job-processing VM fleet.”
3. Cloud TPU — Accelerate supported machine-learning computation
Cloud TPU provides Google-designed Tensor Processing Units for supported ML workloads.
Example: A data-science team trains a transaction-risk model using a compatible framework and model architecture.

What to evaluate
- Framework and operation compatibility.
- Training throughput.
- Input-pipeline performance.
- Time and cost to reach the required model quality.
Interview answer: “TPUs accelerate supported tensor-heavy ML workloads. I would choose them after validating compatibility and benchmarking the actual model.”
4. GPU workloads — Accelerate parallel processing
GPU-equipped instances support workloads such as AI inference, model training, graphics, and scientific computation.
Example: An application classifies uploaded document images. CPUs handle requests and preprocessing; a GPU executes model inference.

Practical checks
| Check | Why it matters |
|---|---|
| GPU memory | The model and its working data must fit |
| Driver compatibility | The framework must work with installed drivers |
| Quota and capacity | Resources must be available in the selected location |
| GPU utilization | Reveals idle capacity or processing bottlenecks |
| Latency and throughput | Shows whether acceleration meets application needs |
Interview answer: “Before deploying GPU workloads, I check memory requirements, machine compatibility, drivers, quota, availability, and performance.”
5. Sole-tenant Nodes — Use dedicated physical hosts
Sole tenancy dedicates physical Compute Engine servers to VMs from your projects. One dedicated host can run multiple VMs.
Example: An enterprise requires dedicated hosts for a particular workload or needs host-level control for a software licensing arrangement.

Practical considerations
- Plan host capacity and placement.
- Verify software licensing conditions.
- Design maintenance and recovery procedures.
- Evaluate whether dedicated capacity justifies its cost.
Dedicated hardware does not independently provide compliance or application availability.
Interview answer: “Sole-tenant nodes provide dedicated physical hosts. I consider them when hardware isolation or licensing requirements justify that architecture.”
6. Confidential VM — Protect data during processing
Confidential VM uses supported hardware-backed technologies to protect data in use, complementing encryption in transit and at rest.
Example: A service processes sensitive customer information in memory.

Understand the three protection stages
| Stage | Meaning |
|---|---|
| At rest | Data stored on disks or storage services |
| In transit | Data moving across a network |
| In use | Data actively processed by an application |
Confidential VM addresses the processing stage. Secure application code, access controls, and secret management remain essential.
Interview answer: “Confidential VM adds hardware-backed protection for data in use. I check supported configurations and workload requirements before selecting it.”
7. Shielded VM — Establish trust in the boot environment
Shielded VM provides Secure Boot, virtual TPM, and integrity-monitoring capabilities.
Example: Following an OS change, the operations team investigates an unexpected boot-integrity signal.

Not every integrity change proves an attack: authorized updates can also change measurements. Investigate against the approved change record.
Interview distinction: “Shielded VM focuses on boot trust and integrity; Confidential VM protects data during processing.”
Check driver and kernel-module compatibility before enabling Secure Boot.
8. VM Manager — Manage operating systems across a fleet
VM Manager provides OS inventory, patch management, and OS policy capabilities.
Example: Apply Linux security updates across an application fleet while preserving service availability.

Practical controls
- Establish maintenance windows.
- Prepare backups and recovery procedures.
- Patch controlled groups.
- Check application health after each group.
- Confirm the final patch status.
Interview answer: “I use VM Manager for OS visibility, patching, and desired configuration. Production changes proceed through testing, controlled execution, and health verification.”
9. Instance Templates — Define repeatable VM configurations
An instance template defines settings used to create VMs, including machine configuration, image, disks, networking, and metadata.
Example: Deploy application version 2 using a new template and a controlled MIG update.

Templates are immutable. Create a new template when the VM configuration changes.
A template is a configuration definition; a VM image contains the operating system and potentially installed software. Creating a template alone does not update running instances.
Interview answer: “I version instance templates and apply them through controlled MIG rollouts, validating health before continuing.”
10. Managed Instance Groups — Operate VMs as an application fleet
A MIG manages instances using a template and supports capabilities such as rolling updates, autohealing, and autoscaling.
Example: Run a stateless application across zones and recreate an instance when an application health check repeatedly fails.

Two health-check purposes
| Check | Purpose |
|---|---|
| Load-balancer health check | Determines whether an instance should receive traffic |
| MIG autohealing health check | Determines whether an instance should be recreated |
Use appropriate thresholds. An overly aggressive autohealing check can repeatedly recreate an application that needs time to start.
Interview answer: “A regional MIG helps distribute instances across zones. Autohealing restores unhealthy instances, while rolling updates control configuration changes.”
11. Autoscaling — Match VM capacity to demand
Autoscaling changes the number of VMs in a MIG using configured signals and limits.
Example configuration for learning
| Setting | Illustrative value |
|---|---|
| Minimum instances | 2 |
| Maximum instances | 8 |
| CPU utilization target | 60% |
| Initialization period | Based on measured application startup |
These values are examples, not production defaults.

A 60% CPU target is a utilization objective used to calculate capacity; it does not mean “immediately add one VM whenever CPU crosses 60%.”
Interview answer: “I select a signal that reflects demand, configure minimum and maximum capacity, and account for startup time. I also investigate downstream bottlenecks because adding VMs cannot resolve every performance issue.”
Putting the services together: an illustrative banking application

In this design:
- Compute Engine hosts the application.
- Templates and MIGs support repeatable deployment and instance recovery.
- Autoscaling adjusts application capacity.
- VM Manager supports OS maintenance.
- Batch handles statement-generation jobs.
- Shielded VM, Confidential VM, and sole tenancy are evaluated against specific security and isolation requirements.
- GPUs or TPUs are added when an ML workload justifies them.
A practical troubleshooting example
Suppose customer traffic rises, the MIG adds instances, but users still experience slow responses.
| Observation | What to investigate |
|---|---|
| New instances remain unhealthy | Startup logs, dependencies, health-check port and path |
| Instance count reaches its maximum | Scaling limits, quota, available capacity |
| CPU remains low but latency rises | Database, connection pools, external APIs |
| Instances repeatedly recreate | Autohealing thresholds and initialization time |
| Only some requests fail | Configuration differences, sessions, application versions |
The useful lesson for readers is to follow the request through the system: scaling increases compute capacity, while troubleshooting identifies the actual bottleneck.
Discover more from DevOps with Patil
Subscribe to get the latest posts sent to your email.