DEV Community

Kaushik Mitra
Kaushik Mitra

Posted on Fully Autonomous

Creating a GKE Cluster and Deploying a Microservice Application (End-to-End)

We've reached the final milestone of our Kubernetes learning journey!

So far, we have:

  1. Set up a local Kubernetes cluster using Docker Desktop.
  2. Built two microservices (add-service and multiply-service) and wired them together using Namespaces, Pods, and ClusterIP Services.
  3. Configured NGINX Ingress Controller for path-based routing (/api/add and /api/multiply) with URL rewriting.
  4. Pushed our container images to Google Cloud Artifact Registry.

Now, it's time to bring everything together into production: we'll provision a real Google Kubernetes Engine (GKE) cluster on Google Cloud, deploy our application, configure cloud ingress with a public Google Cloud Load Balancer, and test our live API on the internet!


Why GKE Autopilot?

Google Kubernetes Engine offers two operation modes:

  1. Standard Mode: You manage and pay for the underlying virtual machines (Compute Engine worker nodes). You handle node provisioning, OS patching, and node pool configurations.
  2. Autopilot Mode: Google manages the entire underlying node infrastructure, node pools, OS upgrades, and security hardening. You only define what your workloads need, and you pay exclusively for the CPU, memory, and storage requested by your actively running Pods.

For modern applications and learning environments, Autopilot is fantastic because it eliminates node management overhead completely while enforcing Kubernetes production best practices out of the box.


Prerequisites

Make sure you have:

  • A GCP account and project (e.g., kkm-kube-practice) with billing enabled.
  • Container images already pushed to Artifact Registry (asia-south1-docker.pkg.dev/...).
  • gcloud and kubectl installed on your machine.
  • The Kubernetes manifests and source code available on GitHub: MitraKumar/kube-prac-calculator-app.

Step 1: Enable the Kubernetes Engine API

Before creating a cluster, ensure the GKE API is active on your project:

gcloud services enable container.googleapis.com
Enter fullscreen mode Exit fullscreen mode

Step 2: Provision the GKE Autopilot Cluster

Let's create an Autopilot cluster named autopilot-cluster-1 in the us-central1 region:

gcloud container clusters create-auto autopilot-cluster-1 \
    --region us-central1 \
    --project kkm-kube-practice
Enter fullscreen mode Exit fullscreen mode

Note: Cluster provisioning usually takes 4 to 7 minutes as Google provisions the regional control planes and private VPC networking. Sit back, grab a coffee, and let Google do the heavy lifting!

Once creation finishes, you'll see:

NAME                 LOCATION     MASTER_VERSION      STATUS
autopilot-cluster-1  us-central1  1.31.x-gke.xxx      RUNNING
Enter fullscreen mode Exit fullscreen mode

Step 3: Connect kubectl to Your GKE Cluster

To instruct your local kubectl to talk to your brand new GKE cluster, download the cluster credentials:

gcloud container clusters get-credentials autopilot-cluster-1 \
    --region us-central1 \
    --project kkm-kube-practice
Enter fullscreen mode Exit fullscreen mode

Verify your active context:

kubectl config current-context
Enter fullscreen mode Exit fullscreen mode

Output:

gke_kkm-kube-practice_us-central1_autopilot-cluster-1
Enter fullscreen mode Exit fullscreen mode

Let's inspect the nodes:

kubectl get nodes
Enter fullscreen mode Exit fullscreen mode

You will see the initial worker nodes provisioned by Autopilot.


Step 4: Important Autopilot Requirement โ€” Resource Requests

One key difference between local Docker Desktop and GKE Autopilot is that Autopilot requires every container to declare CPU and memory requests and limits.

Because Autopilot scales nodes dynamically based on workload demands and bills based on pod resource consumption, it needs to know your pod's requirements.

If you inspect our Pod manifests (k8s/01-add-service-pod.yaml and k8s/03-multiply-service-pod.yaml), you'll see we explicitly declared:

resources:
  requests:
    memory: "128Mi"
    cpu: "250m"
  limits:
    memory: "128Mi"
    cpu: "250m"
Enter fullscreen mode Exit fullscreen mode

๐Ÿ’ก Tip: If you deploy a pod without resources on GKE Autopilot, Autopilot will automatically inject default resource values, but it's best practice to specify them explicitly.


Step 5: Install NGINX Ingress Controller on GKE using Helm

Just like we did locally, we need an Ingress Controller to route external HTTP traffic to our services. We'll install the official chart using Helm:

# 1. Add and update the official repository
helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx
helm repo update

# 2. Install the NGINX Ingress Controller chart into GKE
helm install ingress-nginx ingress-nginx/ingress-nginx \
  --namespace ingress-nginx \
  --create-namespace
Enter fullscreen mode Exit fullscreen mode

Output:

NAME: ingress-nginx
LAST DEPLOYED: Sat Oct  3 20:08:22 2026
NAMESPACE: ingress-nginx
STATUS: deployed
REVISION: 1
TEST SUITE: None
NOTES:
The ingress-nginx controller has been installed.
It may take a few minutes for the LoadBalancer IP to be available.
Enter fullscreen mode Exit fullscreen mode

What happens on Google Cloud?

When Helm installs this chart on GKE:

  1. GKE Autopilot schedules the ingress-nginx-controller pod and allocates compute resources automatically.
  2. GKE detects the type: LoadBalancer Service in the ingress-nginx namespace.
  3. Google Cloud automatically provisions a Google Cloud Network Load Balancer (NLB) in your GCP project!
  4. Google assigns a public, routable IPv4 address to that Load Balancer.

Watch the service until the public EXTERNAL-IP is assigned:

kubectl get svc -n ingress-nginx -w
Enter fullscreen mode Exit fullscreen mode

Initial output:

NAME                       TYPE           CLUSTER-IP      EXTERNAL-IP   PORT(S)
ingress-nginx-controller   LoadBalancer   34.118.235.60   <pending>     80:32234/TCP,443:32055/TCP
Enter fullscreen mode Exit fullscreen mode

After about 1 to 2 minutes, the external IP will appear:

NAME                       TYPE           CLUSTER-IP      EXTERNAL-IP     PORT(S)
ingress-nginx-controller   LoadBalancer   34.118.235.60   136.80.62.73    80:32234/TCP,443:32055/TCP
Enter fullscreen mode Exit fullscreen mode

We now have a dedicated public IP (136.80.62.73) routing internet traffic directly into our cluster!


Step 6: Deploy the Calculator Microservices

Now, let's deploy our manifests:

  1. Namespace: 00-namespace.yaml
  2. Add Service Pod & Service: 01-add-service-pod.yaml, 02-add-service-cluster-ip.yaml
  3. Multiply Service Pod & Service: 03-multiply-service-pod.yaml, 04-multiply-service-cluster-ip.yaml
  4. Docs Service Pod & Service: 06-docs-service-pod.yaml, 07-docs-service-cluster-ip.yaml
  5. Ingress Routing Rules: 05-nginx-ingress-controller.yaml

Apply them sequentially (or all together):

kubectl apply -f k8s/00-namespace.yaml
kubectl apply -f k8s/01-add-service-pod.yaml
kubectl apply -f k8s/02-add-service-cluster-ip.yaml
kubectl apply -f k8s/03-multiply-service-pod.yaml
kubectl apply -f k8s/04-multiply-service-cluster-ip.yaml
kubectl apply -f k8s/06-docs-service-pod.yaml
kubectl apply -f k8s/07-docs-service-cluster-ip.yaml
kubectl apply -f k8s/05-nginx-ingress-controller.yaml
Enter fullscreen mode Exit fullscreen mode

Verify the Deployments

Check the pods and services:

kubectl get all -n calculator-app
Enter fullscreen mode Exit fullscreen mode

Output:

NAME                       READY   STATUS    RESTARTS   AGE
pod/add-service            1/1     Running   0          2m
pod/docs-service           1/1     Running   0          1m
pod/multiply-service       1/1     Running   0          2m

NAME                                  TYPE        CLUSTER-IP       EXTERNAL-IP   PORT(S)    AGE
service/add-service-cluster-ip        ClusterIP   34.118.228.222   <none>        3000/TCP   2m
service/docs-service-cluster-ip       ClusterIP   34.118.228.50    <none>        3000/TCP   1m
service/multiply-service-cluster-ip   ClusterIP   34.118.230.113   <none>        3000/TCP   2m
Enter fullscreen mode Exit fullscreen mode

Check the Ingress resource:

kubectl get ingress -n calculator-app
Enter fullscreen mode Exit fullscreen mode

Output:

NAME                     CLASS   HOSTS   ADDRESS        PORTS   AGE
calculator-api-ingress   nginx   *       136.80.62.73   80      2m
Enter fullscreen mode Exit fullscreen mode

The Ingress is active and bound to our Google Cloud Load Balancer IP (136.80.62.73).


Step 7: Test Your Live Production Application!

Let's test our live microservices from anywhere in the world using curl (or directly from your browser):

1. Test Interactive Swagger UI Documentation:

curl -sI "http://136.80.62.73/"
Enter fullscreen mode Exit fullscreen mode

Response:

HTTP/1.1 200 OK
Content-Type: text/html; charset=utf-8
Enter fullscreen mode Exit fullscreen mode

๐ŸŒ Open http://136.80.62.73/ in your browser to explore the interactive Swagger documentation and test API endpoints right from the UI!

2. Test Addition API:

curl -s "http://136.80.62.73/api/add?a=10&b=20"
Enter fullscreen mode Exit fullscreen mode

Response:

{
  "service": "add-service",
  "operation": "addition",
  "a": 10,
  "b": 20,
  "result": 30
}
Enter fullscreen mode Exit fullscreen mode

3. Test Multiplication API:

curl -s "http://136.80.62.73/api/multiply?a=5&b=6"
Enter fullscreen mode Exit fullscreen mode

Response:

{
  "service": "multiply-service",
  "operation": "multiplication",
  "a": 5,
  "b": 6,
  "result": 30
}
Enter fullscreen mode Exit fullscreen mode

3. Test Health Check:

curl -s "http://136.80.62.73/api/add/health"
Enter fullscreen mode Exit fullscreen mode

Response:

{
  "status": "UP",
  "service": "add-service"
}
Enter fullscreen mode Exit fullscreen mode

It works! The Google Cloud Load Balancer receives the HTTP request on port 80, directs it to the NGINX Ingress Controller, which rewrites the path and dispatches it over the internal cluster network to our respective Node.js pods.


Step 8: Clean Up Resources (Don't Skip This!)

Cloud resources cost money when left running. If you are experimenting or practicing, remember to tear down the cluster once you're done:

gcloud container clusters delete autopilot-cluster-1 \
    --region us-central1 \
    --project kkm-kube-practice \
    --quiet
Enter fullscreen mode Exit fullscreen mode

Deleting the GKE cluster will automatically clean up the associated Google Cloud Network Load Balancer and internal VPC forwarding rules.


Conclusion & Series Retrospective

Over this 5-part journey, we built a complete end-to-end Kubernetes workflow:

  1. Local Setup: Bootstrapped a frictionless single-node cluster using Docker Desktop.
  2. Kubernetes Core: Built multi-service Node.js Express apps and mastered Namespaces, Pods, ClusterIP Services, Labels, and Selectors.
  3. Ingress Routing: Configured the NGINX Ingress Controller to route path-based traffic with regex rewrites.
  4. Artifact Management: Tagged and published production containers to Google Cloud Artifact Registry.
  5. Cloud Deployment: Created a managed GKE Autopilot cluster, provisioned a Cloud Load Balancer, and served live API traffic.

I hope this step-by-step guide helps you navigate Kubernetes and cloud deployments with confidence.

If you followed along or have questions about GKE, Ingress, or microservices, drop a comment belowโ€”I'd love to hear your thoughts!

Top comments (0)