Chapter 3.5: Deploy and access the model on Kubernetes¶
Introduction¶
In this chapter, you will learn how to deploy the model on Kubernetes and access it from a Kubernetes pod using the previous Docker image.
This will allow the model to be used by other applications and services on a public endpoint accessible from anywhere.
In this chapter, you will learn how to:
- Create the Kubernetes cluster
- Validate kubectl can access the Kubernetes cluster
- Create the Kubernetes configuration files
- Deploy the Docker image on Kubernetes
- Access the model
Why Kubernetes instead of a simple VM?
For a single model, a virtual machine would indeed be simpler. But Kubernetes is the de facto standard for running containers in production: the same declarative configuration deploys on a Raspberry Pi cluster or a large cloud cluster, and is understood by any engineer. It also provides self-healing, scaling, and load balancing that you would script by hand on a VM, and the same manifests are reused for CI/CD, training, and monitoring in the next chapters.
Danger
The following steps will create resources on the cloud provider. These resources will be deleted at the end of the guide, but you might be charged for them. Kubernetes clusters are not free and can be expensive. Make sure to delete the resources at the end of the guide.
The following diagram illustrates the control flow of the experiment at the end of this chapter:
flowchart TB
dot_dvc[(.dvc)] <-->|dvc pull
dvc push| s3_storage[(S3 Storage)]
dot_git[(.git)] <-->|git pull
git push| repository[(Repository)]
workspaceGraph <-....-> dot_git
data[data/raw]
subgraph cacheGraph[CACHE]
dot_dvc
dot_git
end
subgraph workspaceGraph[WORKSPACE]
data --> code[*.py]
subgraph dvcGraph["dvc.yaml"]
code
end
params[params.yaml] -.- code
code <--> bento_model[classifier.bentomodel]
subgraph bentoGraph[bentofile.yaml]
bento_model
serve[serve.py] <--> bento_model
end
bento_model <-.-> dot_dvc
end
subgraph remoteGraph[REMOTE]
s3_storage
subgraph gitGraph[Git Remote]
repository <--> |...|action[Action]
end
registry[(Container
registry)]
action --> |bentoml build
bentoml containerize
docker push|registry
subgraph clusterGraph[Kubernetes]
bento_service_cluster[classifierService] --> k8s_fastapi[FastAPI]
end
registry --> |kubectl apply|bento_service_cluster
end
subgraph browserGraph[BROWSER]
k8s_fastapi <--> publicURL["public URL"]
end
style workspaceGraph opacity:0.4,color:#7f7f7f80
style dvcGraph opacity:0.4,color:#7f7f7f80
style cacheGraph opacity:0.4,color:#7f7f7f80
style data opacity:0.4,color:#7f7f7f80
style dot_git opacity:0.4,color:#7f7f7f80
style dot_dvc opacity:0.4,color:#7f7f7f80
style code opacity:0.4,color:#7f7f7f80
style bentoGraph opacity:0.4,color:#7f7f7f80
style serve opacity:0.4,color:#7f7f7f80
style bento_model opacity:0.4,color:#7f7f7f80
style params opacity:0.4,color:#7f7f7f80
style s3_storage opacity:0.4,color:#7f7f7f80
style remoteGraph opacity:0.4,color:#7f7f7f80
style gitGraph opacity:0.4,color:#7f7f7f80
style repository opacity:0.4,color:#7f7f7f80
style action opacity:0.4,color:#7f7f7f80
linkStyle 0 opacity:0.4,color:#7f7f7f80
linkStyle 1 opacity:0.4,color:#7f7f7f80
linkStyle 2 opacity:0.4,color:#7f7f7f80
linkStyle 3 opacity:0.4,color:#7f7f7f80
linkStyle 4 opacity:0.4,color:#7f7f7f80
linkStyle 5 opacity:0.4,color:#7f7f7f80
linkStyle 6 opacity:0.4,color:#7f7f7f80
linkStyle 7 opacity:0.4,color:#7f7f7f80
linkStyle 8 opacity:0.4,color:#7f7f7f80
linkStyle 9 opacity:0.4,color:#7f7f7f80
Steps¶
Install the Kubernetes CLI¶
Install the Kubernetes CLI (kubectl) on your machine with the Google Cloud CLI. You might need to follow the instructions in the terminal and the related documentation.
# Install kubectl with gcloud
gcloud components install kubectl
As per the instructions, you will need to install the gke-gcloud-auth-plugin
authentication plugin as well.
kubectl fails with a client-go credential plugin error?
If kubectl fails with an error about a missing client-go credential plugin,
the gke-gcloud-auth-plugin is not installed. Install it with the following
command:
Create the Kubernetes cluster¶
In order to deploy the model on Kubernetes, you will need a Kubernetes cluster.
Follow the steps below to create one.
Enable the Google Kubernetes Engine API
You must enable the Google Kubernetes Engine API to create Kubernetes clusters on Google Cloud with the following command:
Tip
You can display the available services in your project with the following command:
# Enable the Google Kubernetes Engine API
gcloud services enable container.googleapis.com
Create the Kubernetes cluster
Create the Google Kubernetes cluster with the Google Cloud CLI.
Export the cluster name as an environment variable. Replace <my_cluster_name>
with a cluster name of your choice. It has to be lowercase and words separated
by hyphens.
Warning
The cluster name must be unique across all Google Cloud projects and users.
For example, use mlops-<surname>-cluster, where surname is based on your
name. Change the cluster name if the command fails.
Export the cluster zone as an environment variable. You can view the available
zones at
Regions and zones.
You should ideally select a zone close to where most of the expected traffic
will come from. Replace <my_cluster_zone> with your own zone (ex:
europe-west6-a for Zurich, Switzerland).
Create the Kubernetes cluster. You can also view the available types of machine
with the gcloud compute machine-types list command.
Info
This can take several minutes. Please be patient.
# Create the Kubernetes cluster
gcloud container clusters create \
--machine-type=e2-standard-2 \
--num-nodes=2 \
--zone=$GCP_K8S_CLUSTER_ZONE \
$GCP_K8S_CLUSTER_NAME
The output should be similar to this:
Note: Your Pod address range (`--cluster-ipv4-cidr`) can accommodate at most 1008 node(s).
Creating cluster mlops-surname-cluster in europe-west6-a... Cluster is being health-checked (Kubernetes Control Plane is healthy)...done.
Created [https://container.googleapis.com/v1/projects/mlops-surname-project/zones/europe-west6-a/clusters/mlops-surname-cluster].
To inspect the contents of your cluster, go to: https://console.cloud.google.com/kubernetes/workload_/gcloud/europe-west6-a/mlops-surname-cluster?project=mlops-surname-cluster
kubeconfig entry generated for mlops-surname-cluster.
NAME LOCATION MASTER_VERSION MASTER_IP MACHINE_TYPE NODE_VERSION NUM_NODES STATUS STACK_TYPE
mlops-surname-cluster europe-west6-a 1.35.3-gke.1389002 34.65.125.181 e2-standard-2 1.35.3-gke.1389002 2 RUNNING IPV4
Cluster configuration and cost considerations
This cluster uses 2 nodes and e2-standard-2 machines to prepare for the next
chapters, where we'll demonstrate managing heterogeneous node types (GPU and
CPU). While e2-medium would suffice for this tutorial, the cost difference is
minimal.
For most use cases, high-availability (HA) clusters spanning multiple zones aren't necessary and significantly increase costs. This tutorial uses a cost-effective zonal cluster. In production, optimize costs with smaller machine types, autoscaling, and properly defined resource requirements. Note that GPUs are the primary cost drivers for ML deployments. See the Compute Engine pricing page for details.
Validate kubectl can access the Kubernetes cluster¶
Validate kubectl can access the Kubernetes cluster:
The output should be similar to this:
NAME STATUS AGE
default Active 2m11s
gke-managed-cim Active 96s
gke-managed-networking-dra-driver Active 36s
gke-managed-system Active 77s
gke-managed-volumepopulator Active 72s
gmp-public Active 52s
gmp-system Active 52s
kube-node-lease Active 2m11s
kube-public Active 2m11s
kube-system Active 2m11s
Tip
If you have an existing Kubernetes configuration on your local machine, you can merge your new GCP cluster's credentials into your current kubeconfig using the following command:
Create the Kubernetes configuration files¶
In order to deploy the model on Kubernetes, you will need to create the Kubernetes configuration files. These files describe the deployment and service of the model.
Create a new directory called kubernetes in the root of the project, then
create a new file called deployment.yaml in the kubernetes directory with
the following content.
Tip
Verify your Docker image is available with the following command:
apiVersion: apps/v1
kind: Deployment
metadata:
name: celestial-bodies-classifier-deployment
labels:
app: celestial-bodies-classifier
spec:
replicas: 1
selector:
matchLabels:
app: celestial-bodies-classifier
template:
metadata:
labels:
app: celestial-bodies-classifier
spec:
containers:
- name: celestial-bodies-classifier
image: <docker_image>
Replace the <docker_image> placeholder with the actual Docker image path:
# Replace the placeholder with the actual Docker image
sed -i "s|<docker_image>|$GCP_CONTAINER_REGISTRY_HOST/celestial-bodies-classifier:latest|g" kubernetes/deployment.yaml
Create a new file called service.yaml in the kubernetes directory with the
following content:
apiVersion: v1
kind: Service
metadata:
name: celestial-bodies-classifier-service
spec:
type: LoadBalancer
ports:
- name: http
port: 80
targetPort: 3000
protocol: TCP
selector:
app: celestial-bodies-classifier
-
The
deployment.yamlfile describes the Deployment: the desired state of the pods running the model, such as the number of replicas and the container image to use. Kubernetes keeps that many pods alive and replaces them if they fail. -
The
service.yamlfile describes the Service: a stable entry point that routes traffic to the pods selected by theirapplabel. Pods are ephemeral and their IP addresses change, so the Service abstracts them away. Withtype: LoadBalancer, the cloud provider also provisions an external load balancer with a public IP address.
In short, the deployment runs the workloads while the service exposes them.
Deploy the containerised model on Kubernetes¶
To deploy the containerised Bento model artifact on Kubernetes, you will need to apply the Kubernetes configuration files.
Apply the Kubernetes configuration files with the following commands:
# Apply the deployment
kubectl apply -f kubernetes/deployment.yaml
# Apply the service
kubectl apply -f kubernetes/service.yaml
The output should be similar to this:
deployment.apps/celestial-bodies-classifier-deployment created
service/celestial-bodies-classifier-service created
Open the cluster interface on the cloud provider and check that the model has been deployed.
Open the Kubernetes Engine on the Google Cloud interface and click on your cluster to access the details.
Access the model¶
To access the model, you will need to find the external IP address of the service. You can do so with the following command:
Info
The external IP address of the service can take a few minutes to be available.
# Get the description of the service
kubectl describe services celestial-bodies-classifier
The output should be similar to this:
Name: celestial-bodies-classifier-service
Namespace: default
Labels: <none>
Annotations: cloud.google.com/neg: {"ingress":true}
Selector: app=celestial-bodies-classifier
Type: LoadBalancer
IP Family Policy: SingleStack
IP Families: IPv4
IP: 34.118.228.127
IPs: 34.118.228.127
LoadBalancer Ingress: 34.65.255.92 (VIP)
Port: http 80/TCP
TargetPort: 3000/TCP
NodePort: http 30773/TCP
Endpoints:
Session Affinity: None
External Traffic Policy: Cluster
Internal Traffic Policy: Cluster
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal EnsuringLoadBalancer 91s service-controller Ensuring load balancer
Normal EnsuredLoadBalancer 56s service-controller Ensured load balancer
The LoadBalancer Ingress field contains the external IP address of the
service. In this case, it is 34.65.255.92.
Try to access the model at the port 80 using the external IP address of the
service. You should be able to access the FastAPI documentation page at
http://<load balancer ingress ip>:80. In this case, it is
http://34.65.255.92:80.
Plain HTTP only
The load balancer exposes the service on port 80 over plain HTTP, without TLS.
Make sure to use http:// and not https:// when accessing it. Most browsers
default to HTTPS, which will fail to connect.
A production deployment should expose the service through a domain name, terminate TLS with an Ingress and a certificate, and restrict access to the endpoint.
Check the changes¶
Check the changes with Git to ensure that all the necessary files are tracked:
# Add all the files
git add .
# Check the changes
git status
The output should look similar to this:
On branch main
Your branch is up to date with 'origin/main'.
Changes to be committed:
(use "git restore --staged <file>..." to unstage)
new file: kubernetes/deployment.yaml
new file: kubernetes/service.yaml
Commit the changes to Git¶
Commit the changes to Git.
# Commit the changes
git commit -m "Use Kubernetes to deploy the model"
# Push the changes
git push
Summary¶
Congratulations! You have successfully deployed the model on Kubernetes with BentoML and Docker, accessed it from an external IP address.
You can now use the model from anywhere.
In this chapter, you have successfully:
- Created the Kubernetes configuration files and deployed the BentoML model artifact on Kubernetes
- Accessed the model
You fixed some of the previous issues:
- Model is accessible from the Internet and can be used anywhere
Take away
- Kubernetes provides production-grade model hosting: By deploying your model to Kubernetes, you gain access to enterprise features like automatic scaling, rolling updates, health checks, and high availability that are difficult to implement with simpler hosting solutions.
- Declarative configuration enables reproducible deployments: Kubernetes YAML files describe the desired state of your deployment (replicas, image version, resources), allowing you to version control your infrastructure and reproduce deployments consistently.
- LoadBalancer services expose models to the internet: Kubernetes Service objects with type LoadBalancer automatically provision cloud load balancers and public IPs, making your model accessible from anywhere without manual networking configuration.
- Separation of deployment and service enables flexibility: Keeping deployment configuration (how many pods, which image) separate from service configuration (how to route traffic) allows you to update the model version without changing how clients access it.
State of the MLOps process¶
- Model can be saved and loaded with all required artifacts for future usage
- Model can be easily used outside of the experiment context
- Model publication to the artifact registry is automated
- Model is accessible from the Internet and can be used anywhere
- Model requires manual deployment on the cluster
- Model cannot be trained on hardware other than the local machine
- Model cannot be trained on custom hardware for specific use-cases
Continue to the next chapters to address the remaining items.
Sources¶
Highly inspired by: