Chapter 2.3 - Reproduce the ML experiment in a CI/CD pipeline¶
Introduction¶
At this point, your code, your data and your execution process should be shared with Git and DVC.
Now, it's time to enhance your workflow further by incorporating a CI/CD (Continuous Integration/Continuous Deployment) pipeline. This addition will enable you to execute your ML experiments remotely and reproduce them, ensuring that any changes made to the project won't inadvertently break. This helps eliminate the notorious "but it works on my machine" effect, where code behaves differently across different environments.
In this chapter, you will learn how to:
- Grant access to the S3 bucket on the cloud provider
- Store the cloud provider credentials in the CI/CD configuration
- Create the CI/CD pipeline configuration file
- Push the CI/CD pipeline configuration file to Git
- Visualize the execution of the CI/CD pipeline
The following diagram illustrates the control flow of the experiment at the end of this chapter:
flowchart TB
dot_dvc[(.dvc)] <-->|dvc push
dvc pull| s3_storage[(S3 Storage)]
dot_git[(.git)] <-->|git push
git pull| gitGraph[Git Remote]
workspaceGraph <-....-> dot_git
data[data/raw] <-.-> dot_dvc
subgraph remoteGraph[REMOTE]
s3_storage
subgraph gitGraph[Git Remote]
direction TB
repository[(Repository)] --> action[Action]
action -->|dvc pull| action_data[data/raw]
action_data -->|dvc repro| action_out[metrics & plots]
end
end
subgraph cacheGraph[CACHE]
dot_dvc
dot_git
end
subgraph workspaceGraph[WORKSPACE]
prepare[prepare.py] <-.-> dot_dvc
train[train.py] <-.-> dot_dvc
evaluate[evaluate.py] <-.-> dot_dvc
data --> prepare
subgraph dvcGraph["dvc.yaml (dvc repro)"]
prepare --> train
train --> evaluate
end
params[params.yaml] -.- prepare
params -.- train
params <-.-> dot_dvc
end
style cacheGraph opacity:0.4,color:#7f7f7f80
style workspaceGraph opacity:0.4,color:#7f7f7f80
style dot_git opacity:0.4,color:#7f7f7f80
style dot_dvc opacity:0.4,color:#7f7f7f80
style data opacity:0.4,color:#7f7f7f80
style prepare opacity:0.4,color:#7f7f7f80
style params opacity:0.4,color:#7f7f7f80
style train opacity:0.4,color:#7f7f7f80
style evaluate opacity:0.4,color:#7f7f7f80
style dvcGraph opacity:0.4,color:#7f7f7f80
style s3_storage opacity:0.4,color:#7f7f7f80
style repository opacity:0.4,color:#7f7f7f80
linkStyle 0 opacity:0.4,color:#7f7f7f80
linkStyle 1 opacity:0.4,color:#7f7f7f80
linkStyle 2 opacity:0.4,color:#7f7f7f80
linkStyle 3 opacity:0.4,color:#7f7f7f80
linkStyle 4 opacity:0.4,color:#7f7f7f80
linkStyle 7 opacity:0.4,color:#7f7f7f80
linkStyle 8 opacity:0.4,color:#7f7f7f80
linkStyle 9 opacity:0.4,color:#7f7f7f80
linkStyle 10 opacity:0.4,color:#7f7f7f80
linkStyle 11 opacity:0.4,color:#7f7f7f80
linkStyle 12 opacity:0.4,color:#7f7f7f80
linkStyle 13 opacity:0.4,color:#7f7f7f80
linkStyle 14 opacity:0.4,color:#7f7f7f80
linkStyle 15 opacity:0.4,color:#7f7f7f80
Steps¶
Set up access to the S3 bucket of the cloud provider¶
DVC will need to log in to the S3 bucket of the cloud provider to download the data inside the CI/CD pipeline:
Google Cloud allows the creation of a "Service Account", so you don't have to store/share your own credentials. A Service Account can be deleted, hence revoking all the access it had.
Create the Google Service Account and its associated Google Service Account Key to access Google Cloud without your own credentials.
The key will be stored in your ~/.config/gcloud directory under the name
google-service-account-key.json:
Danger
You must never add and commit this file to your working directory. It is sensitive data that you must keep safe.
# Create the Google Service Account
gcloud iam service-accounts create google-service-account \
--display-name="Google Service Account"
# Set the permissions for the Google Service Account
gcloud projects add-iam-policy-binding $GCP_PROJECT_ID \
--member="serviceAccount:google-service-account@${GCP_PROJECT_ID}.iam.gserviceaccount.com" \
--role="roles/storage.objectViewer"
# Create the Google Service Account Key
gcloud iam service-accounts keys create ~/.config/gcloud/google-service-account-key.json \
--iam-account=google-service-account@${GCP_PROJECT_ID}.iam.gserviceaccount.com
Info
The path ~/.config/gcloud should be created when installing gcloud. If it
does not exist, you can create it by running mkdir -p ~/.config/gcloud
Store the cloud provider credentials in the CI/CD configuration¶
Now that the credentials are created, you need to store them in the CI/CD configuration. Depending on the CI/CD platform you are using, the process will be different:
Display the Google Service Account key
The service account key is stored on your computer as a JSON file. You need to display it and store it as a CI/CD variable in a text format.
Display the Google Service Account key that you have downloaded from Google Cloud:
# Display the Google Service Account key
cat ~/.config/gcloud/google-service-account-key.json
Store the Google Service Account key as a CI/CD variable
Store the output as a CI/CD variable by going to the Settings section from the top header of your GitHub repository.
Select Secrets and variables > Actions and select New repository secret.
Create a new variable named GOOGLE_SERVICE_ACCOUNT_KEY with the output value
of the Google Service Account key file as its value. Save the variable by
selecting Add secret.
Create the CI/CD pipeline configuration file¶
At the root level of your Git repository, create a GitHub Workflow configuration
file .github/workflows/mlops.yaml. Take some time to understand the train job
and its steps:
name: MLOps
on:
# Runs on pushes targeting main branch
push:
branches:
- main
# Allows you to run this workflow manually from the Actions tab
workflow_dispatch:
jobs:
train:
runs-on: ubuntu-latest
steps:
- name: Checkout repository
uses: actions/checkout@v7
- name: Setup Python
uses: actions/setup-python@v6
with:
python-version: '3.13'
cache: pip
- name: Install dependencies
run: pip install -r requirements-freeze.txt
- name: Login to Google Cloud
uses: google-github-actions/auth@v3
with:
credentials_json: '${{ secrets.GOOGLE_SERVICE_ACCOUNT_KEY }}'
- name: Train model
run: dvc repro --pull
Tip
-
Instead of running
dvc pullanddvc reproseparately, you can run them together withdvc repro --pull. -
If you prefer to use uv in CI/CD for faster dependency installation, you can replace the
Setup PythonandInstall dependenciessteps with the astral-sh/setup-uv action and runuv pip install -r requirements-freeze.txt. For simplicity, the rest of this guide continues to use the default Python tools (pipandvenv).
Push the CI/CD pipeline configuration file to Git¶
Push the CI/CD pipeline configuration file to Git:
# Add the configuration file
git add .github/workflows/mlops.yaml
# Commit the changes
git commit -m "Use a pipeline to run my experiment on each push"
# Push the changes
git push
Check the results¶
You can see the pipeline running on the Actions page.
You should see a newly created pipeline. The pipeline should log into Google
Cloud, pull the data and cached results from DVC remote storage, and reproduce
the experiment. If you encounter cache errors, verify that you have pushed all
data to DVC with dvc push.
You may have noticed that DVC was able to skip all stages as its cache is up to date. This caching mechanism allows the pipeline to validate reproducibility quickly, ensuring the experiment can be run (all data and metadata are up to date) and reproduced (the results are the same) without requiring the full computational resources needed for training from scratch.
Understanding GitHub Actions for ML Workloads
The dvc repro --pull command in this pipeline is designed to verify
reproducibility by checking cached results, not to train models from scratch.
GitHub Actions standard runners (2 CPU cores, 6-hour maximum job runtime, no GPU acceleration) are sufficient for reproducing cached DVC experiments, running tests and validation, and checking that your pipeline configuration is correct. However, these resources are not adequate for actual ML training: training large models, running GPU workloads, or executing jobs exceeding 6 hours.
In later chapters, you'll learn to overcome these limitations by using self-hosted runners on powerful cloud infrastructure with GPU support.
This chapter is done, you can check the summary.
Summary¶
Congratulations! You now have a CI/CD pipeline that will run the experiment on each commit.
In this chapter, you have successfully:
- Granted access to the S3 bucket on the cloud provider
- Stored the cloud provider credentials in the CI/CD configuration
- Created the CI/CD pipeline configuration file
- Pushed the CI/CD pipeline configuration file to Git
- Visualized the execution of the CI/CD pipeline
You fixed some of the previous issues:
- The experiment can be executed on a clean machine with the help of a CI/CD pipeline
You have a CI/CD pipeline to ensure the whole experiment can still be reproduced using the data and the commands to run using DVC over time.
Take away
- CI/CD pipelines eliminate "works on my machine" problems: By running experiments in a fresh, clean environment on every commit, you ensure that your code is truly reproducible and doesn't rely on hidden local configurations.
- Secrets management is critical for cloud access: Storing cloud credentials as CI/CD secrets (GitHub Secrets) keeps sensitive information out of your codebase while allowing automated workflows to access required resources.
- Service accounts enable secure automation: Instead of using personal credentials, cloud service accounts provide scoped access that can be safely shared with CI/CD pipelines and easily revoked if compromised.
- dvc repro --pull combines data retrieval and execution: This single command pulls data from DVC remote storage and reproduces the experiment, making CI/CD pipeline configurations simpler and more maintainable.
- Automated validation catches issues early: Running your full pipeline on every push validates that all dependencies, data, and metadata are properly tracked, catching integration issues before they reach production.
State of the MLOps process¶
- Codebase can be shared and improved by multiple developers
- Dataset can be shared among the developers and is placed in the right directory in order to run the experiment
- Experiment can be executed on a clean machine with the help of a CI/CD pipeline
- CI/CD pipeline does not report the results of the experiment
- Changes to model are not thoroughly reviewed and discussed before integration
Continue to the next chapters to address the remaining items.
Sources¶
Highly inspired by:
- Creating and managing service accounts - cloud.google.com
- Create and manage service account keys - cloud.google.com
- IAM basic and predefined roles reference - cloud.google.com
- Using service accounts - dvc.org
- Creating encrypted secrets for a repository - docs.github.com
- Triggering a workflow - docs.github.com