Chapter 1.3 - Initialize Git and DVC for local training¶
Introduction¶
Now that you have a good understanding of the experiment, it's time to improve the code and data sharing process. To share the codebase, you will create a Git repository.
However, when it comes to managing large files, Git has some limitations. Although Git LFS is an option for handling large files in Git repositories, it may not be the most efficient solution.
This is the reason you will use DVC, a version control system specifically designed to share the data and integrates well with Git. DVC utilizes chunking to efficiently store large files and track their changes.
In this chapter, you will learn how to:
- Set up a new Git repository
- Initialize Git in your project directory
- Verify Git tracking for your files
- Exclude experiment results, data, models and Python environment files from Git commits
- Commit your changes to the Git repository
- Install DVC
- Initialize and configure DVC
- Update the gitignore file and add the experiment data to DVC
- Push the data files to DVC
- Commit the metadata files to Git
The following diagram illustrates the control flow of the experiment at the end of this chapter:
flowchart TB
dot_dvc[(.dvc)]
dot_git[(.git)]
data[data/raw] <-.-> dot_dvc
workspaceGraph <-....-> dot_git
subgraph cacheGraph[CACHE]
dot_dvc
dot_git
end
subgraph workspaceGraph[WORKSPACE]
data --> prepare
prepare[prepare.py] --> train
train[train.py] --> evaluate[evaluate.py]
params[params.yaml] -.- prepare
params -.- train
end
style workspaceGraph opacity:0.4,color:#7f7f7f80
style prepare opacity:0.4,color:#7f7f7f80
style train opacity:0.4,color:#7f7f7f80
style evaluate opacity:0.4,color:#7f7f7f80
style params opacity:0.4,color:#7f7f7f80
linkStyle 2 opacity:0.4,color:#7f7f7f80
linkStyle 3 opacity:0.4,color:#7f7f7f80
linkStyle 4 opacity:0.4,color:#7f7f7f80
linkStyle 5 opacity:0.4,color:#7f7f7f80
linkStyle 6 opacity:0.4,color:#7f7f7f80
In future chapters, you will improve the code sharing process by setting up remote Git and DVC repositories to enable easy collaboration with the rest of the team.
Steps¶
Create a new Git repository¶
Initialize Git in your working directory¶
Use the following command to set up a local Git repository in your working directory:
# Initialize Git in your working directory with `main` as the initial branch
git init --initial-branch=main
First-time Git setup? Read this!
If this is your first time using Git on this system (or WSL2 distribution), you need to configure your Git identity. This is required for making commits and should match your GitHub account information:
# Set your Git username (should match your GitHub username)
git config --global user.name "Your Name"
# Set your Git email (should match your GitHub email)
git config --global user.email "[email protected]"
Additionally, configure line ending handling to avoid cross-platform issues:
# Convert CRLF to LF on commit
git config --global core.autocrlf input
You can verify your configuration with:
Check if Git tracks your files¶
Initialize Git in your working directory. Verify available files for committing with this command:
The output should be similar to this:
On branch main
No commits yet
Untracked files:
(use "git add <file>..." to include in what will be committed)
.venv/
README.md
data/...
evaluation/...
model/...
params.yaml
requirements-freeze.txt
requirements.txt
src/__pycache__/...
src/evaluate.py
src/prepare.py
src/train.py
src/utils/__init__.py
src/utils/__pycache__/...
src/utils/seed.py
nothing added to commit but untracked files present (use "git add" to track)
As you can see, no files have been added to Git yet.
Create a .gitignore file¶
Create a .gitignore file to exclude data, models, and Python environment to
improve repository size and clone time. The data and models will be managed by
DVC in the next chapters. Keep the model's evaluation as it doesn't take much
space and you can have a history of the improvements made to your model.
Additionally, this will help to ensure that the repository size and clone time
remain optimized:
# Data used to train the models
data/
# Evaluation results
evaluation/
# The models
model/
## Python
.venv/
# Byte-compiled / optimized / DLL files
__pycache__/
Info
If using macOS, you might want to ignore .DS_Store files as well to avoid
pushing Apple's metadata files to your repository.
Check the changes¶
Check the changes with Git to ensure all wanted files are here with the following commands:
# Add all the available files
git add .
# Check the changes
git status
The output of the git status command should be similar to this:
On branch main
No commits yet
Changes to be committed:
(use "git rm --cached <file>..." to unstage)
new file: .gitignore
new file: README.md
new file: params.yaml
new file: requirements-freeze.txt
new file: requirements.txt
new file: src/evaluate.py
new file: src/prepare.py
new file: src/train.py
new file: src/utils/__init__.py
new file: src/utils/seed.py
Commit the changes¶
Commit the changes to Git:
# Commit the changes
git commit -m "Use Git to version my ML experiment"
Create a DVC repository¶
Install DVC¶
Add the main dvc dependency to the requirements.txt file:
Check the differences with Git to validate the changes:
# Show the differences with Git
git diff requirements.txt
The output should be similar to this:
diff --git a/requirements.txt b/requirements.txt
index bf56600..116c388 100644
--- a/requirements.txt
+++ b/requirements.txt
@@ -2,3 +2,4 @@ tensorflow==2.21.0
matplotlib==3.11.0
scikit-learn==1.9.0
pyyaml==6.0.3
+dvc==3.67.1
Install the dependencies and update the freeze file:
Warning
Prior to running any pip commands, it is crucial to ensure the virtual environment is activated to avoid potential conflicts with system-wide Python packages.
To check its status, simply run which python. If the virtual environment is
active, the output will show the path to the virtual environment's Python
executable. If it is not, you can activate it with source .venv/bin/activate.
Initialize DVC¶
Initialize DVC in the current project.
The dvc init command creates a .dvc directory in the working directory,
which serves as the configuration directory for DVC.
Info
By default, DVC collects anonymized usage analytics to help improve the tool. If you prefer to opt out, you can disable analytics with the following command:
Update the .gitignore file and add the experiment data to DVC¶
With DVC now set up, you can begin adding files to it. The dvc add command
creates a data/raw.dvc file and a data/.gitignore. The .dvc file contains
the metadata DVC uses to download the files and check their integrity. The
.gitignore file tells Git to ignore the files in data/raw.
Try to add the experiment data. Spoiler, it will fail:
# Try to add the experiment data to DVC
dvc add data/raw/
When executing this command, the following output occurs:
DVC tried to create data/raw.dvc, but the data/ rule blocks it. That rule
also keeps data/prepared out of Git, and you still want that: you are only
handing data/raw to DVC for now. Replace data/ with data/prepared/:
# Data used to train the models
data/prepared/
# Evaluation results
evaluation/
# The models
model/
## Python
.venv/
# Byte-compiled / optimized / DLL files
__pycache__/
Info
If using macOS, you might want to ignore .DS_Store files as well to avoid
pushing Apple's metadata files to your repository.
Check the differences with Git to validate the changes:
The output should be similar to this:
diff --git a/.gitignore b/.gitignore
index dc17ed7..1c13140 100644
--- a/.gitignore
+++ b/.gitignore
@@ -1,5 +1,5 @@
# Data used to train the models
-data/
+data/prepared/
# Evaluation results
evaluation/
You can now add the experiment data to DVC without complaint:
The output should be similar to this. You can safely ignore the message:
To track the changes with git, run:
git add data/raw.dvc data/.gitignore
To enable auto staging, run:
dvc config core.autostage true
As expected, DVC created data/raw.dvc and data/.gitignore. Both must be
added to Git.
Various DVC commands will automatically try to update the gitignore files. If a
gitignore file is already present, it will be updated to include the newly
ignored files. As you hand more files and directories to DVC, it will extend
these gitignore files for you. You might need to update existing .gitignore
files accordingly.
Check the changes¶
Check the changes with Git to ensure all wanted files are here.
# Add all the files
git add .
# Check the changes
git status
The output of the git status command should be similar to this.
On branch main
Changes to be committed:
(use "git restore --staged <file>..." to unstage)
new file: .dvc/.gitignore
new file: .dvc/config
new file: .dvcignore
modified: .gitignore
new file: data/.gitignore
new file: data/README.md
new file: data/raw.dvc
modified: requirements-freeze.txt
modified: requirements.txt
Commit the changes to Git¶
You can now commit the changes to Git so the data from DVC is tracked along code changes as well.
# Commit the changes
git commit -m "Use DVC to version the data in my ML experiment"
This chapter is done, you can check the summary.
Summary¶
Congratulations! You now have a codebase and a dataset that is versioned with Git and DVC. At the moment, these tools are only used locally. In the next chapters, you will learn how to share the codebase and the dataset with the rest of the team.
In this chapter, you have successfully:
- Set up a new Git repository
- Initialized Git in your project directory
- Verified Git tracking for your files
- Excluded experiment results, data, models and Python environment files from Git commits
- Committed your changes to the Git repository
- Installed DVC
- Initialized DVC
- Updated the gitignore file and added the experiment data to DVC
- Committed the data files to DVC
- Committed your changes to the Git repository
You fixed some of the previous issues:
- Dataset no longer needs manual download and is placed in the right directory.
- Codebase and dataset are versioned
Take away
- Git and DVC solve different version control problems: Git excels at tracking code changes with its line-by-line diffs, while DVC efficiently handles large files (datasets, models) using chunking and content-addressable storage without bloating your repository.
- The
.dvcfiles are metadata, not data: When youdvc adda file, DVC creates a small.dvcmetadata file that gets committed to Git, while the actual data is stored in the DVC cache. This separation keeps repositories lightweight. - Smart
.gitignoreconfiguration is essential: Configure the main.gitignoreto exclude only what DVC does not manage yet, and let DVC maintain its own.gitignorefiles for the paths it tracks. - Local versioning enables future collaboration: Setting up Git and DVC locally is the foundation. In later chapters, you'll push to remote repositories to enable team collaboration and experiment tracking.
State of the MLOps process¶
- Notebook has been transformed into scripts for production
- Codebase and dataset are versioned
- Model steps rely on verbal communication and may be undocumented
- Changes to model are not easily visualized
Continue to the next chapters to address the remaining items.
Sources¶
Highly inspired by: