Skip to content

Introduction

Part 1 mission patch

Learn how to train a model locally and evaluate it using DVC DVC.

Environment

This guide has been written with Apple macOS and Linux Linux operating systems in mind. If you use Windows, please use the Windows Subsystem for Linux (WSL 2) for optimal results. See the Setup page for step-by-step instructions to prepare your machine.

Requirements

The following requirements are necessary to follow this part:

Using Conda, Anaconda, Poetry, or similar? Read this!

Why we avoid most helper tools

Conda, Anaconda, Poetry, and similar tools add abstraction layers that make debugging significantly harder when things go wrong. Their complexity introduces intricate failure modes that require substantial time and expertise to diagnose, often outweighing any convenience they provide.

This guide is tested and validated with vanilla Python only. If you still want to use Conda/Anaconda/Poetry/etc., be aware that you might encounter issues we cannot help you debug.

Our approach to Python tools

Why we use standard Python

This guide uses pip and venv as its default tools because they are part of the Python standard ecosystem, maintained by the Python Packaging Authority (PyPA) under the Python Software Foundation. They evolve slowly, but they are stable, universally available, and above all easy to debug when issues arise, which is a critical property for a tutorial.

uv: the one exception

Many steps also show the equivalent uv command as an alternative for readers who prefer it. uv is extremely fast, adheres to the relevant PEPs (Python Enhancement Proposals, the community's standardization documents), and its uv pip interface is a drop-in replacement for pip. It also fills a genuine gap in the standard toolchain: it can install and manage Python versions themselves (uv python install), not just packages within an existing interpreter.

Keep in mind that uv is developed by a private company (Astral, now part of OpenAI). While it is open source and production ready today, its long-term roadmap could shift.

For WSL2 users

Ensure you are working in the Linux filesystem, not the mounted Windows filesystem.

Run cd ~ to navigate to your Linux home directory. Avoid working in /mnt/* paths, as cross-filesystem operations cause severe performance bottlenecks when installing packages with pip.

State of the MLOps process

This guide begins with a traditional Jupyter notebook approach. You will first explore the notebook, identify its limitations for production use, and then progressively turn it into a reproducible, versioned machine learning pipeline.