Chapter 1.2 - Adapt and move the Jupyter Notebook to Python scripts¶
Introduction¶
Jupyter Notebooks provide an interactive environment where code can be executed and results can be visualized. They combine code, text explanations, visualizations, and media in a single document, making it a flexible tool to document an ML experiment.
However, they have severe limitations, such as challenges with reproducibility, scalability, experiment tracking, and standardization. Integrating Jupyter Notebooks into Python scripts suitable for running ML experiments in a more modular and reproducible manner can help address these issues and enhance the overall ML development process.
pip is the standard package manager for Python. It is used to install and manage dependencies in a Python environment.
In this chapter, you will learn how to:
- Set up a Python environment using pip
- Adapt the content of the Jupyter Notebook into Python scripts
- Launch the experiment locally
The following diagram illustrates the control flow of the experiment at the end of this chapter:
flowchart LR
subgraph workspaceGraph[WORKSPACE]
prepare[prepare.py] --> train
train[train.py] --> evaluate[evaluate.py]
params[params.yaml] -.- prepare
params -.- train
data[data/raw] --> prepare
end
style data opacity:0.4,color:#7f7f7f80
Let's get started!
Steps¶
Set up a new project directory¶
For the rest of the guide, you will work in a new directory. This will allow you to use the Jupyter Notebook directory as a reference.
Start by ensuring you have left the virtual environment created in the previous chapter:
Next, exit from the current directory and create a new one:
# Move back to the root directory
cd ..
# Create the new working directory
mkdir mlops-guide
# Switch to the new working directory
cd mlops-guide
Set up the dataset¶
You will use the same dataset as in the previous chapter. Copy the data folder
from the previous chapter to your new directory:
# Copy the data folder from the previous chapter
cp -r ../mlops-guide-notebook/data .
Set up a Python environment¶
Firstly, create the virtual environment:
Create a requirements.txt file to list the dependencies:
Install the dependencies:
Create a freeze file to list the dependencies with their versions to ensure that transitive dependencies are also listed. This will help with reproducibility:
Not familiar with freezing dependencies? Read this!
When working on Python projects, managing dependencies is crucial for maintaining a stable and reproducible development environment.
Understanding requirements.txt
The requirements.txt file is a commonly used approach to specify project
dependencies. It lists all the high-level dependencies required for your
project, including their specific versions. Each line in the file typically
follows the format: package_name==version.
Freezing dependencies
Freezing dependencies refers to fixing the versions of all transitive dependencies, ensuring that the same versions are installed consistently across different environments. This is crucial for reproducibility, as it guarantees that everyone working on the project has the exact same dependencies.
Separating high-level and transitive dependencies
To better control and manage your project's dependencies, it's beneficial to separate high-level dependencies from transitive dependencies. This approach allows for clearer identification of the core functionality packages and their required versions, ensuring a more focused and stable development environment.
-
requirements.txt: This file contains the high-level dependencies explicitly required by your project. It should include packages necessary for your project's core functionality while excluding packages that are indirectly required by other dependencies. By isolating the high-level dependencies, you maintain a clear distinction between the essential packages and the ones brought in transitively. -
requirements-freeze.txt: This file includes all the transitive dependencies required by the high-level dependencies. It ensures that all the packages needed for the project, including their versions, are recorded in a separate file. This separation allows for a more flexible and controlled approach when updating transitive dependencies while maintaining the reproducibility of your project.
How to update dependencies
When updating dependencies, it is essential to primarily modify the high-level
requirements.txt file with the desired versions or new packages. Then,
generate an updated requirements-freeze.txt file to capture the updated
transitive dependencies accurately.
Conclusion
Prioritizing stability and reproducibility in your project's dependency management is crucial for minimizing compatibility issues, avoiding unexpected bugs, and ensuring a smooth and reliable development process.
By using separate requirements files for high-level and transitive dependencies, you gain better visibility and control over the dependencies required by your project. This approach promotes a stable and reproducible development environment while allowing you to update specific packages and their versions when needed. By following these practices, you can ensure the long-term success of your Python projects.
# Freeze the dependencies
pip freeze --local --all > requirements-freeze.txt
- The
--localflag ensures that if a virtualenv has global access, it will not output globally-installed packages. - The
--allflag ensures that it does not skip these packages in the output:setuptools,wheel,pip,distribute.
Split the Jupyter Notebook into scripts¶
You will split the Jupyter Notebook in a codebase made of separate Python scripts with a well-defined role. These scripts will be able to be called on the command line, making it ideal for automation tasks.
The following table describes the files that you will create in this codebase:
| File | Description | Input | Output |
|---|---|---|---|
params.yaml |
The parameters to run the ML experiment | - | - |
src/prepare.py |
Prepare the dataset to run the ML experiment | The dataset to prepare in data/raw directory |
The prepared data in data/prepared directory |
src/train.py |
Train the ML model | The prepared dataset | The trained model in the model directory |
src/evaluate.py |
Evaluate the ML model using scikit-learn | The model to evaluate | The results of the model evaluation in evaluation directory |
src/utils/seed.py |
Util function to fix the seed | - | - |
We will refactor the notebook code into modular functions as we move each step to its own script.
Move the parameters to their own file¶
Let's split the parameters to run the ML experiment with in a distinct file:
prepare:
seed: 5241
split: 0.2
image_size: [32, 32]
grayscale: True
batch_size: 32
train:
seed: 5241
lr: 0.0001
epochs: 5
conv_size: 32
dense_size: 64
output_classes: 10
Move the preparation step to its own file¶
The src/prepare.py script prepares the dataset. It loads the raw images,
splits them into a training set and a validation set, copies the images into
data/prepared, and saves the preview plot and the class labels there.
import json
import sys
from pathlib import Path
from typing import List
import keras
import matplotlib.pyplot as plt
import tensorflow as tf
import yaml
from utils.seed import set_seed
def get_preview_plot(ds: tf.data.Dataset, labels: List[str]) -> plt.Figure:
"""Plot a preview of the prepared dataset"""
fig, axes = plt.subplots(2, 5, figsize=(10, 5), tight_layout=True)
for images, label_idxs in ds.take(1):
for ax, image, label_idx in zip(axes.ravel(), images, label_idxs):
ax.imshow(image.numpy().astype("uint8"), cmap="gray")
ax.set_title(labels[label_idx.numpy()])
ax.set_xticks([])
ax.set_yticks([])
return fig
def main() -> None:
if len(sys.argv) != 3:
print("Arguments error. Usage:\n")
print("\tpython3 prepare.py <raw-dataset-folder> <prepared-dataset-folder>\n")
exit(1)
# Load parameters
prepare_params = yaml.safe_load(open("params.yaml"))["prepare"]
raw_dataset_folder = Path(sys.argv[1])
prepared_dataset_folder = Path(sys.argv[2])
seed = prepare_params["seed"]
split = prepare_params["split"]
image_size = prepare_params["image_size"]
grayscale = prepare_params["grayscale"]
batch_size = prepare_params["batch_size"]
# Set seed for reproducibility
set_seed(seed)
# Read data
ds_train, ds_val = keras.utils.image_dataset_from_directory(
raw_dataset_folder,
labels="inferred",
label_mode="int",
color_mode="grayscale" if grayscale else "rgb",
batch_size=batch_size,
image_size=image_size,
shuffle=True,
seed=seed,
validation_split=split,
subset="both",
)
labels = ds_train.class_names
prepared_dataset_folder.mkdir(parents=True, exist_ok=True)
# Save the preview plot
preview_plot = get_preview_plot(ds_train, labels)
preview_plot.savefig(prepared_dataset_folder / "preview.png")
# Normalize the data
normalization_layer = keras.layers.Rescaling(1.0 / 255)
ds_train = ds_train.map(lambda x, y: (normalization_layer(x), y))
ds_val = ds_val.map(lambda x, y: (normalization_layer(x), y))
# Save the prepared dataset
with open(prepared_dataset_folder / "labels.json", "w") as f:
json.dump(labels, f)
tf.data.Dataset.save(ds_train, str(prepared_dataset_folder / "train"))
tf.data.Dataset.save(ds_val, str(prepared_dataset_folder / "val"))
print(f"\nDataset saved at {prepared_dataset_folder.absolute()}")
if __name__ == "__main__":
main()
Move the train step to its own file¶
The src/train.py script trains the ML model and saves it in the model
directory.
import sys
from pathlib import Path
from typing import Tuple
import keras
import numpy as np
import tensorflow as tf
import yaml
from utils.seed import set_seed
def get_model(
image_shape: Tuple[int, int, int],
conv_size: int,
dense_size: int,
output_classes: int,
) -> keras.Model:
"""Create a simple CNN model"""
model = keras.models.Sequential(
[
keras.layers.Input(shape=image_shape),
keras.layers.Conv2D(conv_size, (3, 3), activation="relu"),
keras.layers.MaxPooling2D((3, 3)),
keras.layers.Flatten(),
keras.layers.Dense(dense_size, activation="relu"),
keras.layers.Dense(output_classes),
]
)
return model
def main() -> None:
if len(sys.argv) != 3:
print("Arguments error. Usage:\n")
print("\tpython3 train.py <prepared-dataset-folder> <model-folder>\n")
exit(1)
# Load parameters
params = yaml.safe_load(open("params.yaml"))
prepare_params = params["prepare"]
train_params = params["train"]
prepared_dataset_folder = Path(sys.argv[1])
model_folder = Path(sys.argv[2])
image_size = prepare_params["image_size"]
grayscale = prepare_params["grayscale"]
image_shape = (*image_size, 1 if grayscale else 3)
seed = train_params["seed"]
lr = train_params["lr"]
epochs = train_params["epochs"]
conv_size = train_params["conv_size"]
dense_size = train_params["dense_size"]
output_classes = train_params["output_classes"]
# Set seed for reproducibility
set_seed(seed)
# Load the prepared datasets and shuffle the training set at each epoch
ds_train = tf.data.Dataset.load(str(prepared_dataset_folder / "train"))
ds_train = ds_train.shuffle(
buffer_size=ds_train.cardinality(), seed=seed, reshuffle_each_iteration=True
)
ds_val = tf.data.Dataset.load(str(prepared_dataset_folder / "val"))
# Define the model
model = get_model(image_shape, conv_size, dense_size, output_classes)
model.compile(
optimizer=keras.optimizers.Adam(lr),
loss=keras.losses.SparseCategoricalCrossentropy(from_logits=True),
metrics=[keras.metrics.SparseCategoricalAccuracy()],
)
model.summary()
# Train the model
model.fit(
ds_train,
epochs=epochs,
validation_data=ds_val,
)
# Save the model
model_folder.mkdir(parents=True, exist_ok=True)
model_path = model_folder.absolute() / "model.keras"
model.save(model_path)
# Save the model history
np.save(model_folder.absolute() / "history.npy", model.history.history)
print(f"\nModel saved at {model_folder.absolute()}")
if __name__ == "__main__":
main()
Move the evaluate step to its own file¶
The src/evaluate.py script evaluates the ML model and saves the metrics and
graphs in the evaluation directory.
import json
import sys
from pathlib import Path
from typing import List
import keras
import matplotlib.pyplot as plt
import numpy as np
import tensorflow as tf
from sklearn.metrics import (
ConfusionMatrixDisplay,
f1_score,
precision_score,
recall_score,
)
def get_training_plot(model_history: dict) -> plt.Figure:
"""Plot the training and validation loss"""
epochs = range(1, len(model_history["loss"]) + 1)
fig = plt.figure(figsize=(10, 4))
plt.plot(epochs, model_history["loss"], label="Training loss")
plt.plot(epochs, model_history["val_loss"], label="Validation loss")
plt.xticks(epochs)
plt.title("Training and validation loss")
plt.xlabel("Epochs")
plt.ylabel("Loss")
plt.legend()
plt.grid(True)
return fig
def get_pred_preview_plot(
model: keras.Model, ds_val: tf.data.Dataset, labels: List[str]
) -> plt.Figure:
"""Plot a preview of the predictions"""
fig, axes = plt.subplots(2, 5, figsize=(10, 5), tight_layout=True)
for images, label_idxs in ds_val.take(1):
pred_idxs = np.argmax(model.predict(images, verbose=0), axis=1)
for ax, image, true_idx, pred_idx in zip(
axes.ravel(), images, label_idxs, pred_idxs
):
true_label = labels[true_idx.numpy()]
pred_label = labels[pred_idx]
ax.imshow(image.numpy().squeeze(), cmap="gray")
ax.set_title(f"True: {true_label}\nPred: {pred_label}")
ax.set_xticks([])
ax.set_yticks([])
border_color = "lime" if true_idx.numpy() == pred_idx else "red"
for spine in ax.spines.values():
spine.set_edgecolor(border_color)
spine.set_linewidth(4)
return fig
def get_confusion_matrix_plot(
y_true: np.ndarray, y_pred: np.ndarray, labels: List[str]
) -> plt.Figure:
"""Plot the confusion matrix"""
fig, ax = plt.subplots(figsize=(6, 6), tight_layout=True)
display = ConfusionMatrixDisplay.from_predictions(
y_true,
y_pred,
display_labels=labels,
normalize="true",
cmap="Blues",
values_format=".2f",
ax=ax,
colorbar=True,
)
for value, text in zip(display.confusion_matrix.ravel(), display.text_.ravel()):
text.set_fontsize(7)
if np.isclose(value, 0.0):
text.set_color("lightgray")
ax.set_xticklabels(labels, rotation=90)
ax.set_title("Validation confusion matrix")
return fig
def main() -> None:
if len(sys.argv) != 3:
print("Arguments error. Usage:\n")
print("\tpython3 evaluate.py <model-folder> <prepared-dataset-folder>\n")
exit(1)
model_folder = Path(sys.argv[1])
prepared_dataset_folder = Path(sys.argv[2])
evaluation_folder = Path("evaluation")
plots_folder = Path("plots")
# Create folders
(evaluation_folder / plots_folder).mkdir(parents=True, exist_ok=True)
# Load files
ds_val = tf.data.Dataset.load(str(prepared_dataset_folder / "val"))
with open(prepared_dataset_folder / "labels.json") as f:
labels = json.load(f)
# Load model
model_path = model_folder.absolute() / "model.keras"
model = keras.models.load_model(model_path)
model_history = np.load(
model_folder.absolute() / "history.npy", allow_pickle=True
).item()
# Log metrics
val_loss, val_acc = model.evaluate(ds_val)
preds = model.predict(ds_val)
y_true = tf.concat([y for _, y in ds_val], axis=0).numpy()
y_pred = np.argmax(preds, axis=1)
metrics = {
"val_loss": val_loss,
"val_acc": val_acc,
"precision": precision_score(y_true, y_pred, average="macro", zero_division=0),
"recall": recall_score(y_true, y_pred, average="macro", zero_division=0),
"f1_score": f1_score(y_true, y_pred, average="macro", zero_division=0),
}
print(f"Validation loss: {metrics['val_loss']:.2f}")
print(f"Validation accuracy: {metrics['val_acc'] * 100:.2f}%")
print(f"Precision: {metrics['precision']:.2f}")
print(f"Recall: {metrics['recall']:.2f}")
print(f"F1 score: {metrics['f1_score']:.2f}")
with open(evaluation_folder / "metrics.json", "w") as f:
json.dump(metrics, f)
# Save training history plot
fig = get_training_plot(model_history)
fig.savefig(evaluation_folder / plots_folder / "training_history.png")
# Save predictions preview plot
fig = get_pred_preview_plot(model, ds_val, labels)
fig.savefig(evaluation_folder / plots_folder / "pred_preview.png")
# Save confusion matrix plot
fig = get_confusion_matrix_plot(y_true, y_pred, labels)
fig.savefig(evaluation_folder / plots_folder / "confusion_matrix.png")
print(
f"\nEvaluation metrics and plot files saved at {evaluation_folder.absolute()}"
)
if __name__ == "__main__":
main()
Create the seed helper function¶
Finally, add a module for utils:
# Create the utils module
mkdir src/utils
# Create the __init__.py file to make the utils module a package
touch src/utils/__init__.py
In this module, include src/utils/seed.py to handle the fixing of the seed
parameters. This ensures the results are reproducible:
import os
import keras
import tensorflow as tf
def set_seed(seed: int) -> None:
"""Set all random seeds and enable deterministic operations"""
os.environ["PYTHONHASHSEED"] = str(seed)
keras.utils.set_random_seed(seed)
tf.config.experimental.enable_op_determinism()
tf.config.threading.set_inter_op_parallelism_threads(1)
tf.config.threading.set_intra_op_parallelism_threads(1)
Understanding when to fix random seeds
Fixed seeds make experiments reproducible, which helps compare changes fairly. They also hide natural variance in model performance.
- Fix seeds during development, CI/CD runs, and debugging so differences come from your changes, not randomness.
- Avoid a single fixed seed for final evaluation. Train with multiple seeds and report mean and standard deviation to measure stability.
- Keep seeds configurable (as in
params.yaml) instead of hardcoding them.
Create a README.md file¶
Finally, create a README.md file at the root of the project to describe the
repository. Feel free to use the following template. As you progress through
this guide, you can add your notes in the ## Notes section:
# MLOps - Celestial Body Classification
This repository contains the code from
[A guide to MLOps](https://mlops.swiss-ai-center.ch/).
## Notes
<!-- Enter your notes below -->
Check the results¶
Your working directory should now look like this:
.
โโโ data
โ โโโ raw
โ โ โโโ ...
โ โโโ README.md
โโโ params.yaml # (1)!
โโโ README.md # (2)!
โโโ requirements-freeze.txt # (3)!
โโโ requirements.txt # (4)!
โโโ src # (5)!
โโโ evaluate.py
โโโ prepare.py
โโโ train.py
โโโ utils
โโโ __init__.py
โโโ seed.py
- This is new.
- This is new.
- This is new.
- This is new.
- This, and all its sub-directories, is new.
Run the experiment¶
Awesome! You now have everything you need to run the experiment: the codebase and the dataset are in place, the new virtual environment is set up, and you are ready to run the experiment using scripts for the first time.
You can now follow these steps to reproduce the experiment:
# Prepare the dataset
python3.13 src/prepare.py data/raw data/prepared
# Train the model with the training set and save it
python3.13 src/train.py data/prepared model
# Evaluate the model performance
python3.13 src/evaluate.py model data/prepared
The experiment will take some time to run. Once it is done, you will find the
results in the data/prepared, model, and evaluation directories.
Check the results¶
Your working directory should now be similar to this:
.
โโโ data
โ โโโ prepared # (1)!
โ โ โโโ labels.json
โ โ โโโ preview.png
โ โ โโโ train
โ โ โ โโโ ...
โ โ โโโ val
โ โ โโโ ...
โ โโโ raw
โ โ โโโ ...
โ โโโ README.md
โโโ evaluation # (2)!
โ โโโ metrics.json
โ โโโ plots
โ โโโ confusion_matrix.png
โ โโโ pred_preview.png
โ โโโ training_history.png
โโโ model # (3)!
โ โโโ history.npy
โ โโโ model.keras
โโโ params.yaml
โโโ README.md
โโโ requirements-freeze.txt
โโโ requirements.txt
โโโ src
โโโ evaluate.py
โโโ prepare.py
โโโ train.py
โโโ utils
โโโ __init__.py
โโโ seed.py
- This, and all its sub-directories, is new.
- This, and all its sub-directories, is new.
- This is new.
Here, the following should be noted:
- the
prepare.pyscript created thedata/prepareddirectory and divided the dataset into a training set and a validation set - the
train.pyscript created themodeldirectory and trained the model with the prepared data. - the
evaluate.pyscript created theevaluationdirectory and generated some plots and metrics to evaluate the model
Take some time to get familiar with the scripts and the results.
Summary¶
Congratulations! You have successfully reproduced the experiment on your machine, this time using a modular approach that can be put into production.
In this chapter, you have:
- Set up a Python virtual environment
- Adapted the content of the Jupyter Notebook into Python scripts
- Launched the experiment locally
You fixed some of the previous issues:
- Notebook has been transformed into scripts for production
Take away
- Modular Python scripts enable production-ready ML workflows: Converting
notebooks to separate
prepare.py,train.py, andevaluate.pyscripts creates clear separation of concerns and makes code easier to test, maintain, and integrate into automated pipelines. - Parameters should be externalized: Using a
params.yamlfile to store configuration separates logic from configuration, making it easy to experiment with different hyperparameters without modifying code. - Dependency management requires two levels:
requirements.txttracks high-level dependencies you explicitly need, whilerequirements-freeze.txtcaptures all transitive dependencies with exact versions for reproducibility. - Command-line interfaces make automation possible: Scripts that accept arguments and can be run from the terminal integrate seamlessly with scheduling tools, CI/CD pipelines, and orchestration systems.
- Fixed seeds enable reproducibility but hide variance: Using fixed seeds during development ensures fair comparison of model changes, while final evaluation should use multiple seeds to assess real-world model stability.
State of the MLOps process¶
- Notebook has been transformed into scripts for production
- Codebase and dataset are not versioned
- Model steps rely on verbal communication and may be undocumented
- Changes to model are not easily visualized
Continue to the next chapters to address the remaining items.
Sources¶
Highly inspired by: