OPEN SOURCE · CASE STUDY
Reproducible ML experiments
A practical machine learning pipeline, from raw data to versioned experiments with DVC, XGBoost and CatBoost.
The idea
An end-to-end churn prediction example. Data preparation, model training and evaluation are tracked as a reproducible pipeline.
A reproducible loop
Data preparation, preprocessing, model training and evaluation are defined as explicit stages. The original walkthrough explains how Git and DVC keep the code, datasets and model outputs in sync.