Operations 60 min read

Beyond v2_final: Production-Grade ML Data Versioning & Experiment Tracking with DVC + MLflow

This article demonstrates a complete, reproducible ML pipeline using DVC for data versioning and MLflow for experiment tracking, covering immutable data snapshots, time-based splits, pipeline definitions, model registration gates, canary deployment, rollback procedures, CI integration, and retention policies — all validated with concrete code and failure-mode analysis.

Ray's Galactic Tech
Ray's Galactic Tech
Ray's Galactic Tech
Beyond v2_final: Production-Grade ML Data Versioning & Experiment Tracking with DVC + MLflow

1. Why File-Name Management Fails at Release

A Monday release of a click-prediction model shows inflated CTR on a new channel by Wednesday. The algorithm engineer says they used "last week's data", the data engineer says they only fixed one SQL query, and the platform engineer can only locate features_v2_final.csv. All three statements may be true, yet no one can answer: which exact byte content did the model read? When was the feature query executed? Were labels mature? Which run passed which evaluation set?

Retraining old code does not guarantee the old model. The database has received late-arriving events, historical features were backfilled, and random splits mixed the same time window into the validation set. Even if all files remain, the combination relationships at training time may be lost.

We need a verifiable chain: Data Snapshot → Input Code Commit → Parameters & Runtime Environment → Training Run → Model Version → Release Record . Rollback should restore the verified model artifact; retraining is a separate process.

The article first provides a complete minimal loop without Kubernetes, then explains team collaboration and production scaling. Examples use CSV to lower startup cost; production data can swap to partitioned Parquet but must not omit snapshot identity.

2. Clarifying Responsibilities: Git, DVC, and MLflow

Git stores code, parameters, DVC metadata, lock files. It solves: which commit defined which inputs and outputs. It does not automatically guarantee remote availability of large files.

DVC stores data, preparation results, model directories via content-addressable cache. It solves: restore files by pointer, recompute pipeline by dependencies. It does not automatically guarantee data semantic correctness, cross-service transactions, or service traffic switching.

MLflow stores parameters, metrics, run status, model artifacts, Registry versions. It solves: compare experiments, trace model provenance. It does not automatically guarantee no data leakage, online revenue, or automatic safe rollback.

Release Controller stores fixed model version, artifact verification, instance status, change log. It solves: promote candidate model to production instance. It does not automatically guarantee a single alias updates all services.

DVC 3 cache typically lives in .dvc/cache/files/md5/. Regular files are content-addressed; directories use file manifests. Deduplication granularity is usually per-file — not block-level incremental storage for arbitrary large files. A 100 GB file changed by one line may produce a new full file object. MD5 here serves content identity and change detection; untrusted inputs still need signatures, permissions, and stronger integrity checks. dvc.api.get_url() returns a location convenient for download but cannot serve as version identity alone. This article records Git commit, DVC raw pointer, data SHA-256, parameter digest, and split rule. URLs do not enter the lineage primary key, avoiding temporary signed URLs or credentials in logs.

3. What a Production Data Snapshot Must Contain

Real click data usually comes from impression, click, and historical behavior tables. A file hash only proves file sameness; it cannot prove SQL correctness or that features were available at impression time.

snapshot_id : Unique identifier for a completed export

source_table_snapshot / object_version : Bind to source table transaction snapshot or object version

extraction_query_sha256 : Identify query logic changes

event_time_start / end : Define business event window

extracted_at / watermark : Describe extraction moment and late-event boundary

label_definition / maturity_window : Explain click attribution and label wait time

feature_schema_version : Lock feature meaning, order, and types

row_count / partition_checksums : Verify export completeness

privacy_policy_id : Constrain retention and sensitive field handling

Extraction jobs write to a temporary prefix first, verify, then write a completion manifest; training only consumes snapshots with a completion manifest. A SELECT ... WHERE date=... in the data warehouse changes with backfills and cannot replace an immutable snapshot.

The example's hist_ctr is generated directly from random numbers. Real business must compute from historical windows before impression time using point-in-time joins. Time splitting prevents future samples from leaking into training, but cannot fix features that have already leaked future information . Click labels should wait for the attribution window to mature; immature impressions must not be treated as non-clicks.

4. Minimal Engineering: Copy Files and Run

4.1 Versions and Directory Structure

Uses Python 3.12, Git, DVC 3.59.2, MLflow 2.22.0. Additional constraints: pathspec==0.12.1 to avoid old DVC vs pathspec 1.x interface conflict; SQLAlchemy==2.0.40 to avoid MLflow 2.22 vs SQLAlchemy 2.1 conflict. Only major dependencies and discovered compatibility constraints are pinned; full transitive dependencies still need a lock file. The 2.x baseline ensures API and example consistency — not a long-term production security baseline. Teams should lock dependencies, scan vulnerabilities, and validate compatibility at image build. Upgrading to MLflow 3.x requires re-validating model logging API and server config.

Project contains src/, data/raw/, data/prepared/, models/, reports/. dvc init generates .dvc/ which must enter Git. Derived product ignore rules are added by DVC.

requirements.txt

dvc[s3]==3.59.2
mlflow==2.22.0
numpy==2.2.5
pandas==2.2.3
scikit-learn==1.6.1
PyYAML==6.0.2
pathspec==0.12.1
SQLAlchemy==2.0.40

.gitignore

.venv/
__pycache__/
mlruns/
mlartifacts/
mlflow.db
mlflow.db-*
.env
runtime/

params.yaml

prepare:
  split_time: "2026-01-22T00:00:00Z"
  min_rows: 1000
train:
  n_estimators: 80
  learning_rate: 0.05
  max_depth: 3
  random_state: 42
gate:
  min_auc: 0.70
  max_log_loss: 0.70

src/common.py

import hashlib
import json
import subprocess
from pathlib import Path

FEATURES = ["hist_ctr", "price", "hour"]

def sha256(path):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        for block in iter(lambda: f.read(1024 * 1024), b""):
            h.update(block)
    return h.hexdigest()

def git(*args):
    return subprocess.check_output(["git", *args], text=True).strip()

def write_json(path, value):
    p = Path(path)
    p.parent.mkdir(parents=True, exist_ok=True)
    tmp = p.with_suffix(p.suffix + ".tmp")
    tmp.write_text(json.dumps(value, indent=2, ensure_ascii=False,
                              allow_nan=False), encoding="utf-8")
    tmp.replace(p)

4.2 Generate Reproducible Demo Data

12,000 impressions grow over time; labels generated from three features. ID, time, and label do not enter model features. Each generation uses a fixed seed. This scale verifies the engineering loop; cannot infer TB-level processing performance.

src/make_demo.py

from pathlib import Path
import numpy as np
import pandas as pd

rng = np.random.default_rng(42)
n = 12000
hours = rng.integers(0, 24, n)
ctr = rng.uniform(0.01, 0.50, n)
price = rng.uniform(5, 300, n)
p = 1 / (1 + np.exp(-(-2 + 8 * ctr - 0.004 * price
        + 0.4 * (hours >= 18))))
df = pd.DataFrame({
    "event_id": [f"e{i:08d}" for i in range(n)],
    "event_time": pd.date_range("2026-01-01", periods=n, freq="4min", tz="UTC"),
    "hist_ctr": ctr, "price": price, "hour": hours,
    "label": rng.binomial(1, p),
})
Path("data/raw").mkdir(parents=True, exist_ok=True)
df.to_csv("data/raw/events.csv", index=False)
print(f"demo rows={len(df)}")

4.3 Validation and Time Split

Missing columns, duplicate impression IDs, illegal labels, nulls, infinite values, out-of-range values, too few samples, or single-label sets all cause preparation to fail. Failure stops training. Silent zero-filling or dropping bad rows while only reporting remaining samples is forbidden because it masks upstream quality issues.

Train/validation split by fixed UTC time. ID sorting ensures stable output order for same time. Production should add user-group evaluation, channel stratification, label maturity checks, and feature distribution comparison as needed.

src/prepare.py

import json
from pathlib import Path
import numpy as np
import pandas as pd
import yaml
from common import FEATURES, sha256, write_json

cfg = yaml.safe_load(Path("params.yaml").read_text())["prepare"]
raw = Path("data/raw/events.csv")
df = pd.read_csv(raw)
required = ["event_id", "event_time", *FEATURES, "label"]
if set(df.columns) != set(required):
    raise ValueError("schema mismatch")
if len(df) < cfg["min_rows"] or df[required].isna().any().any():
    raise ValueError("too few rows or null values")
if df.event_id.duplicated().any():
    raise ValueError("duplicate event_id")
for col in [*FEATURES, "label"]:
    df[col] = pd.to_numeric(df[col], errors="raise")
if not np.isfinite(df[[*FEATURES, "label"]].to_numpy()).all():
    raise ValueError("non-finite value")
if not df.label.isin([0, 1]).all():
    raise ValueError("invalid label")
if not df.hist_ctr.between(0, 1).all() or not (df.price >= 0).all():
    raise ValueError("feature range violation")
if not df.hour.between(0, 23).all() or not (df.hour % 1 == 0).all():
    raise ValueError("invalid hour")
df.event_time = pd.to_datetime(df.event_time, utc=True, errors="raise")
if df.event_time.isna().any():
    raise ValueError("invalid event_time")
df = df.sort_values(["event_time", "event_id"])
cut = pd.Timestamp(cfg["split_time"])
train = df[df.event_time < cut]
valid = df[df.event_time >= cut]
for name, part in [("train", train), ("valid", valid)]:
    if len(part) < 100 or part.label.nunique() != 2:
        raise ValueError(f"invalid {name} partition")
Path("data/prepared").mkdir(parents=True, exist_ok=True)
train.to_csv("data/prepared/train.csv", index=False)
valid.to_csv("data/prepared/valid.csv", index=False)
manifest = {
    "schema_version": 1, "raw_sha256": sha256(raw),
    "train_sha256": sha256("data/prepared/train.csv"),
    "valid_sha256": sha256("data/prepared/valid.csv"),
    "split_time": cfg["split_time"], "feature_order": FEATURES,
    "train_rows": len(train), "valid_rows": len(valid),
    "train_label_rate": float(train.label.mean()),
    "valid_label_rate": float(valid.label.mean()),
    "train_time_max": train.event_time.max().isoformat(),
    "valid_time_min": valid.event_time.min().isoformat(),
}
write_json("data/prepared/manifest.json", manifest)
print(json.dumps(manifest, indent=2))

4.4 Define DVC Pipeline

Parameters declared via params. deps list real files or directories; do not write params.yaml:train as a dependency path. Raw data managed by dvc add; preparation results and model outputs managed by stages; two stages cannot own the same output directory. reports/run.json uses cache output because it relates to that training. metrics.json not cached, so it enters Git and code review. dvc.lock generated by DVC, records actual dependencies, parameters, and output hashes — not handwritten.

dvc.yaml

stages:
  prepare:
    cmd: python src/prepare.py
    deps:
      - src/prepare.py
      - src/common.py
      - requirements.txt
      - data/raw/events.csv
    params:
      - prepare
    outs:
      - data/prepared
  train:
    cmd: python src/train.py
    deps:
      - src/train.py
      - src/common.py
      - requirements.txt
      - data/prepared
    params:
      - train
    outs:
      - models
      - reports/run.json
    metrics:
      - reports/metrics.json:
        cache: false

4.5 Training and Lineage Recording

Before training, require source code, parameters, and raw data pointers committed. dvc.lock and metrics changing during training is normal; cannot judge source untrustworthy from whole workspace dirty status. source_commit points to the training input commit. After training, commit dvc.lock and output pointers to get a release_commit. Keep them separate to avoid mistaking post-training commits for the training-time code version.

Model saved first to DVC-managed local directory, then uploaded to MLflow. Service loads explicit model version. Example model signature describes predict output; smoke test uses sklearn flavor's predict_proba for click probability. If using generic pyfunc service expecting probabilities, wrap a probability-output model and redefine signature — cannot assume default pyfunc returns probabilities.

src/train.py

import hashlib
import json
import os
import platform
import subprocess
from pathlib import Path
import mlflow
import mlflow.sklearn
import pandas as pd
import yaml
from mlflow.models import infer_signature
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import roc_auc_score, log_loss
from common import FEATURES, git, sha256, write_json

# Allow DVC-generated lock/metrics changes; forbid uncommitted training source, config, raw pointers.
source_paths = ["src", "params.yaml", "requirements.txt", "dvc.yaml",
                "data/raw/events.csv.dvc"]
if git("status", "--porcelain", "--", *source_paths):
    raise RuntimeError("commit all training inputs first")
commit = git("rev-parse", "HEAD")
params = yaml.safe_load(Path("params.yaml").read_text())
manifest = json.loads(Path("data/prepared/manifest.json").read_text())
for name in ["train", "valid"]:
    if sha256(f"data/prepared/{name}.csv") != manifest[f"{name}_sha256"]:
        raise RuntimeError("prepared data hash mismatch")
raw_pointer = yaml.safe_load(Path("data/raw/events.csv.dvc").read_text())
if sha256("data/raw/events.csv") != manifest["raw_sha256"]:
    raise RuntimeError("raw data hash mismatch")
with open("data/raw/events.csv", "rb") as f:
    h = hashlib.md5()
    for chunk in iter(lambda: f.read(1024 * 1024), b""):
        h.update(chunk)
    if h.hexdigest() != raw_pointer["outs"][0]["md5"]:
        raise RuntimeError("raw file differs from committed DVC pointer")
train = pd.read_csv("data/prepared/train.csv")
valid = pd.read_csv("data/prepared/valid.csv")
identity = {
    "source_commit": commit,
    "params_sha256": sha256("params.yaml"),
    "requirements_sha256": sha256("requirements.txt"),
    "manifest_sha256": sha256("data/prepared/manifest.json"),
    "raw_dvc_out": raw_pointer["outs"][0],
    "dataset": manifest, "python_version": platform.python_version(),
}
mlflow.set_tracking_uri(os.environ["MLFLOW_TRACKING_URI"])
mlflow.set_experiment("ctr-demo-v2")
with mlflow.start_run() as run:
    mlflow.set_tags({
        "source_commit": commit,
        "raw_sha256": manifest["raw_sha256"],
        "valid_sha256": manifest["valid_sha256"],
        "purpose": "candidate"
    })
    mlflow.log_params(params["train"])
    mlflow.log_dict(identity, "lineage/input.json")
    mlflow.log_artifact("params.yaml", "lineage")
    mlflow.log_artifact("requirements.txt", "lineage")
    freeze = subprocess.check_output(
        [os.sys.executable, "-m", "pip", "freeze"], text=True)
    mlflow.log_text(freeze, "lineage/environment.txt")
    model = GradientBoostingClassifier(**params["train"])
    model.fit(train[FEATURES], train.label)
    prob = model.predict_proba(valid[FEATURES])[:, 1]
    metrics = {
        "val_auc": float(roc_auc_score(valid.label, prob)),
        "val_log_loss": float(log_loss(valid.label, prob))}
    mlflow.log_metrics(metrics)
    Path("models").mkdir(exist_ok=True)
    model_path = Path("models/model")
    # When called by DVC, output changes handled by DVC first.
    # Direct repeat runs should not overwrite old model directory.
    mlflow.sklearn.save_model(
        model, str(model_path),
        signature=infer_signature(train[FEATURES], model.predict(train[FEATURES])),
        input_example=train[FEATURES].head(3),
        pip_requirements=["scikit-learn==1.6.1", "pandas==2.2.3", "numpy==2.2.5"],
    )
    mlflow.log_artifacts(str(model_path), "model")
    write_json("reports/metrics.json", metrics)
    record = {**identity, "run_id": run.info.run_id,
              "model_uri": f"runs:/{run.info.run_id}/model",
              "metrics": metrics}
    write_json("reports/run.json", record)
    mlflow.log_artifact("reports/run.json", "lineage")
    print(json.dumps({"run_id": run.info.run_id, **metrics}, indent=2))

4.6 First Execution Order

Install dependencies and start local Tracking Server. Commands run in two terminals; all paths relative to project root. SQLite and local artifact for solo demo; team deployment later.

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

# Terminal A: keep running
mlflow server --backend-store-uri sqlite:///mlflow.db \
  --serve-artifacts --artifacts-destination ./mlartifacts \
  --host 127.0.0.1 --port 5000
# Terminal B: activate same venv
git init
dvc init
dvc config cache.type copy
mkdir -p ../ctr-dvc-remote
dvc remote add -d local ../ctr-dvc-remote
export MLFLOW_TRACKING_URI=http://127.0.0.1:5000

python src/make_demo.py
dvc add data/raw/events.csv
git add .
git commit -m "inputs: lock demo snapshot and training source"

dvc repro
dvc status
dvc push

git add dvc.lock reports/metrics.json .gitignore data/.gitignore
git commit -m "release: record trained outputs and validation metrics"
git tag -a release-demo-001 -m "First reproducible demo candidate"

If DVC-generated ignore paths differ, use git status to verify and commit all DVC-generated .gitignore files; do not add real data files to Git. Successful run yields two CSVs, data manifest, model directory, two report files, and a FINISHED MLflow Run.

Execution order matters: first dvc push to confirm large files available, then publish Git commit and tag. If tag pushed first, collaborators may see pointers but cannot pull files. Two systems lack distributed transactions; release pipeline must make pull verification the next gate.

Local remote relative path convenient for demo, but collaborator's new directory may resolve elsewhere. Shared environments use absolute paths or S3 remote. Demo does not need version_aware; Git history DVC pointers already locate content-addressable cache.

5. Experiment Comparison and History Recovery

5.1 Avoiding Parameter Experiment Overwrites

dvc repro

reuses unchanged stage caches, so second run may not train and produces no new MLflow Run. If business requires new Run per CI, explicitly force re-run training stage — don't mistake missing logs for server failure.

# Current training output managed by DVC. Force re-run only when new training truly needed.
dvc repro --force train

# Formal candidates use independent worktree; modify params, commit, then train.
git worktree add ../ctr-lr-010 -b experiment/lr-010
# In new worktree configure same remote, change params.yaml learning_rate to 0.10.
# Then git add params.yaml && git commit, then dvc pull and dvc repro.
# Compare metrics of two Git revisions.
dvc metrics diff <baseline-release-commit> <candidate-release-commit>

Training script forbids uncommitted parameters; direct training with uncommitted params is intercepted. This is a design choice for formal candidate training. For DVC ad-hoc experiments, provide separate exploration mode: save input diff or experiment revision, parameter digest, tag purpose=exploration, forbid direct promotion. Before formal promotion, write selected params to Git, generate new input commit, then run this pipeline. Don't disable production input checks for exploration convenience.

Parallel training by multiple people uses independent checkouts, worktrees, or containers. Shared content-addressable remote is common collaboration pattern; main risks are concurrent writes to same working directory, erroneous shared cache modifications, and GC competing with training. Don't mistake shared remote storage for inevitable model overwrite.

5.2 Restore Old Data and Reproduce Old Training

# Operate in clean new directory. Never switch files directly in serving model directory.
git clone <your-git-repository> ctr-restore
cd ctr-restore
git checkout --detach release-demo-001
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
dvc pull
dvc status
dvc pull

fetches and restores cache content. dvc checkout primarily restores from local cache; should not be described as network download command. --force allows overwriting workspace files, provides no atomic update guarantee. Restored directory only becomes service's new version after data verification, model loading, and smoke test all pass.

To retrain, set up another working directory, connect Tracking Server, run dvc repro --force, compare against pre-defined metric tolerances. Fixed random seed does not guarantee byte-for-byte consistency across CPU, BLAS, framework, or GPU. Restoring old artifacts and recomputing equivalent models are two different acceptance goals.

6. Model Registration: Offline Pass Is Only Candidate Eligibility

Training script does not auto-promote every experiment to production model. Release script requires clean workspace, explicit release tag, corresponding data and parameter digests, FINISHED Run, and metrics passing gates. Example thresholds are demo config, not universal click-prediction thresholds.

Production gates also need baseline comparison on fixed evaluation snapshot. Only when candidate and champion evaluate on same validation snapshot and label rule can AUC difference be directly compared. Cannot judge win/loss using this week's validation AUC vs last month's different sample AUC.

src/publish.py

import argparse
import json
import math
import os
from pathlib import Path
import mlflow
import yaml
from mlflow import MlflowClient
from common import git, sha256, write_json

p = argparse.ArgumentParser()
p.add_argument("--name", default="ctr-demo-v2")
p.add_argument("--release-tag", required=True)
args = p.parse_args()
if git("status", "--porcelain"):
    raise RuntimeError("publish requires a clean release checkout")
if git("rev-parse", f"{args.release_tag}^{commit}") != git("rev-parse", "HEAD"):
    raise RuntimeError("release tag must point at current HEAD")
record = json.loads(Path("reports/run.json").read_text())
metrics = json.loads(Path("reports/metrics.json").read_text())
cfg = yaml.safe_load(Path("params.yaml").read_text())["gate"]
if not all(math.isfinite(v) for v in metrics.values()):
    raise RuntimeError("invalid metrics")
if metrics["val_auc"] < cfg["min_auc"] or metrics["val_log_loss"] > cfg["max_log_loss"]:
    raise RuntimeError("quality gate failed")
if metrics != record["metrics"] or sha256("params.yaml") != record["params_sha256"]:
    raise RuntimeError("release inputs changed")
if sha256("data/prepared/manifest.json") != record["manifest_sha256"]:
    raise RuntimeError("dataset manifest changed")
for part in ["train", "valid"]:
    if sha256(f"data/prepared/{part}.csv") != record["dataset"][f"{part}_sha256"]:
        raise RuntimeError("release dataset changed")
mlflow.set_tracking_uri(os.environ["MLFLOW_TRACKING_URI"])
client = MlflowClient()
run = client.get_run(record["run_id"])
if run.info.status != "FINISHED":
    raise RuntimeError("run not finished")
for key, value in metrics.items():
    if not math.isclose(run.data.metrics[key], value, abs_tol=1e-12):
        raise RuntimeError("tracking metrics mismatch")
if run.data.tags.get("source_commit") != record["source_commit"]:
    raise RuntimeError("source commit mismatch")
# Register candidate version, do not modify champion. Before retry, query existing version for run_id.
version = mlflow.register_model(record["model_uri"], args.name)
client.set_model_version_tag(args.name, version.version, "release_tag", args.release_tag)
client.set_model_version_tag(args.name, version.version, "release_commit", git("rev-parse", "HEAD"))
client.set_model_version_tag(args.name, version.version, "dvc_lock_sha256", sha256("dvc.lock"))
client.set_model_version_tag(args.name, version.version, "validation_status", "offline_passed")
write_json("runtime/candidate.json", {
    "model_name": args.name, "version": version.version,
    "run_id": record["run_id"], "release_tag": args.release_tag,
    "release_commit": git("rev-parse", "HEAD"),
    "model_uri": f"models:/{args.name}/{version.version}",
})
print(Path("runtime/candidate.json").read_text())

src/smoke.py

import argparse
import os
import mlflow
import mlflow.sklearn
import numpy as np
import pandas as pd
from common import FEATURES

p = argparse.ArgumentParser()
p.add_argument("--model-uri", required=True)
args = p.parse_args()
mlflow.set_tracking_uri(os.environ["MLFLOW_TRACKING_URI"])
model = mlflow.sklearn.load_model(args.model_uri)
x = pd.DataFrame([[0.10, 100.0, 12], [0.40, 20.0, 20]], columns=FEATURES)
y = model.predict_proba(x)[:, 1]
if y.shape != (2,) or not np.isfinite(y).all() or not ((y >= 0) & (y <= 1)).all():
    raise RuntimeError("smoke failed")
print({"probabilities": y.tolist(), "status": "passed"})
python src/publish.py --release-tag release-demo-001
# Fill candidate version from candidate.json output. Don't assume always 1.
python src/smoke.py --model-uri models:/ctr-demo-v2/<candidate-version>
publish.py

is minimal candidate registration, not a full release controller. Registration success but tag write failure may leave unlabeled candidate versions. Production needs to persist release task ID, query existing version by model name and run_id before registration, and use single publisher or external lock to avoid concurrent duplicate creation. Retry must complete tags for same candidate; cannot treat new version creation as retry success.

Model storage duplication is intentional trade-off: DVC for full experiment output recovery, MLflow for model retrieval and service distribution. For large model weights, choose one authoritative artifact store; the other stores immutable URI, version, and checksums to reduce copies, but ensure unified retention policy.

7. Release and Rollback: Separate Model Version from Traffic Change

MLflow Registry alias is a mutable pointer. Service processes typically load model at startup; modifying champion does not auto-refresh running instances. Production release record must include: model name, fixed version, Run ID, source commit, release commit, image digest, artifact digest, evaluation snapshot, release time, and change task ID.

Recommended sequence: register candidate → download and verify → smoke test → new instance ready → small traffic canary → metric observation → expand traffic → update Registry alias and release record. Alias update and traffic cutover are not an atomic transaction; should be coordinated by retryable state machine.

Following commands only show Registry metadata updates; execute after release controller completes instance and traffic validation, not replacing deployment actions:

from mlflow import MlflowClient
client = MlflowClient()
name = "ctr-demo-v2"
new_version = "<validated-version>"
# Persist old version in release audit. First release has no previous champion.
old = client.get_model_version_by_alias(name, "champion")
client.set_registered_model_alias(name, "previous", old.version)
client.set_registered_model_alias(name, "champion", new_version)

Rollback reads fixed version from previous verified release manifest, restores model in new directory, runs smoke.py, deploys old image and old model combo, waits for instance ready, then cuts traffic. Registry champion finally syncs to actual live version. Forbid repeated alias reads in request handling, causing same batch to use different models.

Canary observation should include prediction distribution, service error rate and latency; business CTR may be affected by traffic selection, attribution wait, and experiment bucketing. Offline AUC lift ≠ online CTR lift. When new model incompatible with old feature protocol, must rollback feature config or maintain compatible interfaces simultaneously.

8. Team Deployment: PostgreSQL and S3 Artifact Proxy

8.1 Two Data Channels

DVC client accesses dvc-cache directly; MLflow Server uses separate identity to access mlflow-artifacts. With artifact proxy enabled, training clients upload models via MLflow Server, needing no write permission on model bucket, but still need DVC data bucket permissions.

# Existing S3 bucket with project-specific prefix
dvc remote add -d team s3://your-dvc-bucket/ctr-project
# MinIO or other S3-compatible service sets custom endpoint.
dvc remote modify team endpointurl https://objects.internal.example.com
dvc remote modify team verify true

Prefer short-lived workload identities. MinIO local dev can inject AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY. If credentials must go into DVC config, use --local to keep out of Git, but disk still holds sensitive plaintext. Forbid putting secrets in parameter files, MLflow params, or reports.

8.2 Explicit Server Startup Parameters

Pre-create artifact bucket and PostgreSQL database. Build own fixed-version server image, add PostgreSQL driver and S3 deps, pin image digest at deploy.

FROM python:3.12-slim
RUN pip install --no-cache-dir mlflow==2.22.0 SQLAlchemy==2.0.40 psycopg2-binary==2.9.10 boto3==1.38.23
USER 10001:10001
EXPOSE 5000
ENTRYPOINT ["mlflow", "server"]
# MLFLOW_BACKEND_URI injected via Secret; special chars in password must be URL-encoded.
# Example only shows process config; credentials via workload identity or Secret injection.
mlflow server \
  --backend-store-uri "$MLFLOW_BACKEND_URI" \
  --serve-artifacts \
  --artifacts-destination s3://your-mlflow-bucket/ctr-project \
  --host 0.0.0.0 --port 5000

For MinIO, server side additionally sets MLFLOW_S3_ENDPOINT_URL=https://objects.internal.example.com. Client sets MLFLOW_TRACKING_URI to internal Tracking address. When creating experiment, don't specify artifact location requiring client direct S3 connect, else mixes with proxy mode.

Switching artifact proxy does not auto-migrate existing experiment artifact locations. Create new experiment, verify full upload/download/model load, then plan historical content migration. Database backup and bucket backup both need recovery drills.

In Kubernetes, run server as Deployment; database and object storage use managed or independently operated services. Configure Service, readiness probes, resource limits, termination grace period. TLS, auth, authorization, network isolation should be implemented by validated gateway or service mesh; enabling artifact proxy does not equal tenant isolation.

Article does not provide unverified "official Helm values". Different Charts have non-universal fields; ${PASSWORD} in plain YAML won't auto-interpolate. When using a Chart, pin source and version, verify values schema, and clarify Secret injection method.

9. CI: Separate Training, Archival, and Release

Self-hosted runners need access to internal MLflow and object storage. Pull raw data, then run preparation and training separately, avoiding dvc pull requiring outputs that don't exist for first training. CI output first retained as candidate, not auto-updating champion.

# .github/workflows/train.yml
name: candidate-training
on:
  workflow_dispatch:
  push:
    branches: [main]
    paths:
      - "src/**"
      - "params.yaml"
      - "requirements.txt"
      - "dvc.yaml"
      - "data/raw/*.dvc"
permissions:
  contents: read
concurrency:
  group: ctr-candidate-main
  cancel-in-progress: false
jobs:
  train:
    runs-on: [self-hosted, linux, mlops]
    timeout-minutes: 60
    env:
      MLFLOW_TRACKING_URI: ${{ vars.MLFLOW_TRACKING_URI }}
      AWS_ACCESS_KEY_ID: ${{ secrets.DVC_ACCESS_KEY_ID }}
      AWS_SECRET_ACCESS_KEY: ${{ secrets.DVC_SECRET_ACCESS_KEY }}
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: python -m pip install -r requirements.txt
      - name: Prepare and train
        shell: bash
        run: |
          set -euo pipefail
          dvc pull data/raw/events.csv.dvc
          dvc repro prepare
          dvc repro --force train
          dvc push
      - uses: actions/upload-artifact@v4
        with:
          name: candidate-${{ github.sha }}-${{ github.run_id }}
          path: |
            dvc.lock
            reports/metrics.json
            reports/run.json
          if-no-files-found: error

This CI does not auto-commit outputs, write commit status, call CML, or execute production deployment. Candidate archival for review. Formal release task brings lock file and reports from archive back to matching source commit, generates release commit, does dvc pull verification in isolated directory, pushes release tag, then runs registration and deployment. Don't ignore new lock file produced by training; otherwise Git history cannot recover actual candidate outputs.

Gate parameters and evaluation results in repo still need protection. Production release uses controlled approval, fixed evaluation program, and independent evaluation set; cannot allow candidate code to write AUC=1.0 and pass. Third-party Actions should pin commit SHA in controlled repo; runners use isolated or ephemeral environments. Example uses short-lived credentials as preferred path; fixed Secrets only as easy-to-understand wiring example.

10. Common Failure Diagnosis

Symptom: No new Run. First Check: Did DVC hit cache? Response: Force re-run train when new training truly needed.

Symptom: Pointer exists but download fails. First Check: Remote, permissions, object pushed? Response: Stop release, push missing, verify pull in new directory.

Symptom: Artifact upload fails. First Check: Experiment artifact location vs proxy mode. Response: Validate upload channel with new experiment; don't mix S3 direct.

Symptom: Run search slow. First Check: Pagination, filters, slow SQL, lock waits, connection pool. Response: Optimize per actual execution plan; don't blindly add universal indexes.

Symptom: Local checkout slow. First Check: Link type, file size, disk and filesystem. Response: Use copy for isolation first, then test reflink.

Symptom: Model predicts OK but CTR drops. First Check: Online feature protocol, distribution, label attribution. Response: Locate data and service changes; canary rollback to fixed version.

Symptom: Registration retry creates multiple versions. First Check: Is publish task idempotent? Response: Reuse candidate by run_id; serialize publish and persist state.

MLflow Run metric writes, search, model registration, and LLM traces are different paths. Cannot treat trace async queue params as universal fix for all training logs and SQL bottlenecks. First collect request latency, DB slow queries, connection counts, failure rates, artifact upload times, then decide optimization direction.

DVC cache reflink,copy can use copy-on-write when supported; copy better for initial verification and one-off CI. Hardlink or symlink share cache content; in-place workspace file modification risks cache corruption. Before choosing link type, verify whether tools write files in-place and follow DVC protect/unprotect workflow.

11. Retention Policy: Don't Let "Cost Cutting" Delete the Last Rollback Model

DVC content-addressable objects and S3 bucket versioning are different mechanisms. Ordinary content-addressable remote doesn't need version_aware enabled just to restore history. Cloud version-aware mode suits workflows needing to record and access cloud object versions; verify behavior and support scope per official docs separately.

Deleting entire DVC cache by current object age is dangerous. A feature file written three years ago and never changed may still be referenced by latest release. Bucket versioning provides extra recovery chance but cannot replace reference retention policy; delete markers, lifecycle policies, and permissions can also make objects inaccessible.

Production retention set must include: live release, previous rollback-capable release, regulatory-required versions, important experiments, and running tasks. Shared remote GC must consider all repos and experiments referencing it, not just reachable objects from one checkout.

Suggest: build object reference manifest, compute candidate deletion set, isolate pending-delete objects, complete recovery drill, then reclaim in maintenance window. DVC GC and MLflow artifact cleanup are destructive actions; cannot go in ordinary training pipeline. Privacy deletion may require revoking historical training sets; then explicitly mark some releases as non-retrainable, keep deletion audit, cannot promise permanent full reproducibility simultaneously.

12. Big Data Evolution: Fix Data Layout Before Stacking Platforms

Small teams start with local remote and SQLite to validate process. Multi-person collaboration moves to S3-compatible storage, PostgreSQL, centralized Tracking. As tasks grow, add queues, isolate runners, quotas, observability. Scale evolution depends on data volume, throughput, permission boundaries — not directly on team headcount.

TB-level training sets prioritize Parquet partitioned by date, business, or stable ID; incremental production only rewrites affected partitions. Use immutable snapshot manifest to describe partition list; don't rewrite one big CSV daily. Many small files increase manifest and object request overhead; merge files appropriately at export and measure actual workload.

Data warehouse transaction snapshots, lakeFS, or other data version systems can handle source-data-level snapshots and branching; DVC binds exports and pipeline artifacts; MLflow manages experiments. Whether to introduce extra systems depends on multi-person data branching, cross-table consistency, and query needs. Don't promise unverified "MLflow federation", "auto cross-site sync", or specific recovery seconds.

13. Production Acceptance and Fault Injection

Experiment: Raw CSV adds duplicate event_id. Expected Behavior: Prepare fails, train does not run.

Experiment: Feature has NaN, Inf, illegal label. Expected Behavior: Prepare rejects input.

Experiment: Time boundary produces single-class validation set. Expected Behavior: Explicit failure; cannot fake AUC.

Experiment: Parameters changed but not committed. Expected Behavior: Train rejects formal candidate training.

Experiment: Modify prepared files without recomputing manifest. Expected Behavior: Train SHA-256 verification fails.

Experiment: Pause MLflow Server. Expected Behavior: Train fails, does not advance release.

Experiment: Revoke object write permission. Expected Behavior: dvc push fails, no formal release.

Experiment: Clear new directory local cache then restore old tag. Expected Behavior: Can pull from remote and pass model smoke test.

Experiment: Two release tasks start simultaneously. Expected Behavior: External release lock allows only one to update live state.

Experiment: Canary period online metrics exceed bounds. Expected Behavior: Restore fixed version and traffic from previous release manifest.

Acceptance by evidence: retain command exit codes, data digests, Run status, lock files, artifact load logs, and rollback task records. Don't turn example thresholds, synthetic data performance, or single local timing into production gains.

14. Actual Validation Scope of This Code Version

This version installed and verified DVC 3.59.2, MLflow 2.22.0, and main dependencies under Python 3.12. Tests used isolated temp directories, local DVC remote, MLflow Tracking and Registry directly connected to SQLite; no real business data used.

Executed Check: Generate and prepare data. Result: 12,000 rows, 7,560 train, 4,440 valid.

Executed Check: Duplicate ID, illegal label, infinite values. Result: All rejected.

Executed Check: DVC prepare → train → push. Result: Passed.

Executed Check: MLflow Run, candidate registration, fixed version smoke test. Result: Passed.

Executed Check: Repeat dvc repro. Result: Reused outputs, Run ID unchanged.

Executed Check: Force retrain train. Result: Generated new Run ID.

Executed Check: Restore release lock file, delete local cache, dvc pull. Result: Restored old model and reports from local remote, passed Registry model smoke test.

Validation AUC on synthetic data ≈ 0.7712, log loss ≈ 0.5686; only shows example can train and pass demo gates. Verification actually discovered and fixed pathspec 1.x vs old DVC, SQLAlchemy 2.1 vs MLflow 2.22 compatibility issues.

HTTP artifact proxy, S3/MinIO, PostgreSQL, Docker image build, GitHub Actions, and Kubernetes deployment were not executed in this validation; related configs are deployment examples per official interfaces, needing further acceptance on target infrastructure. Pinned major dependency versions ≠ full environment lock; production should save transitive dependency lock files and image digests.

15. Delivery Boundaries and Conclusion

This minimal engineering covers data generation, quality validation, time splitting, input identity recording, training, model artifacts, candidate registration, and smoke test. Production still needs: source snapshot export, authentication, dependency image locking, release idempotency, independent evaluation, canary control, and disaster recovery.

Truly usable data version management should let a team answer three questions after an online degradation: Which inputs did this model version actually use? Can we still recover that verified artifact? Which release record can bring the service back to stable state? DVC and MLflow are key tools for this chain; whether they land depends on whether input freezing, artifact retention, and release ordering are actually enforced.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

CI/CDdata versioningMLOpspipelineMLflowDVCmodel registrydata snapshotsexperiment trackingreproducible ML
Ray's Galactic Tech
Written by

Ray's Galactic Tech

Practice together, never alone. We cover programming languages, development tools, learning methods, and pitfall notes. We simplify complex topics, guiding you from beginner to advanced. Weekly practical content—let's grow together!

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.