Every ML team I have worked with keeps a drawer full of dead experiment trackers. A spreadsheet someone maintained for three months. A folder with train_final_v2_REAL.ipynb in it. A Slack channel where people pasted screenshots of loss curves. Eventually a real tool gets adopted, and eighteen months later it is just another layer in the drawer.
I have spent the last year owning ML platform roadmaps where MLflow is the tracking layer, and before that I ran ML infrastructure on cloud I was personally responsible for keeping alive.
The question is not “which tracker is best”
The market wants you to compare feature checklists. That is the wrong axis. Every one of these tools logs a metric and draws a line. What actually separates them is three questions:
- Where is your data allowed to live?
- Do you need only tracking, or a whole lifecycle?
- Who is going to run this thing?
Answer those honestly and the choice mostly makes itself.
What each tool actually is
| Tool | Licence & governance | Self-host | Scope |
|---|---|---|---|
| MLflow | Apache-2.0 [1], Linux Foundation [2] | First-class it is the default deployment [3] | Tracking, model registry, evaluation, GenAI tracing, AI gateway |
| ClearML | Apache-2.0 core + paid tiers [8][9] | Yes (ClearML Server, free) [8] | Tracking + orchestration + data + pipelines + serving |
| Aim | Apache-2.0 (AimStack) [10] | Yes that is the whole model [10] | Tracking and a fast comparison UI, nothing else |
| DVC / DVCLive | Apache-2.0; maintained by lakeFS since Nov 2025 [12] | N/A your logs are files in your Git repo [13] | Data/model versioning + Git-committed metric logging |
| Weights & Biases | Proprietary [15] | Yes commercial licence (Self-Managed / Dedicated Cloud) [16] | Tracking, sweeps, artifacts, reports, registry, Weave (LLM) |
| Neptune | Proprietary [17] | Yes commercial licence; self-hosting is first-class [17] | Tracking + registry, tuned for large training runs |
| Comet | Proprietary platform; Opik (its LLM eval layer) is open source [19] | Enterprise for the platform [18]; Opik is self-hostable [19] | Tracking, registry, production monitoring, Opik for LLM |
What you get if you self-host
| Tool | Model registry | Artifact store you control | GenAI tracing / eval | Multi-user auth | What it costs |
|---|---|---|---|---|---|
| MLflow | Yes, native (champion/challenger aliases) [5] | S3 / MinIO / GCS / Azure / local [4] | Yes MLflow 3, OpenTelemetry-based, auto-traces LangChain, LlamaIndex, OpenAI, Anthropic, CrewAI [6] | Basic HTTP auth; full RBAC (roles, admin UI) from 3.13 [7] | Free + a VM + Postgres + object storage + your ops time |
| ClearML | Yes | Yes | Partial | Basic in OSS; SSO/RBAC behind the paid tier [9] | Free self-host; you pay for SSO and scale |
| Aim | No | Local / mounted volume | Minimal | None built-in TLS only, no user auth [11] | Free + ops time |
| DVCLive | Via Git tags + DVC (GitOps style) [14] | Your DVC remote | No | Whatever your Git host enforces | Free; the DVC Studio dashboard is paid |
| W&B / Neptune / Comet | Yes | On the self-managed tier commercial licence [16][17][18] | W&B: yes (Weave) · Neptune: limited · Comet: yes (Opik, open source [19]) | SSO / RBAC (paid / enterprise tier) [16][17][18] | Per-seat + usage; self-hosting is a sales conversation |
Two things jump out of that second table. The open-source tools give you the artifact store and the data residency for free; the proprietary ones make it the thing you have to buy the enterprise plan for. And “open source” is a gradient, not a checkbox ClearML and MLflow are both Apache-2.0, but ClearML puts single sign-on behind a paywall and MLflow does not.
The decision, as a flowchart
1. Do you have a hard data-residency, air-gap, or "no third party touches our runs"
requirement? (regulated finance, health, defense, public sector)
|- YES -> self-hosted only. Skip W&B / Neptune / Comet SaaS entirely.
| Go to 2.
|- NO -> you may use SaaS, but keep reading the lock-in is real.
Go to 2.
2. What do you actually need beyond "log a metric"?
|- Tracking only, small team, everything already in Git
| -> DVCLive (or Aim if you want a richer UI and no Git noise)
|- Tracking + a model registry + a path to serving
| -> MLflow
|- Tracking + orchestration + data + pipelines, one platform
-> ClearML
3. Are you building LLM / agent apps need traces and evaluation, not just scalars?
|- YES -> MLflow 3 (self-host) or Opik (open source, self-hostable).
| W&B Weave only if you are already committed to W&B.
|- NO -> any of the above stands.
4. Who operates it?
|- You have a platform or infra person
| -> self-host MLflow or ClearML. Cheap, and you own everything.
|- You have nobody and a company card
-> SaaS. Accept the per-seat bill and the lock-in
and test the export button before you commit.
DEFAULT: MLflow, self-hosted, Postgres + object storage.
Widest integration surface, a real registry, a credible GenAI story,
Linux Foundation governance, and your data never leaves infrastructure
you control. The price is that you run it.
Why “self-host” is the part that matters
Ivan Illich called a tool convivial when the people using it stay in control of it, they can inspect it, repair it, and walk away from it without losing what they built. Most of the AI stack fails that test badly. Experiment tracking does not have to.
A licence is necessary but not sufficient. What actually protects you is governance plus the ability to run it yourself. MLflow sits at the Linux Foundation [2], which means no acquisition can relicense it or quietly kill it. Compare that with DVC, which was acquired by lakeFS in November 2025 [12], still Apache-2.0, still maintained, but the point is you found out after the fact. With a vendor SaaS you do not even get that: you get an email with 90 days’ notice.
Self-hosting is the practical expression of the same idea. Your runs live in your Postgres. Your artifacts live in your object storage. If the project loses momentum tomorrow, you lose a UI, not a decade of institutional memory. You are not renting your own history back from anyone.
There is a modest ecological version of this argument too. A self-hosted tracking server is one small always-on container and a database, shared by the whole team. The SaaS model is a paid seat per person and a data pipeline shipping every metric you log to someone else’s region. The cheaper, more legible option is usually the one you host.
Where the brain analogy breaks
It is tempting to call a tracking server the team’s shared memory, and that is roughly right. It is more like a lab notebook everyone can actually read. But the analogy breaks in a way worth naming. Biological memory is reconstructive and forgets on purpose; it keeps what matters and lets the rest decay. A tracking server hoards perfectly and forgets nothing. Ten thousand runs that no one will ever open again is not memory, it is a landfill with a search bar. Which is why #3 in this series is about conventions and pruning, not storage.
What MLflow does not give you
It gives you a place to put facts. That is all. It does not give you reproducibility nor good metrics. Theses things comes with discipline, real data versioning and evaluation design. And it certainly does not give you a model that works.
Finally self-hosting has a tail: backups, version upgrades, a Postgres that grows without bound and an auth story that was frankly an afterthought until the RBAC work landed in 3.13 [7]. If you need true multi-tenancy today, budget for an OAuth2 proxy and Keycloak in front of it, or wait for 3.13’s RBAC to settle. None of this is a reason not to self-host. It is the honest cost of doing so.
Go deeper
- MLflow self-hosting the official guide (tracking server + backend store + artifact store) and the tracking docs.
- MLflow, going further Model Registry & aliases, RBAC (3.13+), GenAI tracing.
- The projects themselves, to verify every claim above: mlflow/mlflow, clearml/clearml-server, aimhubio/aim, iterative/dvc + dvclive, comet-ml/opik.
- Tutorials & neutral maps: DVC’s Get Started: Experiment Tracking, W&B’s self-managed setup, the LF AI & Data Landscape, and awesome-ml-experiment-management.
Next in this series: a production-shaped self-hosted MLflow you can stand up in an afternoon Postgres, object storage, artifact proxy, auth with the compose file and Terraform in a repo you can fork.
References
- MLflow LICENSE (Apache-2.0). github.com/mlflow/mlflow
- “The MLflow Project Joins the Linux Foundation.” Linux Foundation press release, 2020. linuxfoundation.org
- MLflow Documentation Self-Hosting Overview. mlflow.org/docs/latest/self-hosting
- MLflow Documentation Artifact Stores. mlflow.org/docs/latest/self-hosting/architecture/artifact-store
- MLflow Documentation Model Registry (aliases; stages deprecated). mlflow.org/docs/latest/ml/model-registry
- MLflow Documentation Automatic Tracing / integrations. mlflow.org/docs/latest/genai/tracing/integrations
- MLflow 3.13.0 Release Notes Role-Based Access Control. mlflow.org/releases/3.13.0
- ClearML Server (GitHub, Apache-2.0). github.com/clearml/clearml-server
- ClearML Pricing (SSO on Scale / Enterprise tiers). clear.ml/pricing
- Aim LICENSE (Apache-2.0). github.com/aimhubio/aim
- Aim Documentation Track experiments with a remote server (TLS only). aimstack.readthedocs.io
- “DVC Joins lakeFS: Your Questions Answered.” dvc.org, November 2025. dvc.org/blog
- DVCLive Documentation. dvc.org/doc/dvclive
- GTO Git Tag Ops (DVC’s GitOps model registry). dvc.org/doc/gto
- Weights & Biases Master Service Agreement. wandb.ai/site/terms
- Weights & Biases Documentation Self-Managed hosting. docs.wandb.ai
- neptune.ai Deployment options (self-hosted). neptune.ai/product/deployment-options
- Comet Enterprise (self-hosted / VPC / on-prem). comet.com/site/enterprise
- Opik by Comet (GitHub, Apache-2.0, self-hostable). github.com/comet-ml/opik

Leave a Reply