A practitioner’s comparison of MLflow against Weights & Biases, Neptune, ClearML, Aim, Comet and DVCLive — and the case for running your experiment tracker on infrastructure you control.

By

—

Why MLflow?

Every ML team I have worked with keeps a drawer full of dead experiment trackers. A spreadsheet someone maintained for three months. A folder with train_final_v2_REAL.ipynb in it. A Slack channel where people pasted screenshots of loss curves. Eventually a real tool gets adopted, and eighteen months later it is just another layer in the drawer.

I have spent the last year owning ML platform roadmaps where MLflow is the tracking layer, and before that I ran ML infrastructure on cloud I was personally responsible for keeping alive.

The question is not “which tracker is best”

The market wants you to compare feature checklists. That is the wrong axis. Every one of these tools logs a metric and draws a line. What actually separates them is three questions:

  1. Where is your data allowed to live?
  2. Do you need only tracking, or a whole lifecycle?
  3. Who is going to run this thing?

Answer those honestly and the choice mostly makes itself.

What each tool actually is

Tool Licence & governance Self-host Scope
MLflow Apache-2.0 [1], Linux Foundation [2] First-class it is the default deployment [3] Tracking, model registry, evaluation, GenAI tracing, AI gateway
ClearML Apache-2.0 core + paid tiers [8][9] Yes (ClearML Server, free) [8] Tracking + orchestration + data + pipelines + serving
Aim Apache-2.0 (AimStack) [10] Yes that is the whole model [10] Tracking and a fast comparison UI, nothing else
DVC / DVCLive Apache-2.0; maintained by lakeFS since Nov 2025 [12] N/A your logs are files in your Git repo [13] Data/model versioning + Git-committed metric logging
Weights & Biases Proprietary [15] Yes commercial licence (Self-Managed / Dedicated Cloud) [16] Tracking, sweeps, artifacts, reports, registry, Weave (LLM)
Neptune Proprietary [17] Yes commercial licence; self-hosting is first-class [17] Tracking + registry, tuned for large training runs
Comet Proprietary platform; Opik (its LLM eval layer) is open source [19] Enterprise for the platform [18]; Opik is self-hostable [19] Tracking, registry, production monitoring, Opik for LLM

What you get if you self-host

Tool Model registry Artifact store you control GenAI tracing / eval Multi-user auth What it costs
MLflow Yes, native (champion/challenger aliases) [5] S3 / MinIO / GCS / Azure / local [4] Yes MLflow 3, OpenTelemetry-based, auto-traces LangChain, LlamaIndex, OpenAI, Anthropic, CrewAI [6] Basic HTTP auth; full RBAC (roles, admin UI) from 3.13 [7] Free + a VM + Postgres + object storage + your ops time
ClearML Yes Yes Partial Basic in OSS; SSO/RBAC behind the paid tier [9] Free self-host; you pay for SSO and scale
Aim No Local / mounted volume Minimal None built-in TLS only, no user auth [11] Free + ops time
DVCLive Via Git tags + DVC (GitOps style) [14] Your DVC remote No Whatever your Git host enforces Free; the DVC Studio dashboard is paid
W&B / Neptune / Comet Yes On the self-managed tier commercial licence [16][17][18] W&B: yes (Weave) · Neptune: limited · Comet: yes (Opik, open source [19]) SSO / RBAC (paid / enterprise tier) [16][17][18] Per-seat + usage; self-hosting is a sales conversation

Two things jump out of that second table. The open-source tools give you the artifact store and the data residency for free; the proprietary ones make it the thing you have to buy the enterprise plan for. And “open source” is a gradient, not a checkbox ClearML and MLflow are both Apache-2.0, but ClearML puts single sign-on behind a paywall and MLflow does not.

The decision, as a flowchart

1. Do you have a hard data-residency, air-gap, or "no third party touches our runs"
   requirement?  (regulated finance, health, defense, public sector)

   |- YES -> self-hosted only. Skip W&B / Neptune / Comet SaaS entirely.
   |         Go to 2.
   |- NO  -> you may use SaaS, but keep reading  the lock-in is real.
             Go to 2.

2. What do you actually need beyond "log a metric"?

   |- Tracking only, small team, everything already in Git
   |     -> DVCLive  (or Aim if you want a richer UI and no Git noise)
   |- Tracking + a model registry + a path to serving
   |     -> MLflow
   |- Tracking + orchestration + data + pipelines, one platform
         -> ClearML

3. Are you building LLM / agent apps need traces and evaluation, not just scalars?

   |- YES -> MLflow 3 (self-host) or Opik (open source, self-hostable).
   |         W&B Weave only if you are already committed to W&B.
   |- NO  -> any of the above stands.

4. Who operates it?

   |- You have a platform or infra person
   |     -> self-host MLflow or ClearML. Cheap, and you own everything.
   |- You have nobody and a company card
         -> SaaS. Accept the per-seat bill and the lock-in 
            and test the export button before you commit.

DEFAULT: MLflow, self-hosted, Postgres + object storage.
Widest integration surface, a real registry, a credible GenAI story,
Linux Foundation governance, and your data never leaves infrastructure
you control. The price is that you run it.

Why “self-host” is the part that matters

Ivan Illich called a tool convivial when the people using it stay in control of it, they can inspect it, repair it, and walk away from it without losing what they built. Most of the AI stack fails that test badly. Experiment tracking does not have to.

A licence is necessary but not sufficient. What actually protects you is governance plus the ability to run it yourself. MLflow sits at the Linux Foundation [2], which means no acquisition can relicense it or quietly kill it. Compare that with DVC, which was acquired by lakeFS in November 2025 [12], still Apache-2.0, still maintained, but the point is you found out after the fact. With a vendor SaaS you do not even get that: you get an email with 90 days’ notice.

Self-hosting is the practical expression of the same idea. Your runs live in your Postgres. Your artifacts live in your object storage. If the project loses momentum tomorrow, you lose a UI, not a decade of institutional memory. You are not renting your own history back from anyone.

There is a modest ecological version of this argument too. A self-hosted tracking server is one small always-on container and a database, shared by the whole team. The SaaS model is a paid seat per person and a data pipeline shipping every metric you log to someone else’s region. The cheaper, more legible option is usually the one you host.

Where the brain analogy breaks

It is tempting to call a tracking server the team’s shared memory, and that is roughly right. It is more like a lab notebook everyone can actually read. But the analogy breaks in a way worth naming. Biological memory is reconstructive and forgets on purpose; it keeps what matters and lets the rest decay. A tracking server hoards perfectly and forgets nothing. Ten thousand runs that no one will ever open again is not memory, it is a landfill with a search bar. Which is why #3 in this series is about conventions and pruning, not storage.

What MLflow does not give you

It gives you a place to put facts. That is all. It does not give you reproducibility nor good metrics. Theses things comes with discipline, real data versioning and evaluation design. And it certainly does not give you a model that works.

Finally self-hosting has a tail: backups, version upgrades, a Postgres that grows without bound and an auth story that was frankly an afterthought until the RBAC work landed in 3.13 [7]. If you need true multi-tenancy today, budget for an OAuth2 proxy and Keycloak in front of it, or wait for 3.13’s RBAC to settle. None of this is a reason not to self-host. It is the honest cost of doing so.

Go deeper

Next in this series: a production-shaped self-hosted MLflow you can stand up in an afternoon Postgres, object storage, artifact proxy, auth with the compose file and Terraform in a repo you can fork.

References

  1. MLflow LICENSE (Apache-2.0). github.com/mlflow/mlflow
  2. “The MLflow Project Joins the Linux Foundation.” Linux Foundation press release, 2020. linuxfoundation.org
  3. MLflow Documentation Self-Hosting Overview. mlflow.org/docs/latest/self-hosting
  4. MLflow Documentation Artifact Stores. mlflow.org/docs/latest/self-hosting/architecture/artifact-store
  5. MLflow Documentation Model Registry (aliases; stages deprecated). mlflow.org/docs/latest/ml/model-registry
  6. MLflow Documentation Automatic Tracing / integrations. mlflow.org/docs/latest/genai/tracing/integrations
  7. MLflow 3.13.0 Release Notes Role-Based Access Control. mlflow.org/releases/3.13.0
  8. ClearML Server (GitHub, Apache-2.0). github.com/clearml/clearml-server
  9. ClearML Pricing (SSO on Scale / Enterprise tiers). clear.ml/pricing
  10. Aim LICENSE (Apache-2.0). github.com/aimhubio/aim
  11. Aim Documentation Track experiments with a remote server (TLS only). aimstack.readthedocs.io
  12. “DVC Joins lakeFS: Your Questions Answered.” dvc.org, November 2025. dvc.org/blog
  13. DVCLive Documentation. dvc.org/doc/dvclive
  14. GTO Git Tag Ops (DVC’s GitOps model registry). dvc.org/doc/gto
  15. Weights & Biases Master Service Agreement. wandb.ai/site/terms
  16. Weights & Biases Documentation Self-Managed hosting. docs.wandb.ai
  17. neptune.ai Deployment options (self-hosted). neptune.ai/product/deployment-options
  18. Comet Enterprise (self-hosted / VPC / on-prem). comet.com/site/enterprise
  19. Opik by Comet (GitHub, Apache-2.0, self-hostable). github.com/comet-ml/opik

Leave a Reply

Discover more from Neur_It

Subscribe now to keep reading and get access to the full archive.

Continue reading