# Model flavors The *flavor* is how this runtime is told to deserialize a model. It is **configuration, never discovery**: the value comes from the model document in the `models` collection and the code never reads the artifact's `MLmodel` manifest to guess it. ```json { "model_config": { "transform_flavor": "sklearn", "predict_flavor": "joblib", "retention_minutes": 60, "target": "SE" } } ``` `predict_flavor` drives `load_predict_model` (artifact `prediction_model`); `transform_flavor` drives `load_transform_model` (artifact `data_model`). Both default to `sklearn` when the key is absent. ## Accepted values | Flavor | Read | Write (retrain) | How it loads | |---|---|---|---| | `sklearn` | yes | yes | `mlflow.sklearn.load_model()` — the loader fetches what it needs | | `pyfunc` | yes | yes | `mlflow.pyfunc.load_model()`; the wrapper path downloads the artifact instead and unwraps `._model_impl.python_model` | | `joblib` | yes | yes | downloads the same artifact and `joblib.load`s its `/model.pkl` — **no MLflow flavor loader involved** | Anything else raises `ValueError` with `INVALID_FLAVOR_MESSAGE`, **before** any artifact is resolved or downloaded. The same message and the same three values govern reading and writing. `pytorch` was accepted until this change and is now rejected like any other unknown value: `torch` is not installed in the runtime, so `mlflow.pytorch.load_model` could never have loaded anything there. **There is no fallback between flavors.** A failing loader propagates: one attempt, one cause. A model configured with the wrong flavor fails visibly instead of being rescued silently. ## Same artifact, three readers The flavor changes the reader, never the path. All three resolve the same two artifact directories of the Production run: ``` /artifacts/ ├── data_model/ ← transform_flavor loads this one │ ├── MLmodel ← parsed by sklearn/pyfunc; ignored by joblib │ ├── model.pkl ← what joblib.load reads │ └── code/ ← put on sys.path by every flavor, when present ├── prediction_model/ ← predict_flavor loads this one └── model_card/, *.csv ``` The run's `artifacts/` root holds **no** `.pkl` — every model pickle lives one level down, inside its artifact directory. Predict resolves that directory through the registry URI (`models://production`, which MLflow points at the same `prediction_model`); transform builds the run artifact URI directly with `get_model_uri`, since the run is already resolved. Observability is the same for the three: `MODEL_READ_*` on load, `MODEL_WRITE_*` on `log_model`, under the base `operation_type` — no per-flavor metric, operation type or label. ## `code/` on `sys.path`, for every flavor An artifact may ship the modules its own classes live in, under `/code/`. Without that directory on `sys.path`, unpickling raises `ModuleNotFoundError` — the artifact is intact, the class is simply not importable. MLflow's loaders do add it, but **from the manifest**: `_add_code_from_conf_to_system_path` acts only when the `MLmodel` flavor config declares a `code` entry. An artifact carrying `code/` on disk without that entry — code logged with `log_artifacts`, a re-log that dropped the config, anything a `joblib` retrain wrote — loads with no module path at all. So this runtime reads the **directory**, which no manifest can misdescribe: - `prepend_artifact_code_dir` inserts `/code` at the front of `sys.path`, **only when that directory exists**, and only once (an entry already present is not duplicated). - `download_artifact_code_dir` is what feeds it on the download paths, and it is deliberately cheap: the remote listing (`list_artifacts(run_id, '')`) decides first, so a run without `code` downloads **nothing**, and a run with it downloads only `/code` — a handful of `.py` files, never the model's weights. - `download_model` runs that step first, for every flavor and both of its branches, *before* any model is loaded — the insert has to be in place before the loader unpickles. - The models themselves are not pre-downloaded: `mlflow.sklearn.load_model` and `mlflow.pyfunc.load_model` are handed the model URI and fetch what they need. `joblib` downloads its artifact because it must read the pickle, and adds the entry again inside `_load_model_pickle` through the same helper, so the flavor works when called directly too. - When the manifest *does* declare `code`, MLflow inserts the same directory a second time. Harmless. The entry is never removed: a loaded model may import lazily, long after the load returned. The directory `download_artifact_code_dir` downloaded into is also what `download_model` returns next to the model — `tuple[Any, str | None]`, the second element being the artifact directory the retrain path re-logs `code/` from, or `None` when the artifact has no `code/`. It falls back to whatever the loader downloaded on its own, which is a real artifact directory for `joblib` and for the wrapper branch. `load_predict_model` and `load_transform_model` stay the **single** place where the flavor is decided: `download_model` has no `joblib` branch, it delegates by `model_type`. Its only remaining branch is `load_wrapper`, which is derived from `flavor == 'pyfunc'` but is otherwise orthogonal to the flavor. > `sys.path` is process-global: two models embedding a package with the same name (`utils/`) on the > same worker resolve to whichever entered first. Mitigation, if it bites: one worker per model. ## `sklearn` Default. A standard MLflow sklearn artifact: `mlflow.sklearn.log_model` writes `MLmodel` + `model.pkl`, `mlflow.sklearn.load_model` reads it, and every outside consumer (MLflow UI, MLflow serving) recognizes it as a model. ## `pyfunc` For models logged as `mlflow.pyfunc`, including the wrapper models this pipeline produces. Retrain writes through the wrapper's own `store_model(artifact_path=, code_path=…)`. Two read paths, and the second is **orthogonal to the flavor**: - Plain: `mlflow.pyfunc.load_model()`, called as `predict(context, model_input)`. - Wrapper (`download_model(load_wrapper=True)`): downloads to `./tmp/artifacts//` via `dowload_artifacts`, loads as pyfunc, then unwraps `._model_impl.python_model`. Inside such an artifact the real pickle is **nested** — `/artifacts/stacking_model.pkl` for predict, `/artifacts/training_transformer.pkl` for transform (`PREDICTION_COMPRESSED_PATH` / `TRANSFORMED_COMPRESSED_PATH`). That nesting is why the `joblib` listing does not recurse. ## `joblib` For models whose `.pkl` is in the compressed form `joblib.dump` writes, which MLflow's `pickle.load` does not read — the artifact is intact, but no flavor loader opens it. - **Read**: `client.download_artifacts(run_id, )`, then `joblib.load(/model.pkl)`. `model.pkl` is `MODEL_PICKLE_NAME`, the same constant `log_model` writes; another `.pkl` directly inside the artifact is a fallback, in name order. - **Write**: `joblib.dump(model, '/model.pkl', compress=3)`, plus a re-log of the source artifact's `code/` to `/code`. - **`code/` goes on `sys.path`** before the load, when present — see the section above; the joblib flavor is where that started, and it is now shared by all three. - **No `MLmodel` is written**, so only this runtime reads the artifact back — `mlflow.sklearn.load_model` and `models://production` both fail on it. A model retrained in joblib stays joblib. - **The pickled object must expose `predict(data)`**, plus `fit(data)` if retrained. A pyfunc `PythonModel` wrapper expecting `predict(context, model_input)` raises `TypeError`. ## Switching a model to `joblib` The runtime does not migrate documents; editing `models` is an operator step. 1. **Inventory before the deploy** — run predict/transform and note which models raise. There is no discovery afterwards: the failure names the loader, not the fix. 2. **Edit the document** — `predict_flavor` and/or `transform_flavor` to `'joblib'`. Imported models arrive ready: `import_model` writes it itself, see [`model-import.md`](model-import.md) § 12. 3. **Check the first retrain** — the new run must carry `prediction_model/model.pkl` and, when the source had one, `prediction_model/code/`. 4. **Rollback is not just a revert** — artifacts a joblib retrain already wrote stop being loadable, so it also means promoting the previous version back and reverting the documents' flavors. ## Known gap: how the run id is resolved `get_model_run_id` takes `source.split('/')[2]` of the registered version's `source`. **Not confirmed against the real tracking server**: no scheme reproduced locally puts the run id there (a file store yields `''`, S3 the bucket, `mlflow-artifacts:/` the experiment id), yet predict works in production. Every flavor that resolves a run goes through this line — confirm it with a **read-only** query before relying on it.