157 lines
9.0 KiB
Markdown
157 lines
9.0 KiB
Markdown
# Model flavors
|
|
|
|
The *flavor* is how this runtime is told to deserialize a model. It is **configuration, never
|
|
discovery**: the value comes from the model document in the `models` collection and the code never
|
|
reads the artifact's `MLmodel` manifest to guess it.
|
|
|
|
```json
|
|
{
|
|
"model_config": {
|
|
"transform_flavor": "sklearn",
|
|
"predict_flavor": "joblib",
|
|
"retention_minutes": 60,
|
|
"target": "SE"
|
|
}
|
|
}
|
|
```
|
|
|
|
`predict_flavor` drives `load_predict_model` (artifact `prediction_model`); `transform_flavor` drives
|
|
`load_transform_model` (artifact `data_model`). Both default to `sklearn` when the key is absent.
|
|
|
|
## Accepted values
|
|
|
|
| Flavor | Read | Write (retrain) | How it loads |
|
|
|---|---|---|---|
|
|
| `sklearn` | yes | yes | `mlflow.sklearn.load_model(<uri>)` — the loader fetches what it needs |
|
|
| `pyfunc` | yes | yes | `mlflow.pyfunc.load_model(<uri>)`; the wrapper path downloads the artifact instead and unwraps `._model_impl.python_model` |
|
|
| `joblib` | yes | yes | downloads the same artifact and `joblib.load`s its `<artifact>/model.pkl` — **no MLflow flavor loader involved** |
|
|
|
|
Anything else raises `ValueError` with `INVALID_FLAVOR_MESSAGE`, **before** any artifact is resolved
|
|
or downloaded. The same message and the same three values govern reading and writing.
|
|
|
|
`pytorch` was accepted until this change and is now rejected like any other unknown value: `torch` is
|
|
not installed in the runtime, so `mlflow.pytorch.load_model` could never have loaded anything there.
|
|
|
|
**There is no fallback between flavors.** A failing loader propagates: one attempt, one cause. A
|
|
model configured with the wrong flavor fails visibly instead of being rescued silently.
|
|
|
|
## Same artifact, three readers
|
|
|
|
The flavor changes the reader, never the path. All three resolve the same two artifact directories of
|
|
the Production run:
|
|
|
|
```
|
|
<run>/artifacts/
|
|
├── data_model/ ← transform_flavor loads this one
|
|
│ ├── MLmodel ← parsed by sklearn/pyfunc; ignored by joblib
|
|
│ ├── model.pkl ← what joblib.load reads
|
|
│ └── code/ ← put on sys.path by every flavor, when present
|
|
├── prediction_model/ ← predict_flavor loads this one
|
|
└── model_card/, *.csv
|
|
```
|
|
|
|
The run's `artifacts/` root holds **no** `.pkl` — every model pickle lives one level down, inside its
|
|
artifact directory. Predict resolves that directory through the registry URI
|
|
(`models:/<name>/production`, which MLflow points at the same `prediction_model`); transform builds
|
|
the run artifact URI directly with `get_model_uri`, since the run is already resolved. Observability
|
|
is the same for the three: `MODEL_READ_*` on load, `MODEL_WRITE_*` on `log_model`, under the base
|
|
`operation_type` — no per-flavor metric, operation type or label.
|
|
|
|
## `code/` on `sys.path`, for every flavor
|
|
|
|
An artifact may ship the modules its own classes live in, under `<artifact>/code/`. Without that
|
|
directory on `sys.path`, unpickling raises `ModuleNotFoundError` — the artifact is intact, the class
|
|
is simply not importable.
|
|
|
|
MLflow's loaders do add it, but **from the manifest**: `_add_code_from_conf_to_system_path` acts only
|
|
when the `MLmodel` flavor config declares a `code` entry. An artifact carrying `code/` on disk
|
|
without that entry — code logged with `log_artifacts`, a re-log that dropped the config, anything a
|
|
`joblib` retrain wrote — loads with no module path at all. So this runtime reads the **directory**,
|
|
which no manifest can misdescribe:
|
|
|
|
- `prepend_artifact_code_dir` inserts `<artifact>/code` at the front of `sys.path`, **only when that
|
|
directory exists**, and only once (an entry already present is not duplicated).
|
|
- `download_artifact_code_dir` is what feeds it on the download paths, and it is deliberately cheap:
|
|
the remote listing (`list_artifacts(run_id, '<artifact>')`) decides first, so a run without `code`
|
|
downloads **nothing**, and a run with it downloads only `<artifact>/code` — a handful of `.py`
|
|
files, never the model's weights.
|
|
- `download_model` runs that step first, for every flavor and both of its branches, *before* any
|
|
model is loaded — the insert has to be in place before the loader unpickles.
|
|
- The models themselves are not pre-downloaded: `mlflow.sklearn.load_model` and
|
|
`mlflow.pyfunc.load_model` are handed the model URI and fetch what they need. `joblib` downloads
|
|
its artifact because it must read the pickle, and adds the entry again inside `_load_model_pickle`
|
|
through the same helper, so the flavor works when called directly too.
|
|
- When the manifest *does* declare `code`, MLflow inserts the same directory a second time. Harmless.
|
|
|
|
The entry is never removed: a loaded model may import lazily, long after the load returned.
|
|
|
|
The directory `download_artifact_code_dir` downloaded into is also what `download_model` returns
|
|
next to the model — `tuple[Any, str | None]`, the second element being the artifact directory the
|
|
retrain path re-logs `code/` from, or `None` when the artifact has no `code/`. It falls back to
|
|
whatever the loader downloaded on its own, which is a real artifact directory for `joblib` and for
|
|
the wrapper branch. `load_predict_model` and `load_transform_model` stay the **single** place where
|
|
the flavor is decided: `download_model` has no `joblib` branch, it delegates by `model_type`. Its
|
|
only remaining branch is `load_wrapper`, which is derived from `flavor == 'pyfunc'` but is otherwise
|
|
orthogonal to the flavor.
|
|
|
|
> `sys.path` is process-global: two models embedding a package with the same name (`utils/`) on the
|
|
> same worker resolve to whichever entered first. Mitigation, if it bites: one worker per model.
|
|
|
|
## `sklearn`
|
|
|
|
Default. A standard MLflow sklearn artifact: `mlflow.sklearn.log_model` writes `MLmodel` +
|
|
`model.pkl`, `mlflow.sklearn.load_model` reads it, and every outside consumer (MLflow UI, MLflow
|
|
serving) recognizes it as a model.
|
|
|
|
## `pyfunc`
|
|
|
|
For models logged as `mlflow.pyfunc`, including the wrapper models this pipeline produces. Retrain
|
|
writes through the wrapper's own `store_model(artifact_path=<artifact>, code_path=…)`. Two read
|
|
paths, and the second is **orthogonal to the flavor**:
|
|
|
|
- Plain: `mlflow.pyfunc.load_model(<uri>)`, called as `predict(context, model_input)`.
|
|
- Wrapper (`download_model(load_wrapper=True)`): downloads to `./tmp/artifacts/<model>/<artifact>`
|
|
via `dowload_artifacts`, loads as pyfunc, then unwraps `._model_impl.python_model`. Inside such an
|
|
artifact the real pickle is **nested** — `<artifact>/artifacts/stacking_model.pkl` for predict,
|
|
`<artifact>/artifacts/training_transformer.pkl` for transform (`PREDICTION_COMPRESSED_PATH` /
|
|
`TRANSFORMED_COMPRESSED_PATH`). That nesting is why the `joblib` listing does not recurse.
|
|
|
|
## `joblib`
|
|
|
|
For models whose `.pkl` is in the compressed form `joblib.dump` writes, which MLflow's `pickle.load`
|
|
does not read — the artifact is intact, but no flavor loader opens it.
|
|
|
|
- **Read**: `client.download_artifacts(run_id, <artifact>)`, then
|
|
`joblib.load(<artifact>/model.pkl)`. `model.pkl` is `MODEL_PICKLE_NAME`, the same constant
|
|
`log_model` writes; another `.pkl` directly inside the artifact is a fallback, in name order.
|
|
- **Write**: `joblib.dump(model, '<artifact>/model.pkl', compress=3)`, plus a re-log of the source
|
|
artifact's `code/` to `<artifact>/code`.
|
|
- **`code/` goes on `sys.path`** before the load, when present — see the section above; the joblib
|
|
flavor is where that started, and it is now shared by all three.
|
|
- **No `MLmodel` is written**, so only this runtime reads the artifact back —
|
|
`mlflow.sklearn.load_model` and `models:/<name>/production` both fail on it. A model retrained in
|
|
joblib stays joblib.
|
|
- **The pickled object must expose `predict(data)`**, plus `fit(data)` if retrained. A pyfunc
|
|
`PythonModel` wrapper expecting `predict(context, model_input)` raises `TypeError`.
|
|
|
|
## Switching a model to `joblib`
|
|
|
|
The runtime does not migrate documents; editing `models` is an operator step.
|
|
|
|
1. **Inventory before the deploy** — run predict/transform and note which models raise. There is no
|
|
discovery afterwards: the failure names the loader, not the fix.
|
|
2. **Edit the document** — `predict_flavor` and/or `transform_flavor` to `'joblib'`. Imported models
|
|
arrive ready: `import_model` writes it itself, see [`model-import.md`](model-import.md) § 12.
|
|
3. **Check the first retrain** — the new run must carry `prediction_model/model.pkl` and, when the
|
|
source had one, `prediction_model/code/`.
|
|
4. **Rollback is not just a revert** — artifacts a joblib retrain already wrote stop being loadable,
|
|
so it also means promoting the previous version back and reverting the documents' flavors.
|
|
|
|
## Known gap: how the run id is resolved
|
|
|
|
`get_model_run_id` takes `source.split('/')[2]` of the registered version's `source`. **Not confirmed
|
|
against the real tracking server**: no scheme reproduced locally puts the run id there (a file store
|
|
yields `''`, S3 the bucket, `mlflow-artifacts:/` the experiment id), yet predict works in production.
|
|
Every flavor that resolves a run goes through this line — confirm it with a **read-only** query
|
|
before relying on it.
|