This commit is contained in:
156
docs/model-flavors.md
Normal file
156
docs/model-flavors.md
Normal file
@@ -0,0 +1,156 @@
|
||||
# Model flavors
|
||||
|
||||
The *flavor* is how this runtime is told to deserialize a model. It is **configuration, never
|
||||
discovery**: the value comes from the model document in the `models` collection and the code never
|
||||
reads the artifact's `MLmodel` manifest to guess it.
|
||||
|
||||
```json
|
||||
{
|
||||
"model_config": {
|
||||
"transform_flavor": "sklearn",
|
||||
"predict_flavor": "joblib",
|
||||
"retention_minutes": 60,
|
||||
"target": "SE"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`predict_flavor` drives `load_predict_model` (artifact `prediction_model`); `transform_flavor` drives
|
||||
`load_transform_model` (artifact `data_model`). Both default to `sklearn` when the key is absent.
|
||||
|
||||
## Accepted values
|
||||
|
||||
| Flavor | Read | Write (retrain) | How it loads |
|
||||
|---|---|---|---|
|
||||
| `sklearn` | yes | yes | `mlflow.sklearn.load_model(<uri>)` — the loader fetches what it needs |
|
||||
| `pyfunc` | yes | yes | `mlflow.pyfunc.load_model(<uri>)`; the wrapper path downloads the artifact instead and unwraps `._model_impl.python_model` |
|
||||
| `joblib` | yes | yes | downloads the same artifact and `joblib.load`s its `<artifact>/model.pkl` — **no MLflow flavor loader involved** |
|
||||
|
||||
Anything else raises `ValueError` with `INVALID_FLAVOR_MESSAGE`, **before** any artifact is resolved
|
||||
or downloaded. The same message and the same three values govern reading and writing.
|
||||
|
||||
`pytorch` was accepted until this change and is now rejected like any other unknown value: `torch` is
|
||||
not installed in the runtime, so `mlflow.pytorch.load_model` could never have loaded anything there.
|
||||
|
||||
**There is no fallback between flavors.** A failing loader propagates: one attempt, one cause. A
|
||||
model configured with the wrong flavor fails visibly instead of being rescued silently.
|
||||
|
||||
## Same artifact, three readers
|
||||
|
||||
The flavor changes the reader, never the path. All three resolve the same two artifact directories of
|
||||
the Production run:
|
||||
|
||||
```
|
||||
<run>/artifacts/
|
||||
├── data_model/ ← transform_flavor loads this one
|
||||
│ ├── MLmodel ← parsed by sklearn/pyfunc; ignored by joblib
|
||||
│ ├── model.pkl ← what joblib.load reads
|
||||
│ └── code/ ← put on sys.path by every flavor, when present
|
||||
├── prediction_model/ ← predict_flavor loads this one
|
||||
└── model_card/, *.csv
|
||||
```
|
||||
|
||||
The run's `artifacts/` root holds **no** `.pkl` — every model pickle lives one level down, inside its
|
||||
artifact directory. Predict resolves that directory through the registry URI
|
||||
(`models:/<name>/production`, which MLflow points at the same `prediction_model`); transform builds
|
||||
the run artifact URI directly with `get_model_uri`, since the run is already resolved. Observability
|
||||
is the same for the three: `MODEL_READ_*` on load, `MODEL_WRITE_*` on `log_model`, under the base
|
||||
`operation_type` — no per-flavor metric, operation type or label.
|
||||
|
||||
## `code/` on `sys.path`, for every flavor
|
||||
|
||||
An artifact may ship the modules its own classes live in, under `<artifact>/code/`. Without that
|
||||
directory on `sys.path`, unpickling raises `ModuleNotFoundError` — the artifact is intact, the class
|
||||
is simply not importable.
|
||||
|
||||
MLflow's loaders do add it, but **from the manifest**: `_add_code_from_conf_to_system_path` acts only
|
||||
when the `MLmodel` flavor config declares a `code` entry. An artifact carrying `code/` on disk
|
||||
without that entry — code logged with `log_artifacts`, a re-log that dropped the config, anything a
|
||||
`joblib` retrain wrote — loads with no module path at all. So this runtime reads the **directory**,
|
||||
which no manifest can misdescribe:
|
||||
|
||||
- `prepend_artifact_code_dir` inserts `<artifact>/code` at the front of `sys.path`, **only when that
|
||||
directory exists**, and only once (an entry already present is not duplicated).
|
||||
- `download_artifact_code_dir` is what feeds it on the download paths, and it is deliberately cheap:
|
||||
the remote listing (`list_artifacts(run_id, '<artifact>')`) decides first, so a run without `code`
|
||||
downloads **nothing**, and a run with it downloads only `<artifact>/code` — a handful of `.py`
|
||||
files, never the model's weights.
|
||||
- `download_model` runs that step first, for every flavor and both of its branches, *before* any
|
||||
model is loaded — the insert has to be in place before the loader unpickles.
|
||||
- The models themselves are not pre-downloaded: `mlflow.sklearn.load_model` and
|
||||
`mlflow.pyfunc.load_model` are handed the model URI and fetch what they need. `joblib` downloads
|
||||
its artifact because it must read the pickle, and adds the entry again inside `_load_model_pickle`
|
||||
through the same helper, so the flavor works when called directly too.
|
||||
- When the manifest *does* declare `code`, MLflow inserts the same directory a second time. Harmless.
|
||||
|
||||
The entry is never removed: a loaded model may import lazily, long after the load returned.
|
||||
|
||||
The directory `download_artifact_code_dir` downloaded into is also what `download_model` returns
|
||||
next to the model — `tuple[Any, str | None]`, the second element being the artifact directory the
|
||||
retrain path re-logs `code/` from, or `None` when the artifact has no `code/`. It falls back to
|
||||
whatever the loader downloaded on its own, which is a real artifact directory for `joblib` and for
|
||||
the wrapper branch. `load_predict_model` and `load_transform_model` stay the **single** place where
|
||||
the flavor is decided: `download_model` has no `joblib` branch, it delegates by `model_type`. Its
|
||||
only remaining branch is `load_wrapper`, which is derived from `flavor == 'pyfunc'` but is otherwise
|
||||
orthogonal to the flavor.
|
||||
|
||||
> `sys.path` is process-global: two models embedding a package with the same name (`utils/`) on the
|
||||
> same worker resolve to whichever entered first. Mitigation, if it bites: one worker per model.
|
||||
|
||||
## `sklearn`
|
||||
|
||||
Default. A standard MLflow sklearn artifact: `mlflow.sklearn.log_model` writes `MLmodel` +
|
||||
`model.pkl`, `mlflow.sklearn.load_model` reads it, and every outside consumer (MLflow UI, MLflow
|
||||
serving) recognizes it as a model.
|
||||
|
||||
## `pyfunc`
|
||||
|
||||
For models logged as `mlflow.pyfunc`, including the wrapper models this pipeline produces. Retrain
|
||||
writes through the wrapper's own `store_model(artifact_path=<artifact>, code_path=…)`. Two read
|
||||
paths, and the second is **orthogonal to the flavor**:
|
||||
|
||||
- Plain: `mlflow.pyfunc.load_model(<uri>)`, called as `predict(context, model_input)`.
|
||||
- Wrapper (`download_model(load_wrapper=True)`): downloads to `./tmp/artifacts/<model>/<artifact>`
|
||||
via `dowload_artifacts`, loads as pyfunc, then unwraps `._model_impl.python_model`. Inside such an
|
||||
artifact the real pickle is **nested** — `<artifact>/artifacts/stacking_model.pkl` for predict,
|
||||
`<artifact>/artifacts/training_transformer.pkl` for transform (`PREDICTION_COMPRESSED_PATH` /
|
||||
`TRANSFORMED_COMPRESSED_PATH`). That nesting is why the `joblib` listing does not recurse.
|
||||
|
||||
## `joblib`
|
||||
|
||||
For models whose `.pkl` is in the compressed form `joblib.dump` writes, which MLflow's `pickle.load`
|
||||
does not read — the artifact is intact, but no flavor loader opens it.
|
||||
|
||||
- **Read**: `client.download_artifacts(run_id, <artifact>)`, then
|
||||
`joblib.load(<artifact>/model.pkl)`. `model.pkl` is `MODEL_PICKLE_NAME`, the same constant
|
||||
`log_model` writes; another `.pkl` directly inside the artifact is a fallback, in name order.
|
||||
- **Write**: `joblib.dump(model, '<artifact>/model.pkl', compress=3)`, plus a re-log of the source
|
||||
artifact's `code/` to `<artifact>/code`.
|
||||
- **`code/` goes on `sys.path`** before the load, when present — see the section above; the joblib
|
||||
flavor is where that started, and it is now shared by all three.
|
||||
- **No `MLmodel` is written**, so only this runtime reads the artifact back —
|
||||
`mlflow.sklearn.load_model` and `models:/<name>/production` both fail on it. A model retrained in
|
||||
joblib stays joblib.
|
||||
- **The pickled object must expose `predict(data)`**, plus `fit(data)` if retrained. A pyfunc
|
||||
`PythonModel` wrapper expecting `predict(context, model_input)` raises `TypeError`.
|
||||
|
||||
## Switching a model to `joblib`
|
||||
|
||||
The runtime does not migrate documents; editing `models` is an operator step.
|
||||
|
||||
1. **Inventory before the deploy** — run predict/transform and note which models raise. There is no
|
||||
discovery afterwards: the failure names the loader, not the fix.
|
||||
2. **Edit the document** — `predict_flavor` and/or `transform_flavor` to `'joblib'`. Imported models
|
||||
arrive ready: `import_model` writes it itself, see [`model-import.md`](model-import.md) § 12.
|
||||
3. **Check the first retrain** — the new run must carry `prediction_model/model.pkl` and, when the
|
||||
source had one, `prediction_model/code/`.
|
||||
4. **Rollback is not just a revert** — artifacts a joblib retrain already wrote stop being loadable,
|
||||
so it also means promoting the previous version back and reverting the documents' flavors.
|
||||
|
||||
## Known gap: how the run id is resolved
|
||||
|
||||
`get_model_run_id` takes `source.split('/')[2]` of the registered version's `source`. **Not confirmed
|
||||
against the real tracking server**: no scheme reproduced locally puts the run id there (a file store
|
||||
yields `''`, S3 the bucket, `mlflow-artifacts:/` the experiment id), yet predict works in production.
|
||||
Every flavor that resolves a run goes through this line — confirm it with a **read-only** query
|
||||
before relying on it.
|
||||
Reference in New Issue
Block a user