Refactor ModelMetrics to utilize DriftAnalysis for drift detection - Replaced ModelAnalysis with DriftAnalysis in the ModelMetrics class to enhance drift detection capabilities. - Updated method signatures and documentation to reflect the changes in target_name and return values. - Adjusted data handling to ensure compatibility with the new analysis methods and improved clarity in the drift metrics dataframe preparation.
8.1 KiB
E2E test run report
Date: 2026-05-08
Command: source venv/bin/activate && rtk pytest e2e/ -v --tb=short
Environment: Linux, Python 3.11.15, pytest 9.0.3
Summary
| Metric | Count |
|---|---|
| Collected | 43 |
| Passed | 37 |
| Failed | 6 |
Full pytest output (compressed by rtk) was written to:
~/.local/share/rtk/tee/1778268193_pytest.log
Failed tests (6)
e2e/test_drift.py::test_drift_happy_path_persists_all_columns_with_reference_datae2e/test_drift.py::test_drift_uses_30pct_fallback_when_reference_unavailablee2e/test_drift.py::test_drift_chunk_period_seconds_preserves_seconds_in_chunk_start_datee2e/test_predictions_batch_prediction_process.py::test_scenario_2_1_3_input_gate_triggers_repeate2e/test_predictions_batch_prediction_process.py::test_scenario_2_2_3_transform_gate_triggers_repeate2e/test_predictions_batch_prediction_process.py::test_scenario_2_3_3_predict_gate_triggers_repeat
Follow-up: Jensen–Shannon NULLs for the drift happy-path scenario are fully traced (histogram out-of-range + density=True, vs NannyML’s leftover bin) in DRIFT_JS_NULL_INVESTIGATION.md.
Failure group A — Drift: drift_metrics.value NOT NULL (2 tests)
Error
psycopg2.errors.NotNullViolation: null value in column value of relation sientia_data.drift_metrics violates not-null constraint.
Example failing row (from logs): feature=sensor_2, method=jensen_shannon, value=null, with kolmogorov_smirnov / wasserstein populated for the same chunk.
The bulk INSERT built by Activities.export_data_to_postgres includes parameters such as 'value__4': None for jensen_shannon on a given chunk.
Root cause
calculate_drift (real DriftAnalysis + ModelMetrics.get_drift_metrics) can emit NaN / missing values for some metric methods (here Jensen–Shannon) on some features/chunks. Pandas/SQLAlchemy turns that into SQL NULL, while the E2E schema (mirroring production) defines:
value numeric NOT NULL
in e2e/db_schema.sql for sientia_data.drift_metrics.
Recommended fixes (pick one consistent with product rules)
- Application layer (preferred if NULLs are never valid in production): Before export, sanitize the drift dataframe — e.g. drop rows where
valueis null/NaN, or replace with a defined sentinel (only if product agrees), or skip emitting that method row when the statistic is undefined. - Analytics layer: Harden the Jensen–Shannon (and similar) paths so they always return a finite float for the supported inputs, or explicitly map “undefined” to an agreed numeric convention.
- Schema (only if product allows missing metrics): Align DDL with reality by making
valuenullable — only if production and downstream consumers already expect missing metrics; the E2E comment intest_drift.pysuggestsfeature/timestampnullable cases exist, butvalueis still listed inNON_NULL_DRIFT_COLUMNS.
Tests affected
test_drift_happy_path_persists_all_columns_with_reference_datatest_drift_uses_30pct_fallback_when_reference_unavailable
Failure group B — Drift: chunk period seconds (1 test)
Error
AssertionError: expected at least one drift row to be persisted (e2e/test_drift.py:493).
The workflow run completed without failing the test via WorkflowFailureError, but no rows were found in sientia_data.drift_metrics for the model.
Likely cause
In Drift.run, persistence runs only when if drift_data: is truthy (laborious/workflows/drift.py). An empty drift result skips export_data_to_postgres, so the table stays empty.
Probable reasons:
- With
chunk_period='s'and only three target rows (30 s spacing),calculate_drift/DriftAnalysismay produce no output rows (insufficient data per chunk or internal filters). - Less likely here: time-window mismatch — timestamps are built from
datetime.now(UTC)withinterval: 60minutes fromdrift_base.json, so data should still fall in the window.
Recommended fixes
- Test data: Increase the number of second-spaced points (and/or span multiple chunk boundaries) so the analyzer reliably emits at least one chunk row.
- Product code: If sub-minute chunking is required to always produce metrics when any data exists, adjust
ModelMetrics/DriftAnalysisintegration for small-N second buckets. - Diagnostics: Run the same scenario with
pytest -sand confirm logs for “emptydrift_data” vs export errors.
Solução a ser aplicada: Dividir em dois testes:
- Teste com dados suficientes para produzir pelo menos uma linha de drift
- Teste com dados insuficientes para produzir pelo menos uma linha de drift, mas ja esperando os erros e validando que nao foi persistido nada
Test affected
test_drift_chunk_period_seconds_preserves_seconds_in_chunk_start_date
Failure group C — Prediction REPEAT path: unique constraint on predictions (3 tests)
Error
psycopg2.errors.UniqueViolation: duplicate key value violates unique constraint unique_model_id_timestamp on sientia_data.predictions (model_id, timestamp).
Root cause
REPEAT is handled by Activities.repeat_last_prediction (registered from sientia_do.temporal.activities.postgres_sync.Postgres, via Storage inheritance in laborious/activities/storage.py). The E2E tests seed a prior row with a fixed timestamp:
# e2e/test_predictions_batch_prediction_process.py — insert_sample_prediction
VALUES ({model_id}, '2024-01-01 12:00:00+00:00', ...)
The helper assert_repeat expects two rows with the same (model_id, prediction, prediction_confidence, prediction_status) but compares only those columns — not timestamp (e2e/helpers.py). So the intended behavior is: duplicate business payload, not necessarily duplicate primary unique key (model_id, timestamp).
If repeat_last_prediction INSERTs a copy using the same timestamp as the last prediction, Postgres correctly rejects the second insert.
Recommended fixes
- Implement REPEAT as UPSERT: Use
ON CONFLICT (model_id, timestamp) DO UPDATE(or the project’s existingexport_data_to_postgreson_conflictpattern used inFormatAndExportPrediction) when writing the repeated prediction — if the product definition of REPEAT is “refresh same logical slot.” - Insert with the current batch timestamp: Copy numeric/status fields from the last row but set
timestampto the new batch instant (e.g. the workflow’slast_timestamp/ slice timestamp). This matchesassert_repeat, which does not assert ontimestamp. - Test-only change (weakest): Relax assertions or change seed data — only if production behavior is “duplicate key is expected” (unlikely).
Solução a ser aplicada:
Alterar o teste para usar o timestamp da ultima execucao do batch, ao inves do timestamp do primeiro batch.
2. Insert with the current batch timestamp: Copy numeric/status fields from the last row but set timestamp to the new batch instant (e.g. the workflow’s last_timestamp / slice timestamp). This matches assert_repeat, which does not assert on timestamp.
Tests affected
test_scenario_2_1_3_input_gate_triggers_repeattest_scenario_2_2_3_transform_gate_triggers_repeattest_scenario_2_3_3_predict_gate_triggers_repeat
Passing areas (sanity check)
All scenarios in test_child_workflows_e2e.py, test_minimal_retrain.py, test_minio_offload.py, test_predictions_batch_format_export.py, test_predictions_batch_main_workflow.py (except the three REPEAT cases above), and test_simple_metrics.py passed in this run.
Suggested order of work
- Fix drift
valueNULL — unblocks two high-value drift E2Es and may clarify the chunk-seconds scenario if exports start succeeding consistently. - Fix
repeat_last_predictionuniqueness — unblocks three prediction-process E2Es; implementation likely lives insientia_doPostgres activities, not in this repo. - Revisit
test_drift_chunk_period_secondsdata volume / expectations after drift export is stable.