SIENTIAPDE-1430: Update README to document new features, expanded validation, and enhanced testing.

*   Documents support for Polynomial Regression, per-variable lag configuration, date range filtering, and NaN treatment options.
*   Includes details on RCE drift metrics and comprehensive business validation rules for training parameters.
*   Describes new JSON-based integration test scenarios with batch execution and updated 100% code coverage.
*   Refines architecture overview and example workflow parameters.
This commit is contained in:
Bruno Domingues
2025-12-19 11:54:08 -03:00
parent 44dca9b66b
commit 9a77257920

124
README.md
View File

@@ -79,6 +79,7 @@ An enterprise-grade ML model training orchestration platform built on Temporal.
### Core Functionality ### Core Functionality
- **ML Model Training Pipeline**: Complete training workflow from validation to deployment using MLFlow - **ML Model Training Pipeline**: Complete training workflow from validation to deployment using MLFlow
- **Polynomial Regression Support**: Configurable polynomial degree with interaction terms and mandatory scaler validation
- **Automated File Cleanup**: Scheduled cleanup of stale files from MinIO and local filesystem - **Automated File Cleanup**: Scheduled cleanup of stale files from MinIO and local filesystem
- **Temporal Workflow Orchestration**: Robust workflow management with granular retry policies and fault tolerance - **Temporal Workflow Orchestration**: Robust workflow management with granular retry policies and fault tolerance
- **Parameter Validation**: Defense-in-depth validation with business rules and type checking - **Parameter Validation**: Defense-in-depth validation with business rules and type checking
@@ -93,14 +94,19 @@ An enterprise-grade ML model training orchestration platform built on Temporal.
- **Notification System**: Integrated alerting and notification management via MongoDB - **Notification System**: Integrated alerting and notification management via MongoDB
- **Scalable Architecture**: Kubernetes-ready deployment with horizontal scaling support - **Scalable Architecture**: Kubernetes-ready deployment with horizontal scaling support
- **MLFlow Integration**: Seamless model and artifact persistence to MLFlow tracking server - **MLFlow Integration**: Seamless model and artifact persistence to MLFlow tracking server
- **Per-Variable Lag Configuration**: Flexible lag settings for each variable independently
- **Date Range Filtering**: Filter training data by start/end dates and removed intervals
- **NaN Treatment Options**: Configurable handling of missing values (drop, linear interpolation)
- **RCE Drift Metrics**: Reduced Coulomb Energy metrics for drift detection
### Development & Quality Assurance ### Development & Quality Assurance
- **Code Quality Tools**: Ruff (linting/formatting), mypy (type checking), Bandit (security analysis) - **Code Quality Tools**: Ruff (linting/formatting), mypy (type checking), Bandit (security analysis)
- **Automated Validation**: Pre-commit validation script (`validate.sh`) and CI/CD integration - **Automated Validation**: Pre-commit validation script (`validate.sh`) and CI/CD integration
- **Comprehensive Testing**: pytest with async support and **99%+ code coverage** 🎯 - **Comprehensive Testing**: pytest with async support and **100% code coverage** 🎯
- **Type Safety**: Static type checking with mypy for improved code reliability - **Type Safety**: Static type checking with mypy for improved code reliability
- **Coverage Visualization**: Integration with Coverage Gutters for real-time coverage feedback - **Coverage Visualization**: Integration with Coverage Gutters for real-time coverage feedback
- **Automated Versioning**: Semantic versioning based on branch patterns (release/*, feature/*, fix/*, rc/*) - **Automated Versioning**: Semantic versioning based on branch patterns (release/*, feature/*, fix/*, rc/*)
- **Integration Test Scenarios**: JSON-based test scenarios with batch execution support
## Architecture ## Architecture
@@ -179,6 +185,8 @@ The Model Manager system uses a Temporal-based workflow architecture with clear
- Integration with TrainingRepository for business logic separation - Integration with TrainingRepository for business logic separation
- MLFlow model saving and artifact management - MLFlow model saving and artifact management
- MinIO object storage operations - MinIO object storage operations
- **Polynomial Regression**: Support for configurable degree and interaction terms
- **Training Predictions**: Calculates y_train_pred before denormalization for accurate metrics
- **Cleanup**: File and directory cleanup operations - **Cleanup**: File and directory cleanup operations
- `cleanup_minio_files()`: Removes stale files from MinIO based on timestamp prefixes - `cleanup_minio_files()`: Removes stale files from MinIO based on timestamp prefixes
- `cleanup_temp_directories()`: Cleans local temporary directories - `cleanup_temp_directories()`: Cleans local temporary directories
@@ -198,8 +206,8 @@ The Model Manager system uses a Temporal-based workflow architecture with clear
- `training_repository.py`: Training business logic and operations - `training_repository.py`: Training business logic and operations
- `model_repository.py`: MLFlow artifact generation and model persistence - `model_repository.py`: MLFlow artifact generation and model persistence
- **Models**: Data models and schemas - **Models**: Data models and schemas
- `train_model_params.py`: Training parameters model - `train_model_params.py`: Training parameters model with comprehensive validation
- `train_model_result.py`: Training result model - `train_model_result.py`: Training result model (includes y_train_pred)
- `experiment_status.py`: Experiment status enum - `experiment_status.py`: Experiment status enum
- **Key Features**: - **Key Features**:
- Environment variable-based configuration with sensible defaults - Environment variable-based configuration with sensible defaults
@@ -276,7 +284,7 @@ The **TrainModel** workflow orchestrates the complete ML model training pipeline
{ {
"experiment_run_id": 123, "experiment_run_id": 123,
"target_variable": "price", "target_variable": "price",
"variable_columns": ["feature1", "feature2", "price"], "variable_columns": ["feature1", "feature2"],
"train_size": 80, "train_size": 80,
"shuffle": true, "shuffle": true,
"use_scaler": true, "use_scaler": true,
@@ -285,14 +293,22 @@ The **TrainModel** workflow orchestrates the complete ML model training pipeline
"file_name": "training_data.csv", "file_name": "training_data.csv",
"line_separator": "\n", "line_separator": "\n",
"decimal_separator": ".", "decimal_separator": ".",
"lag_train": 5, "lag_train": {"feature1": 0, "feature2": 2},
"lag_val": 3, "lag_val": {"feature1": 0, "feature2": 1},
"rem_static_win": false, "rem_static_win": false,
"low_lim": {"feature1": 0.0, "feature2": 0.0, "price": 0.0}, "low_lim": {"feature1": 0.0, "feature2": 0.0},
"upp_lim": {"feature1": 100.0, "feature2": 100.0, "price": 1000.0}, "upp_lim": {"feature1": 100.0, "feature2": 100.0},
"window": 10, "window": 10,
"experiment_name": "production_model_v1", "experiment_name": "production_model_v1",
"removed_intervals": [] "removed_intervals": [],
"model_name": "Linear Regression",
"degree": 1,
"interaction_only": false,
"nan_treatment": "drop",
"start_date": null,
"end_date": null,
"scaler_name": "Standard Scaler",
"support_filters": {}
} }
``` ```
@@ -328,14 +344,21 @@ The workflow implements 5 different retry policies optimized for each operation
#### Business Validation Rules #### Business Validation Rules
The workflow validates 10 business rules beyond type checking: The workflow validates comprehensive business rules beyond type checking:
1. **train_size**: Must be between 1-99% 1. **train_size**: Must be between 10-100%
2. **variable_columns**: Cannot be empty 2. **variable_columns**: Cannot be empty
3. **lag_train, lag_val, window**: Must be positive integers 3. **lag_train, lag_val**: Per-variable dictionaries with non-negative values
4. **low_lim/upp_lim**: Must have same keys and low < upp for each variable 4. **window**: Must be non-negative integer
5. **target_variable**: Must be in variable_columns 5. **low_lim/upp_lim**: Must have same keys and low < upp for each variable
6. **bucket_name, file_name, experiment_name**: Cannot be empty or whitespace 6. **target_variable**: Cannot be empty
7. **bucket_name, file_name, experiment_name**: Cannot be empty or whitespace
8. **degree**: Must be at least 1; must be >= 2 for Polynomial Regression
9. **nan_treatment**: Must be one of 'drop', 'linear interpolation', 'fill linear'
10. **scaler_name**: Must be 'Standard Scaler' or 'None'
11. **model_name**: Must be 'Linear Regression' or 'Polynomial Regression'
12. **Polynomial Regression requires Scaler**: Models with degree > 1 must have a scaler to prevent numerical overflow
13. **Linear Regression requires degree 1**: Linear models must have degree = 1
### Cleanup Files Workflow (`cleanup_files.py`) ### Cleanup Files Workflow (`cleanup_files.py`)
@@ -884,14 +907,58 @@ The project includes integration tests that validate the complete training workf
#### Running Integration Tests #### Running Integration Tests
```bash ```bash
# List available scenarios
python scripts/run_training_test.py --list
# Run a specific test scenario # Run a specific test scenario
python scripts/run_training_test.py --scenario 01-linear-regression-basic python scripts/run_training_test.py --scenario 01-linear-regression-basic
# Run with custom data file # Run with custom CSV data file
python scripts/run_training_test.py --scenario 03-polynomial-regression-degree2 --data-file /path/to/data.csv python scripts/run_training_test.py --scenario 03-polynomial-regression-degree2 --csv /path/to/data.csv
# List available scenarios # Run ALL scenarios sequentially with summary report
ls docs/test-scenarios/ python scripts/run_training_test.py --all
# Run all scenarios with custom CSV
python scripts/run_training_test.py --all --csv docs/custom-data.csv
```
#### Batch Execution Output
When running all scenarios with `--all`, the script provides:
- Progress indicators for each scenario (`[1/10] Running scenario: ...`)
- Status symbols (✓ for passed, ✗ for failed)
- Final summary with total/passed/failed counts
- Detailed error messages for failed scenarios
- Exit code 0 if all pass, 1 if any fail
Example output:
```
Running 10 scenarios...
[1/10] Running scenario: 01-linear-regression-basic
Loaded scenario: 01-linear-regression-basic
Uploaded CSV to MinIO: test-model-data-20231219-120000.csv
Created experiment_run with ID: 42
Workflow started: train-model-test-abc123
[1/10] ✓ 01-linear-regression-basic
...
============================================================
SUMMARY
============================================================
Total: 10 | Passed: 9 | Failed: 1
============================================================
✓ PASSED:
- 01-linear-regression-basic
- 02-linear-regression-with-scaler
...
✗ FAILED:
- 05-linear-regression-with-lags
Error: Failed to start Temporal workflow: connection refused
``` ```
#### Test Scenarios #### Test Scenarios
@@ -905,7 +972,7 @@ Test scenarios are defined as JSON files in `docs/test-scenarios/`. Each scenari
| `03-polynomial-regression-degree2` | Polynomial regression (degree 2) | Requires scaler (mandatory) | | `03-polynomial-regression-degree2` | Polynomial regression (degree 2) | Requires scaler (mandatory) |
| `04-polynomial-regression-degree3` | Polynomial regression (degree 3) | Requires scaler (mandatory) | | `04-polynomial-regression-degree3` | Polynomial regression (degree 3) | Requires scaler (mandatory) |
| `05-linear-regression-with-lags` | Linear regression with lag features | Lag train/val configuration | | `05-linear-regression-with-lags` | Linear regression with lag features | Lag train/val configuration |
| `06-linear-regression-nan-interpolation` | Linear regression with NaN handling | `nanTreatment: "interpolate"` | | `06-linear-regression-nan-interpolation` | Linear regression with NaN handling | `nanTreatment: "linear interpolation"` |
| `07-linear-regression-static-window-removal` | Linear regression with static window removal | `remStaticWin: true` | | `07-linear-regression-static-window-removal` | Linear regression with static window removal | `remStaticWin: true` |
| `08-linear-regression-with-limits` | Linear regression with variable limits | `lowLim`/`uppLim` configuration | | `08-linear-regression-with-limits` | Linear regression with variable limits | `lowLim`/`uppLim` configuration |
| `09-polynomial-degree2-with-scaler-and-lags` | Complete polynomial scenario | Scaler + lags + degree 2 | | `09-polynomial-degree2-with-scaler-and-lags` | Complete polynomial scenario | Scaler + lags + degree 2 |
@@ -977,6 +1044,11 @@ The Model Manager system exposes comprehensive Prometheus metrics for operationa
- Training success/failure rates through notification system - Training success/failure rates through notification system
- Model save performance metrics - Model save performance metrics
- Experiment status tracking - Experiment status tracking
- **RCE Drift Metrics**: Reduced Coulomb Energy (RCE) for drift detection
- `silverman_radius`: Optimal bandwidth for kernel density estimation
- `rce_reference`: RCE value for reference data
- `rce_current`: RCE value for current data
- `rce_drift`: Drift score between reference and current distributions
### Cleanup Metrics ### Cleanup Metrics
- Cleanup execution success/failure rates - Cleanup execution success/failure rates
@@ -1060,7 +1132,7 @@ These timeouts control how long each activity in workflows can run before timing
### Code Quality & Testing ### Code Quality & Testing
The project maintains **99%+ code coverage** with comprehensive unit and integration tests. The project maintains **100% code coverage** with comprehensive unit and integration tests.
#### Running Tests #### Running Tests
@@ -1161,8 +1233,8 @@ sientia-dataops-model-manager/
│ ├── sientia/ # Sientia-specific implementations │ ├── sientia/ # Sientia-specific implementations
│ │ ├── __init__.py │ │ ├── __init__.py
│ │ ├── exceptions.py # Custom exceptions │ │ ├── exceptions.py # Custom exceptions
│ │ ├── metrics.py # Business metrics │ │ ├── metrics.py # Business metrics (includes RCE drift detection)
│ │ ├── models.py # ML model implementations │ │ ├── models.py # ML model implementations (Linear & Polynomial Regression)
│ │ ├── model_serving.py # Model serving utilities │ │ ├── model_serving.py # Model serving utilities
│ │ ├── reports.py # Report generation │ │ ├── reports.py # Report generation
│ │ └── utils.py # Utility functions │ │ └── utils.py # Utility functions
@@ -1173,7 +1245,7 @@ sientia-dataops-model-manager/
│ └── __init__.py │ └── __init__.py
├── scripts/ # Test and utility scripts ├── scripts/ # Test and utility scripts
│ ├── run_cleanup_test.py # Manual cleanup workflow test │ ├── run_cleanup_test.py # Manual cleanup workflow test
│ └── run_training_test.py # Manual training workflow test │ └── run_training_test.py # Training test with scenario support (--all for batch)
├── tests/ # Test suite ├── tests/ # Test suite
│ ├── activities/ # Activity tests │ ├── activities/ # Activity tests
│ ├── workflows/ # Workflow tests │ ├── workflows/ # Workflow tests
@@ -1182,6 +1254,8 @@ sientia-dataops-model-manager/
│ ├── schedules/ # Schedule tests │ ├── schedules/ # Schedule tests
│ └── sientia/ # Sientia module tests │ └── sientia/ # Sientia module tests
├── docs/ # Documentation and test data ├── docs/ # Documentation and test data
│ ├── test-scenarios/ # JSON test scenario files for integration tests
│ └── test-model-data.csv # Sample CSV data for testing
├── .github/workflows/ # CI/CD workflows ├── .github/workflows/ # CI/CD workflows
│ ├── quality-gate.yml # PR quality checks │ ├── quality-gate.yml # PR quality checks
│ └── deploy.yml # Deployment workflow │ └── deploy.yml # Deployment workflow
@@ -1207,7 +1281,7 @@ sientia-dataops-model-manager/
### Test Coverage Guidelines ### Test Coverage Guidelines
- **Minimum coverage**: 80% (enforced by CI/CD) - **Minimum coverage**: 80% (enforced by CI/CD)
- **Current coverage**: 99%+ 🎯 - **Current coverage**: 100% 🎯
- **Test all branches**: Use Coverage Gutters to identify uncovered lines - **Test all branches**: Use Coverage Gutters to identify uncovered lines
- **Mock external dependencies**: Use `unittest.mock` for external services - **Mock external dependencies**: Use `unittest.mock` for external services
- **Async testing**: Use `pytest-asyncio` for async activities and workflows - **Async testing**: Use `pytest-asyncio` for async activities and workflows