SIENTIAPDE-1478
Update version in init_orchestration.ipynb to 3.11.14 and enhance README.md with detailed architecture principles, execution flows, and improved filtering mechanisms in workflows. Refactor orchestrator_functions.py for clarity in filter configuration and add support for PI Web API output configuration in predictions_batch. Update tests to reflect new configurations.
This commit is contained in:
131
README.md
131
README.md
@@ -60,15 +60,15 @@ The SIENTIA DataOps Orchestrator uses a Temporal-based workflow architecture wit
|
||||
### Architecture Principles
|
||||
|
||||
#### 1. **Separation of Concerns**
|
||||
- **Worker Layer**: Manages Temporal workers, task queues, and application lifecycle
|
||||
- **Workflow Layer**: Orchestrates business logic and process coordination
|
||||
- **Activity Layer**: Implements specific operations and external system interactions
|
||||
- **Data Layer**: Handles data persistence, caching, and external service connections
|
||||
- **Worker Layer**: Manages Temporal workers, task queues, Prometheus metrics, and application lifecycle
|
||||
- **Workflow Layer**: Orchestrates business logic and process coordination with subworkflows for modularity
|
||||
- **Activity Layer**: Implements specific operations and external system interactions with error handling
|
||||
- **Data Layer**: Handles data persistence (MongoDB, PostgreSQL), caching (Redis), and external service connections
|
||||
|
||||
#### 2. **Task Queue Isolation**
|
||||
- **Orchestrator Queue**: Pipeline and resource management workflows
|
||||
- **Alerts Queue**: Real-time error notification workflows
|
||||
- **Reports Queue**: Scheduled reporting and summary workflows
|
||||
- **Orchestrator Queue**: Pipeline and resource management workflows (orchestrator workflow)
|
||||
- **Alerts Queue**: Real-time error notification workflows (alerts workflow, load_notification_package, process_notifications subworkflows)
|
||||
- **Reports Queue**: Scheduled reporting and summary workflows (reports workflow, load_notification_package, process_notifications subworkflows)
|
||||
|
||||
#### 3. **Fault Tolerance & Resilience**
|
||||
- **Automatic Retry Policies**: Configurable retry strategies for transient failures
|
||||
@@ -97,12 +97,11 @@ The **Orchestrator** workflow is the main coordination workflow that manages pip
|
||||
- **Infrastructure Management**: Creates, updates, and deletes workflow schedules
|
||||
|
||||
**Execution Flow**:
|
||||
1. **Configuration Loading**: Retrieves pipeline and OPC server configurations from MongoDB
|
||||
2. **Resource Assessment**: Loads current OPC slots and active ingestors from Redis
|
||||
3. **Schedule Processing**: Formats configurations for different workflow types
|
||||
4. **Deployment Operations**: Creates, updates, or deletes Temporal schedules
|
||||
5. **Resource Updates**: Updates OPC slots and MongoDB timestamps
|
||||
6. **Reporting**: Generates comprehensive orchestration reports
|
||||
1. **Parallel Data Loading**: Concurrently loads pipelines, OPC servers, orchestrated schedules, OPC slots, and active ingestors
|
||||
2. **Parallel Processing**: Formats orchestrated schedules and processes new schedules and slots
|
||||
3. **Parallel Config Creation**: Creates schedule and slot action configurations, normalizes schedules, and creates collections with TTL indexes
|
||||
4. **Parallel Operations**: Executes slot deletions, slot updates, schedule deletions, schedule creations, and schedule updates concurrently
|
||||
5. **Parallel Reports & Timestamps**: Generates orchestration reports and updates MongoDB timestamps for created/updated/deleted pipelines
|
||||
|
||||
**Input Parameters**:
|
||||
```json
|
||||
@@ -188,12 +187,13 @@ The **Alerts** workflow processes and sends real-time error notifications to con
|
||||
- **Persistent Monitoring**: Tracks and escalates persistent issues
|
||||
|
||||
**Execution Flow**:
|
||||
1. **Notification Loading**: Retrieves ERROR-level notifications from MongoDB
|
||||
2. **Timestamp Filtering**: Applies incremental processing using Redis timestamps
|
||||
3. **Group Filtering**: Filters notifications by user group configurations
|
||||
4. **TTL Processing**: Checks notification cache to prevent duplicate alerts
|
||||
5. **Email Generation**: Creates HTML email content for each group
|
||||
6. **Delivery & Logging**: Sends emails and logs results to PostgreSQL
|
||||
1. **Notification Loading**: Subworkflow loads ERROR-level notifications from MongoDB with timestamp filtering
|
||||
2. **Configuration Loading**: Loads active receiver group configurations from MongoDB
|
||||
3. **Alert Filtering**: Applies TTL-based filtering to identify core_alerts (new) and persistent_alerts (ongoing issues)
|
||||
4. **Group Filtering**: Filters notifications by user group content policies and ignore lists
|
||||
5. **Email Processing**: Subworkflow generates HTML emails and sends to receiver groups
|
||||
6. **Cache Storage**: Stores sent notification cache in Redis with configurable TTL to prevent duplicates
|
||||
7. **Audit Logging**: Logs delivery results to PostgreSQL for monitoring
|
||||
|
||||
**Input Parameters**:
|
||||
```json
|
||||
@@ -231,12 +231,11 @@ The **Reports** workflow generates and sends scheduled comprehensive reports to
|
||||
- **Audit Trail**: Complete logging of report delivery
|
||||
|
||||
**Execution Flow**:
|
||||
1. **Data Collection**: Loads all notifications from MongoDB (any level)
|
||||
2. **Timestamp Processing**: Uses incremental loading with Redis timestamps
|
||||
3. **Group Processing**: Applies user group filtering for report customization
|
||||
4. **Report Generation**: Creates HTML reports with comprehensive summaries
|
||||
5. **Distribution**: Sends reports to configured recipients
|
||||
6. **Audit Logging**: Records delivery status in PostgreSQL
|
||||
1. **Notification Loading**: Subworkflow loads all notifications from MongoDB (any level) with timestamp filtering
|
||||
2. **Configuration Loading**: Loads active receiver group configurations from MongoDB
|
||||
3. **Report Filtering**: Filters notifications by user group content policies (must include "reports") and ignore lists
|
||||
4. **Email Processing**: Subworkflow generates HTML reports organized by notification level and model, then sends to receiver groups
|
||||
5. **Audit Logging**: Logs delivery results to PostgreSQL for monitoring and tracking
|
||||
|
||||
**Input Parameters**:
|
||||
```json
|
||||
@@ -288,10 +287,10 @@ flowchart LR
|
||||
- `sending_configs` (list[dict]): Active receiver group configurations
|
||||
|
||||
**Key Activities**:
|
||||
- `get_last_data_timestamp`: Retrieves last processed timestamp from Redis
|
||||
- `find_documents_in_mongodb`: Loads receiver group configurations
|
||||
- `load_latest_data`: Loads notifications with timestamp filtering
|
||||
- `put_last_data_timestamp`: Updates last processed timestamp
|
||||
- `get_last_data_timestamp`: Retrieves last processed timestamp from Redis for incremental processing
|
||||
- `find_documents_in_mongodb`: Loads active receiver group configurations from MongoDB
|
||||
- `load_latest_data`: Loads notifications with timestamp filtering from notification_queue collection
|
||||
- `put_last_data_timestamp`: Updates last processed timestamp in Redis with 5-hour TTL
|
||||
|
||||
**Architecture**:
|
||||
|
||||
@@ -343,10 +342,10 @@ flowchart LR
|
||||
- `log_report` (list[dict]): Detailed delivery status for each notification
|
||||
|
||||
**Key Activities**:
|
||||
- `build_email_html`: Generates HTML content using Jinja2 templates
|
||||
- `send_email`: Delivers emails to receiver groups with error handling
|
||||
- `format_log_report`: Formats delivery results for database storage
|
||||
- `export_data_to_postgres`: Stores audit logs in PostgreSQL
|
||||
- `build_email_html`: Generates HTML content using Jinja2 templates organized by notification level and model
|
||||
- `send_email`: Delivers emails to receiver groups with attachment support and automatic SMTP reconnection
|
||||
- `format_log_report`: Formats delivery results into DataFrame structure for database storage, aggregating by notification ID and trigger
|
||||
- `export_data_to_postgres`: Stores audit logs in PostgreSQL with timestamp conversion
|
||||
|
||||
**Architecture**:
|
||||
|
||||
@@ -373,10 +372,11 @@ The orchestrator includes an advanced notification filtering system that prevent
|
||||
- **Ignore Lists**: Supports notification exclusion per group
|
||||
|
||||
### **Report Filtering (`filter_notification_reports`)**
|
||||
- **Purpose**: Filters all notification levels for comprehensive reports
|
||||
- **Purpose**: Filters all notification levels for comprehensive scheduled reports
|
||||
- **Comprehensive Coverage**: Includes ERROR, WARNING, INFO, and DEBUG levels
|
||||
- **Group Customization**: Applies different content policies per receiver group
|
||||
- **Scheduled Processing**: Designed for regular report generation
|
||||
- **Group Customization**: Applies different content policies per receiver group with ignore list support
|
||||
- **Scheduled Processing**: Designed for regular report generation without TTL-based duplicate prevention
|
||||
- **Duplicate Prevention**: Prevents duplicate notifications within the same report using trigger and notification ID keys
|
||||
|
||||
### **Notification Caching (`store_notification_cache`)**
|
||||
- **Purpose**: Manages Redis-based notification cache for TTL enforcement
|
||||
@@ -389,32 +389,40 @@ The orchestrator includes an advanced notification filtering system that prevent
|
||||
#### **Worker (`orchestrator/worker/worker.py`)**
|
||||
- **Purpose**: Main application orchestrator managing Temporal workers and task queues
|
||||
- **Responsibilities**:
|
||||
- Temporal client initialization and connection management
|
||||
- Temporal client initialization and connection management with SDK metrics
|
||||
- Worker lifecycle management and graceful shutdown
|
||||
- Task queue configuration (orchestrator, alerts, reports)
|
||||
- Prometheus metrics server initialization
|
||||
- Task queue configuration (orchestrator, alerts, reports) with dedicated workers
|
||||
- Prometheus metrics server initialization on HTTP_METRICS_PORT
|
||||
- Temporal SDK metrics server initialization on HTTP_SDK_METRICS_PORT
|
||||
- Notification handler setup and configuration
|
||||
- **Key Features**:
|
||||
- Multi-queue worker management with automatic scaling
|
||||
- Health check endpoints for Kubernetes liveness/readiness probes
|
||||
- Graceful shutdown with cleanup procedures
|
||||
- Multi-queue worker management with three dedicated workers (orchestrator, alerts, reports)
|
||||
- Application health metrics (app_up gauge) for Kubernetes liveness/readiness probes
|
||||
- Graceful shutdown with cleanup procedures for all connections
|
||||
- Comprehensive error handling and metrics collection
|
||||
- Parallel worker execution using asyncio.gather
|
||||
|
||||
#### **Activities (`orchestrator/activities/`)**
|
||||
- **Activities**: Main activity orchestrator combining all operations
|
||||
- **TemporalManager**: Temporal schedule CRUD operations across namespaces
|
||||
- **SlotManager**: Redis-based OPC slot and cache management with notification filtering
|
||||
- **MongoDB**: Document operations, aggregations, and TTL management
|
||||
- **Email**: SMTP operations with HTML generation and attachment support
|
||||
- **Formatters**: Configuration processing, slot distribution algorithms, and notification filtering for reports
|
||||
- **Activities**: Main activity orchestrator combining all operations (TemporalManager, SlotManager, Formatters, MongoDB, Email, Postgres)
|
||||
- **TemporalManager**: Temporal schedule CRUD operations across scouter and laborious namespaces with search attributes
|
||||
- **SlotManager**: Redis-based OPC slot and cache management with notification filtering and TTL-based duplicate prevention
|
||||
- **MongoDB**: Document operations, aggregations, timestamp management, and TTL index creation
|
||||
- **Email**: SMTP operations with HTML generation, attachment support, and automatic reconnection handling
|
||||
- **Formatters**: Configuration processing, slot distribution algorithms, notification filtering for reports, and schedule/slot orchestration reporting
|
||||
- **Postgres**: PostgreSQL operations for audit logging and data export (via sientia-dataops-library)
|
||||
|
||||
#### **Utilities (`orchestrator/utils/`)**
|
||||
- **Connectors Configuration**: Database and service configuration management
|
||||
- **Email Builder**: HTML email template generation and formatting using Jinja2 templates
|
||||
- **Orchestrator Functions**: Pipeline configuration transformation utilities for scouter, predictions_batch, minimal_retrain, drift, and simple_metrics workflows
|
||||
- **Converters**: Data type conversion and validation utilities including frequency parsing
|
||||
- **Templates**: HTML email templates for alerts and reports
|
||||
- **Connectors Configuration**: Database and service configuration management from environment variables
|
||||
- **Email Builder**: HTML email template generation and formatting using Jinja2 templates with support for alerts and reports
|
||||
- **Orchestrator Functions**: Pipeline configuration transformation utilities supporting:
|
||||
- `scouter`: OPC UA data collection workflows
|
||||
- `pi_web_api_scouter`: PI Web API data collection workflows
|
||||
- `predictions_batch`: ML prediction workflows with OPC write-back and multi-stage filtering
|
||||
- `minimal_retrain`: Model retraining workflows with SQL queries
|
||||
- `drift`: Data drift detection workflows
|
||||
- `simple_metrics`: Model performance metrics computation workflows
|
||||
- **Converters**: Data type conversion and validation utilities including frequency parsing for Temporal schedules
|
||||
- **Templates**: HTML email templates for alerts and reports (email_template.html, general_template.html)
|
||||
|
||||
## 📋 Prerequisites
|
||||
|
||||
@@ -621,18 +629,19 @@ Temporal input configuration sample:
|
||||
The Orchestrator system exposes comprehensive Prometheus metrics:
|
||||
|
||||
### Application Metrics
|
||||
- `app_up`: Application health status (1=healthy, 0=unhealthy)
|
||||
- `email_sent_count`: Email delivery operation count by group
|
||||
- `app_up`: Application health status gauge (1=healthy, 0=unhealthy) labeled by pod_id
|
||||
- `email_sent_count`: Email delivery counter labeled by pod_id, model_name, pipeline_name, and email_group
|
||||
|
||||
### Workflow Metrics
|
||||
- Schedule creation, update, and deletion success rates
|
||||
- Notification processing times and error rates
|
||||
- Resource allocation and slot management metrics
|
||||
- Schedule creation, update, and deletion success rates (via notification reports)
|
||||
- Notification processing times and error rates (via PostgreSQL audit logs)
|
||||
- Resource allocation and slot management metrics (via notification reports)
|
||||
- Temporal SDK metrics exposed on HTTP_SDK_METRICS_PORT (default: 9091)
|
||||
|
||||
### Database Metrics
|
||||
- MongoDB query performance and connection health
|
||||
- Redis operation counts and response times
|
||||
- PostgreSQL export operations and audit log metrics
|
||||
- MongoDB query performance and connection health (via sientia-dataops-library)
|
||||
- Redis operation counts and response times (via sientia-dataops-library)
|
||||
- PostgreSQL export operations and audit log metrics (via sientia-dataops-library)
|
||||
|
||||
## 🧪 Testing
|
||||
|
||||
|
||||
Reference in New Issue
Block a user