# Test Scenarios for PI Web API Scouter Workflow This document describes all possible test scenarios for the `pi_web_api_scouter` workflow and its child workflow `core_scouter`. ## Workflow Overview The `pi_web_api_scouter` workflow: 1. Retrieves tag values from PI Web API 2. Delegates processing to `core_scouter` child workflow which: - Applies data quality gates - Aggregates data - Groups and holds data in Redis - Exports to PostgreSQL - Writes metrics - Optionally stores debug data package --- ## 1. PI Web API Scouter - Main Workflow Scenarios ### 1.1 Success Scenarios #### Scenario 1.1.1: Happy Path - Complete Success **Description**: Workflow completes successfully with valid data from PI Web API **Input**: - Valid `model_name`, `model_id`, `schedule_name` - Valid `pi_web_api_query` with endpoint, period, max_count, api_timeout - Valid `model_tags` with webids and configurations - Valid filters, schema, table_name, retention_time **Expected Behavior**: - `get_tag_values` returns non-empty list of records - Workflow proceeds to `core_scouter` - All activities execute successfully - Data is stored in PostgreSQL - Metrics are written - Workflow completes without errors **Assertions**: - PI Web API client called once with correct parameters - Data exists in SQLite (PostgreSQL substitute) - Data cached in Redis - Metrics written - No errors raised --- #### Scenario 1.1.2: Success with Multiple Tags **Description**: Workflow processes multiple tags successfully **Input**: - Multiple tags in `model_tags` (3+ tags) - Each tag has valid webid, aggr_function, data_range, frequency **Expected Behavior**: - All tags retrieved from PI Web API - All tags processed through quality gates - All tags aggregated correctly - All tags stored in database **Assertions**: - Number of records matches number of tags - All tags present in final data - Aggregation applied per tag configuration --- #### Scenario 1.1.3: Success with Debug Data Package Enabled **Description**: Workflow completes with `debug_data_package=True` **Input**: - All standard input - `debug_data_package: True` **Expected Behavior**: - Normal workflow execution - `store_data_package` activity called - Data package stored in Redis **Assertions**: - `store_data_package` called once - Data package key exists in Redis - Package contains both `data` and `held_data` --- ### 1.2 Early Exit Scenarios #### Scenario 1.2.1: Empty Data from PI Web API **Description**: PI Web API returns empty data **Input**: - Valid configuration - PI Web API returns empty DataFrame or empty list **Expected Behavior**: - `get_tag_values` returns empty list `[]` - Workflow checks `if not data:` and returns early - `core_scouter` is NOT called - Workflow completes without error **Assertions**: - PI Web API called once - `core_scouter` NOT called - No data in PostgreSQL - No data in Redis (except possibly from previous runs) --- #### Scenario 1.2.2: None Returned from PI Web API **Description**: PI Web API returns None **Input**: - Valid configuration - PI Web API returns None **Expected Behavior**: - `get_tag_values` returns None - Workflow checks `if not data:` and returns early - `core_scouter` is NOT called **Assertions**: - PI Web API called once - `core_scouter` NOT called - Workflow completes without error --- ### 1.3 Error Scenarios #### Scenario 1.3.1: PI Web API Connection Error **Description**: PI Web API client raises connection error **Input**: - Valid configuration - PI Web API client raises `PIMSRequestError` or connection exception **Expected Behavior**: - `get_tag_values` catches exception - Sends notification with `PI_WEB_API_REQUEST_ERROR` - Raises exception (workflow fails after retries) **Assertions**: - Notification sent with correct error details - Exception propagated to workflow - Workflow fails (after retry policy exhausted) - `core_scouter` NOT called --- #### Scenario 1.3.2: PI Web API Timeout **Description**: PI Web API request times out **Input**: - Valid configuration - `api_timeout` set to low value - PI Web API takes longer than timeout **Expected Behavior**: - Request times out - Exception raised - Notification sent - Workflow fails after retries **Assertions**: - Timeout exception caught - Notification sent - Workflow fails --- #### Scenario 1.3.3: Invalid Endpoint **Description**: Invalid PI Web API endpoint provided **Input**: - Invalid endpoint path in `pi_web_api_query` **Expected Behavior**: - PI Web API client raises error - Notification sent - Workflow fails **Assertions**: - Error notification sent - Workflow fails --- #### Scenario 1.3.4: Missing Required Input Fields **Description**: Missing required input fields **Input**: - Missing `model_id`, `model_name`, `schedule_name`, or `pi_web_api_query` **Expected Behavior**: - KeyError raised when accessing missing fields - Workflow fails immediately **Assertions**: - KeyError or similar exception - Workflow fails before any activity execution --- ## 2. CoreScouter - Child Workflow Scenarios ### 2.1 Success Scenarios #### Scenario 2.1.1: Complete Processing Success **Description**: All stages complete successfully **Input**: - Valid data from parent workflow - Valid filters, model_tags, schema, table_name - `fill_missing_tags: False` - `debug_data_package: False` **Expected Behavior**: - `data_quality_gate` filters data - `aggregate_data` aggregates by tag - `group_and_hold_data` stores in Redis - `export_data_to_postgres` writes to database - `write_metrics` records metrics - Workflow completes **Assertions**: - All activities called in correct order - Data in PostgreSQL - Data in Redis - Metrics written - `store_data_package` NOT called --- #### Scenario 2.1.2: Success with Data Quality Filters **Description**: Data quality filters applied successfully **Input**: - Data with some quality issues - Filters configured with `NULL_VALUES_FILTER` or `OUT_OF_BOUNDS_FILTER` - Policy set to `DISCARD` or `WARN` **Expected Behavior**: - Quality gate identifies issues - Notification sent (WARNING level) - If policy is `DISCARD`, bad rows removed - Remaining data processed normally **Assertions**: - Quality issues detected - Notification sent - Bad data discarded if policy is `DISCARD` - Good data processed --- #### Scenario 2.1.3: Success with Different Aggregation Functions **Description**: Different aggregation functions applied correctly **Input**: - Multiple tags with different `aggr_function`: `avg`, `mdn`, `max`, `min`, `lts` - Time-series data with multiple points per tag **Expected Behavior**: - Each tag aggregated with its configured function - Aggregated values correct for each function type **Assertions**: - `avg` calculates mean correctly - `mdn` calculates median correctly - `max` returns maximum value - `min` returns minimum value - `lts` returns latest value --- #### Scenario 2.1.4: Success with Fill Missing Tags **Description**: Missing tags filled with None **Input**: - `fill_missing_tags: True` - Some tags missing from data **Expected Behavior**: - Missing tags added to `data_hold` with value `None` - All expected tags present in final data **Assertions**: - Missing tags present with `None` value - All model_tags represented in output --- ### 2.2 Early Exit Scenarios #### Scenario 2.2.1: Empty Data After Grouping **Description**: `group_and_hold_data` returns empty dict **Input**: - Data that results in empty `held_data` after grouping **Expected Behavior**: - `group_and_hold_data` returns `{}` - Workflow checks `if held_data == {}:` and returns early - `export_data_to_postgres` NOT called - `write_metrics` NOT called - `store_data_package` NOT called **Assertions**: - Early return after grouping - No database export - No metrics written - Workflow completes without error --- #### Scenario 2.2.2: Zero Affected Rows After Export **Description**: PostgreSQL export returns zero affected rows **Input**: - Data that results in `affected_rows: 0` from export **Expected Behavior**: - `export_data_to_postgres` returns `{'affected_rows': 0}` - Workflow checks `if data_exported.get('affected_rows', 0) <= 0:` and returns early - `write_metrics` NOT called - `store_data_package` NOT called **Assertions**: - Early return after export - No metrics written - Workflow completes without error --- ### 2.3 Error Scenarios #### Scenario 2.3.1: Data Quality Gate Error **Description**: Error during quality gate processing **Input**: - Invalid filter configuration - Filter function raises exception **Expected Behavior**: - Exception caught in quality gate - Notification sent with `DATA_QUALITY_GATE_ISSUES` - Exception propagated (workflow fails after retries) **Assertions**: - Error notification sent - Workflow fails --- #### Scenario 2.3.2: Aggregation Error **Description**: Error during data aggregation **Input**: - Invalid aggregation function - Data format issues **Expected Behavior**: - Invalid function sends notification with `AGGREGATION_ISSUES` - Returns `'continue'` for invalid function (skips that tag) - Other errors raise exception **Assertions**: - Invalid function handled gracefully - Other errors cause workflow failure --- #### Scenario 2.3.3: Redis Connection Error **Description**: Redis unavailable during `group_and_hold_data` **Input**: - Valid data - Redis connection fails **Expected Behavior**: - `redis_repository.get()` or `redis_repository.set()` raises exception - Notification sent with `REDIS_GET_ERROR` or `REDIS_SET_ERROR` - Exception propagated (workflow fails after retries) **Assertions**: - Error notification sent - Workflow fails --- #### Scenario 2.3.4: PostgreSQL Connection Error **Description**: PostgreSQL unavailable during export **Input**: - Valid data - PostgreSQL connection fails **Expected Behavior**: - `export_data_to_postgres` raises exception - Notification sent with `ERROR_EXPORTING_DATA_TO_POSTGRES` - Exception propagated (workflow fails after retries) **Assertions**: - Error notification sent - Workflow fails --- #### Scenario 2.3.5: PostgreSQL Unique Constraint Violation **Description**: Duplicate data violates unique constraint **Input**: - Data with duplicate `model_id`, `timestamp`, `variable` combination - `on_conflict: 'ignore'` configured **Expected Behavior**: - PostgreSQL handles conflict with `ON CONFLICT DO NOTHING` - `affected_rows` may be 0 for duplicates - Workflow continues normally **Assertions**: - No exception raised - Duplicates ignored - Workflow continues --- ## 3. Activity-Specific Scenarios ### 3.1 get_tag_values Activity #### Scenario 3.1.1: Success with Valid WebIds **Input**: All webids valid and present **Expected**: Returns list of records with timestamp, name, value, tag #### Scenario 3.1.2: Some WebIds are None **Input**: Some webids in `model_tags` are `None` **Expected**: None webids filtered out, only valid webids queried #### Scenario 3.1.3: DataFrame with NaN Values **Input**: PI Web API returns DataFrame with NaN values **Expected**: NaN values handled, data normalized correctly #### Scenario 3.1.4: Timestamp Normalization **Input**: Multiple timestamps in response **Expected**: All timestamps normalized to max timestamp value --- ### 3.2 data_quality_gate Activity #### Scenario 3.2.1: No Filters Configured **Input**: Empty `filters: {}` **Expected**: Data passes through unchanged, filtered by model_tags only #### Scenario 3.2.2: NULL_VALUES_FILTER with DISCARD Policy **Input**: Data with null values, policy `DISCARD` **Expected**: Null rows removed, notification sent #### Scenario 3.2.3: OUT_OF_BOUNDS_FILTER with WARN Policy **Input**: Data outside range, policy `WARN` **Expected**: Notification sent, data kept #### Scenario 3.2.4: Unknown Filter Type **Input**: Filter name not in `quality_gate_filters` **Expected**: Warning logged, filter skipped, processing continues #### Scenario 3.2.5: Filter Removes All Data **Input**: Filter that removes all rows **Expected**: Empty DataFrame returned, processing continues --- ### 3.3 aggregate_data Activity #### Scenario 3.3.1: Single Value Per Tag **Input**: One data point per tag **Expected**: Fast path returns value directly #### Scenario 3.3.2: Multiple Values - Latest (lts) **Input**: Multiple points, `aggr_function: 'lts'` **Expected**: Returns last value in sorted order #### Scenario 3.3.3: Multiple Values with NaN **Input**: Some NaN values in series **Expected**: NaN values dropped before aggregation #### Scenario 3.3.4: All NaN Values **Input**: All values are NaN **Expected**: Returns `None`, tag skipped #### Scenario 3.3.5: Invalid Aggregation Function **Input**: Unknown `aggr_function` **Expected**: Notification sent, returns `'continue'`, tag skipped #### Scenario 3.3.6: Empty DataFrame After Filtering **Input**: No data after quality gate **Expected**: Returns empty DataFrame dict --- ### 3.4 group_and_hold_data Activity #### Scenario 3.4.1: First Run - No Existing Data **Input**: No existing data in Redis for key **Expected**: Creates new `data_hold` dict, stores in Redis #### Scenario 3.4.2: Subsequent Run - Existing Data **Input**: Existing `data_hold` in Redis **Expected**: Merges new data with existing, updates timestamp #### Scenario 3.4.3: Removed Tags Cleanup **Input**: Tags removed from `model_tags` **Expected**: Removed tags deleted from `data_hold` #### Scenario 3.4.4: Empty Input Data **Input**: Empty DataFrame **Expected**: Returns empty dict, warning logged #### Scenario 3.4.5: Redis Get Error **Input**: Redis get operation fails **Expected**: Notification sent, exception raised #### Scenario 3.4.6: Redis Set Error **Input**: Redis set operation fails **Expected**: Notification sent, exception raised --- ### 3.5 export_data_to_postgres Activity #### Scenario 3.5.1: Successful Insert **Input**: Valid data, no conflicts **Expected**: Data inserted, `affected_rows > 0` #### Scenario 3.5.2: Conflict with Ignore Policy **Input**: Duplicate data, `on_conflict: 'ignore'` **Expected**: Duplicates ignored, `affected_rows` may be less than total #### Scenario 3.5.3: Conflict with Replace Policy **Input**: Duplicate data, `on_conflict: 'replace'` **Expected**: Duplicates updated, `affected_rows` includes updates #### Scenario 3.5.4: Timestamp Conversion **Input**: String timestamps in data **Expected**: Timestamps converted to datetime format #### Scenario 3.5.5: Database Connection Error **Input**: Database unavailable **Expected**: Exception raised, notification sent --- ### 3.6 write_metrics Activity #### Scenario 3.6.1: Success with Valid Values **Input**: Data with non-None values **Expected**: Metrics written for all non-None values #### Scenario 3.6.2: Some None Values **Input**: Some values are None **Expected**: None values skipped, only non-None values written #### Scenario 3.6.3: All None Values **Input**: All values are None **Expected**: No metrics written, activity completes --- ### 3.7 store_data_package Activity #### Scenario 3.7.1: Success **Input**: Valid data and held_data **Expected**: Package stored in Redis with TTL 120 #### Scenario 3.7.2: Redis Error **Input**: Redis set fails **Expected**: Notification sent, exception raised --- ## 4. Integration Scenarios ### 4.1 End-to-End Scenarios #### Scenario 4.1.1: Complete Happy Path **Description**: Full workflow from API to database **Flow**: 1. PI Web API returns data 2. Quality gate passes 3. Aggregation succeeds 4. Redis storage succeeds 5. PostgreSQL export succeeds 6. Metrics written 7. Debug package stored (if enabled) **Assertions**: - All activities called - Data in all storage layers - No errors --- #### Scenario 4.1.2: Partial Failure with Retry **Description**: Activity fails, retries succeed **Flow**: 1. First attempt fails (e.g., Redis timeout) 2. Retry policy triggers 3. Second attempt succeeds 4. Workflow continues **Assertions**: - Retry policy applied - Workflow eventually succeeds - Error logged but not fatal --- #### Scenario 4.1.3: Complete Failure After Retries **Description**: Activity fails after all retries exhausted **Flow**: 1. Activity fails repeatedly 2. Retry policy exhausted 3. Workflow fails **Assertions**: - All retries attempted - Workflow fails with error - Error notification sent --- ## 5. Edge Cases and Boundary Conditions ### 5.1 Data Edge Cases #### Scenario 5.1.1: Very Large Dataset **Input**: Thousands of data points **Expected**: Handles efficiently, all processed #### Scenario 5.1.2: Single Data Point **Input**: One tag, one data point **Expected**: Processes correctly #### Scenario 5.1.3: Extreme Values **Input**: Very large or very small numeric values **Expected**: Handled correctly, no overflow #### Scenario 5.1.4: Special Characters in Tag Names **Input**: Tag names with special characters **Expected**: Handled correctly --- ### 5.2 Configuration Edge Cases #### Scenario 5.2.1: Very Short Retention Time **Input**: `retention_time: 1` (1 second) **Expected**: Data expires quickly but workflow completes #### Scenario 5.2.2: Very Long Retention Time **Input**: `retention_time: 86400` (1 day) **Expected**: Data persists for full duration #### Scenario 5.2.3: Max Count = 1 **Input**: `max_count: 1` **Expected**: Only latest value retrieved #### Scenario 5.2.4: Max Count = Large Number **Input**: `max_count: 10000` **Expected**: Many values retrieved and processed --- ### 5.3 Concurrent Execution Scenarios #### Scenario 5.3.1: Multiple Workflows Same Schedule **Input**: Two workflows with same `schedule_name` running concurrently **Expected**: Both complete, data merged correctly in Redis #### Scenario 5.3.2: Multiple Workflows Different Schedules **Input**: Multiple workflows with different `schedule_name` **Expected**: Each uses separate Redis keys, no interference --- ## 6. Performance Scenarios ### 6.1 Load Scenarios #### Scenario 6.1.1: High Throughput **Input**: Many tags, frequent execution **Expected**: Handles load efficiently #### Scenario 6.1.2: Large Payload **Input**: Large amount of data per tag **Expected**: Processes within timeout limits --- ## 7. Test Data Requirements ### 7.1 Valid Test Data Structure ```python { 'model_name': 'test_model', 'model_id': 'test_model_id', 'schedule_name': 'test_schedule', 'model_tags': { 'tag1': { 'webid': 'webid1', 'aggr_function': 'avg', 'data_range': [0, 100], 'frequency': 60000, }, }, 'trigger_laborious': False, 'filters': {}, 'schema': 'test_schema', 'table_name': 'test_table', 'retention_time': 3600, 'fill_missing_tags': False, 'debug_data_package': False, 'pi_web_api_query': { 'endpoint': '/streamsets/recorded', 'period': '*-1d', 'max_count': 10, 'api_timeout': 30, }, } ``` ### 7.2 Mock PI Web API Response ```python DataFrame({ 'timestamp': ['2024-01-01 12:00:00+0000', ...], 'name': ['tag1', 'tag2', ...], 'value': [10.5, 20.3, ...], 'tag': ['webid1', 'webid2', ...], }) ``` --- ## 8. Test Implementation Notes ### 8.1 Test Organization - Group tests by scenario category - Use descriptive test names matching scenario IDs - Share fixtures for common setup - Use parametrized tests for similar scenarios ### 8.2 Assertions Checklist For each scenario, verify: - [ ] Correct activities called - [ ] Correct parameters passed - [ ] Expected data in storage (SQLite/Redis) - [ ] Expected notifications sent - [ ] Expected metrics written - [ ] No unexpected errors - [ ] Workflow state correct ### 8.3 Mock Configuration - Mock PI Web API client responses - Use fake Redis (fakeredis) - Use fake MongoDB (mongomock) - Use SQLite for PostgreSQL - Mock notification handler - Mock metrics controller --- ## 9. Priority Scenarios ### High Priority (Must Test) 1. Scenario 1.1.1: Happy Path 2. Scenario 1.2.1: Empty Data 3. Scenario 1.3.1: API Connection Error 4. Scenario 2.1.1: Complete Processing 5. Scenario 2.2.1: Empty After Grouping 6. Scenario 2.3.4: PostgreSQL Error ### Medium Priority (Should Test) 1. Scenario 1.1.3: Debug Package 2. Scenario 2.1.2: Quality Filters 3. Scenario 2.1.3: Different Aggregations 4. Scenario 3.3.5: Invalid Aggregation 5. Scenario 3.5.2: Conflict Ignore ### Low Priority (Nice to Have) 1. Scenario 4.1.2: Retry Success 2. Scenario 5.1.1: Large Dataset 3. Scenario 5.3.1: Concurrent Execution