Update requirements-dev.txt to add E2E testing dependencies: fakeredis and mongomock for in-memory testing, and include testcontainers for PostgreSQL support.
20 KiB
Test Scenarios for PI Web API Scouter Workflow
This document describes all possible test scenarios for the pi_web_api_scouter workflow and its child workflow core_scouter.
Workflow Overview
The pi_web_api_scouter workflow:
- Retrieves tag values from PI Web API
- Delegates processing to
core_scouterchild workflow which:- Applies data quality gates
- Aggregates data
- Groups and holds data in Redis
- Exports to PostgreSQL
- Writes metrics
- Optionally stores debug data package
1. PI Web API Scouter - Main Workflow Scenarios
1.1 Success Scenarios
Scenario 1.1.1: Happy Path - Complete Success
Description: Workflow completes successfully with valid data from PI Web API
Input:
- Valid
model_name,model_id,schedule_name - Valid
pi_web_api_querywith endpoint, period, max_count, api_timeout - Valid
model_tagswith webids and configurations - Valid filters, schema, table_name, retention_time
Expected Behavior:
get_tag_valuesreturns non-empty list of records- Workflow proceeds to
core_scouter - All activities execute successfully
- Data is stored in PostgreSQL
- Metrics are written
- Workflow completes without errors
Assertions:
- PI Web API client called once with correct parameters
- Data exists in SQLite (PostgreSQL substitute)
- Data cached in Redis
- Metrics written
- No errors raised
Scenario 1.1.2: Success with Multiple Tags
Description: Workflow processes multiple tags successfully
Input:
- Multiple tags in
model_tags(3+ tags) - Each tag has valid webid, aggr_function, data_range, frequency
Expected Behavior:
- All tags retrieved from PI Web API
- All tags processed through quality gates
- All tags aggregated correctly
- All tags stored in database
Assertions:
- Number of records matches number of tags
- All tags present in final data
- Aggregation applied per tag configuration
Scenario 1.1.3: Success with Debug Data Package Enabled
Description: Workflow completes with debug_data_package=True
Input:
- All standard input
debug_data_package: True
Expected Behavior:
- Normal workflow execution
store_data_packageactivity called- Data package stored in Redis
Assertions:
store_data_packagecalled once- Data package key exists in Redis
- Package contains both
dataandheld_data
1.2 Early Exit Scenarios
Scenario 1.2.1: Empty Data from PI Web API
Description: PI Web API returns empty data
Input:
- Valid configuration
- PI Web API returns empty DataFrame or empty list
Expected Behavior:
get_tag_valuesreturns empty list[]- Workflow checks
if not data:and returns early core_scouteris NOT called- Workflow completes without error
Assertions:
- PI Web API called once
core_scouterNOT called- No data in PostgreSQL
- No data in Redis (except possibly from previous runs)
Scenario 1.2.2: None Returned from PI Web API
Description: PI Web API returns None
Input:
- Valid configuration
- PI Web API returns None
Expected Behavior:
get_tag_valuesreturns None- Workflow checks
if not data:and returns early core_scouteris NOT called
Assertions:
- PI Web API called once
core_scouterNOT called- Workflow completes without error
1.3 Error Scenarios
Scenario 1.3.1: PI Web API Connection Error
Description: PI Web API client raises connection error
Input:
- Valid configuration
- PI Web API client raises
PIMSRequestErroror connection exception
Expected Behavior:
get_tag_valuescatches exception- Sends notification with
PI_WEB_API_REQUEST_ERROR - Raises exception (workflow fails after retries)
Assertions:
- Notification sent with correct error details
- Exception propagated to workflow
- Workflow fails (after retry policy exhausted)
core_scouterNOT called
Scenario 1.3.2: PI Web API Timeout
Description: PI Web API request times out
Input:
- Valid configuration
api_timeoutset to low value- PI Web API takes longer than timeout
Expected Behavior:
- Request times out
- Exception raised
- Notification sent
- Workflow fails after retries
Assertions:
- Timeout exception caught
- Notification sent
- Workflow fails
Scenario 1.3.3: Invalid Endpoint
Description: Invalid PI Web API endpoint provided
Input:
- Invalid endpoint path in
pi_web_api_query
Expected Behavior:
- PI Web API client raises error
- Notification sent
- Workflow fails
Assertions:
- Error notification sent
- Workflow fails
Scenario 1.3.4: Missing Required Input Fields
Description: Missing required input fields
Input:
- Missing
model_id,model_name,schedule_name, orpi_web_api_query
Expected Behavior:
- KeyError raised when accessing missing fields
- Workflow fails immediately
Assertions:
- KeyError or similar exception
- Workflow fails before any activity execution
2. CoreScouter - Child Workflow Scenarios
2.1 Success Scenarios
Scenario 2.1.1: Complete Processing Success
Description: All stages complete successfully
Input:
- Valid data from parent workflow
- Valid filters, model_tags, schema, table_name
fill_missing_tags: Falsedebug_data_package: False
Expected Behavior:
data_quality_gatefilters dataaggregate_dataaggregates by taggroup_and_hold_datastores in Redisexport_data_to_postgreswrites to databasewrite_metricsrecords metrics- Workflow completes
Assertions:
- All activities called in correct order
- Data in PostgreSQL
- Data in Redis
- Metrics written
store_data_packageNOT called
Scenario 2.1.2: Success with Data Quality Filters
Description: Data quality filters applied successfully
Input:
- Data with some quality issues
- Filters configured with
NULL_VALUES_FILTERorOUT_OF_BOUNDS_FILTER - Policy set to
DISCARDorWARN
Expected Behavior:
- Quality gate identifies issues
- Notification sent (WARNING level)
- If policy is
DISCARD, bad rows removed - Remaining data processed normally
Assertions:
- Quality issues detected
- Notification sent
- Bad data discarded if policy is
DISCARD - Good data processed
Scenario 2.1.3: Success with Different Aggregation Functions
Description: Different aggregation functions applied correctly
Input:
- Multiple tags with different
aggr_function:avg,mdn,max,min,lts - Time-series data with multiple points per tag
Expected Behavior:
- Each tag aggregated with its configured function
- Aggregated values correct for each function type
Assertions:
avgcalculates mean correctlymdncalculates median correctlymaxreturns maximum valueminreturns minimum valueltsreturns latest value
Scenario 2.1.4: Success with Fill Missing Tags
Description: Missing tags filled with None
Input:
fill_missing_tags: True- Some tags missing from data
Expected Behavior:
- Missing tags added to
data_holdwith valueNone - All expected tags present in final data
Assertions:
- Missing tags present with
Nonevalue - All model_tags represented in output
2.2 Early Exit Scenarios
Scenario 2.2.1: Empty Data After Grouping
Description: group_and_hold_data returns empty dict
Input:
- Data that results in empty
held_dataafter grouping
Expected Behavior:
group_and_hold_datareturns{}- Workflow checks
if held_data == {}:and returns early export_data_to_postgresNOT calledwrite_metricsNOT calledstore_data_packageNOT called
Assertions:
- Early return after grouping
- No database export
- No metrics written
- Workflow completes without error
Scenario 2.2.2: Zero Affected Rows After Export
Description: PostgreSQL export returns zero affected rows
Input:
- Data that results in
affected_rows: 0from export
Expected Behavior:
export_data_to_postgresreturns{'affected_rows': 0}- Workflow checks
if data_exported.get('affected_rows', 0) <= 0:and returns early write_metricsNOT calledstore_data_packageNOT called
Assertions:
- Early return after export
- No metrics written
- Workflow completes without error
2.3 Error Scenarios
Scenario 2.3.1: Data Quality Gate Error
Description: Error during quality gate processing
Input:
- Invalid filter configuration
- Filter function raises exception
Expected Behavior:
- Exception caught in quality gate
- Notification sent with
DATA_QUALITY_GATE_ISSUES - Exception propagated (workflow fails after retries)
Assertions:
- Error notification sent
- Workflow fails
Scenario 2.3.2: Aggregation Error
Description: Error during data aggregation
Input:
- Invalid aggregation function
- Data format issues
Expected Behavior:
- Invalid function sends notification with
AGGREGATION_ISSUES - Returns
'continue'for invalid function (skips that tag) - Other errors raise exception
Assertions:
- Invalid function handled gracefully
- Other errors cause workflow failure
Scenario 2.3.3: Redis Connection Error
Description: Redis unavailable during group_and_hold_data
Input:
- Valid data
- Redis connection fails
Expected Behavior:
redis_repository.get()orredis_repository.set()raises exception- Notification sent with
REDIS_GET_ERRORorREDIS_SET_ERROR - Exception propagated (workflow fails after retries)
Assertions:
- Error notification sent
- Workflow fails
Scenario 2.3.4: PostgreSQL Connection Error
Description: PostgreSQL unavailable during export
Input:
- Valid data
- PostgreSQL connection fails
Expected Behavior:
export_data_to_postgresraises exception- Notification sent with
ERROR_EXPORTING_DATA_TO_POSTGRES - Exception propagated (workflow fails after retries)
Assertions:
- Error notification sent
- Workflow fails
Scenario 2.3.5: PostgreSQL Unique Constraint Violation
Description: Duplicate data violates unique constraint
Input:
- Data with duplicate
model_id,timestamp,variablecombination on_conflict: 'ignore'configured
Expected Behavior:
- PostgreSQL handles conflict with
ON CONFLICT DO NOTHING affected_rowsmay be 0 for duplicates- Workflow continues normally
Assertions:
- No exception raised
- Duplicates ignored
- Workflow continues
3. Activity-Specific Scenarios
3.1 get_tag_values Activity
Scenario 3.1.1: Success with Valid WebIds
Input: All webids valid and present Expected: Returns list of records with timestamp, name, value, tag
Scenario 3.1.2: Some WebIds are None
Input: Some webids in model_tags are None
Expected: None webids filtered out, only valid webids queried
Scenario 3.1.3: DataFrame with NaN Values
Input: PI Web API returns DataFrame with NaN values Expected: NaN values handled, data normalized correctly
Scenario 3.1.4: Timestamp Normalization
Input: Multiple timestamps in response Expected: All timestamps normalized to max timestamp value
3.2 data_quality_gate Activity
Scenario 3.2.1: No Filters Configured
Input: Empty filters: {}
Expected: Data passes through unchanged, filtered by model_tags only
Scenario 3.2.2: NULL_VALUES_FILTER with DISCARD Policy
Input: Data with null values, policy DISCARD
Expected: Null rows removed, notification sent
Scenario 3.2.3: OUT_OF_BOUNDS_FILTER with WARN Policy
Input: Data outside range, policy WARN
Expected: Notification sent, data kept
Scenario 3.2.4: Unknown Filter Type
Input: Filter name not in quality_gate_filters
Expected: Warning logged, filter skipped, processing continues
Scenario 3.2.5: Filter Removes All Data
Input: Filter that removes all rows Expected: Empty DataFrame returned, processing continues
3.3 aggregate_data Activity
Scenario 3.3.1: Single Value Per Tag
Input: One data point per tag Expected: Fast path returns value directly
Scenario 3.3.2: Multiple Values - Latest (lts)
Input: Multiple points, aggr_function: 'lts'
Expected: Returns last value in sorted order
Scenario 3.3.3: Multiple Values with NaN
Input: Some NaN values in series Expected: NaN values dropped before aggregation
Scenario 3.3.4: All NaN Values
Input: All values are NaN
Expected: Returns None, tag skipped
Scenario 3.3.5: Invalid Aggregation Function
Input: Unknown aggr_function
Expected: Notification sent, returns 'continue', tag skipped
Scenario 3.3.6: Empty DataFrame After Filtering
Input: No data after quality gate Expected: Returns empty DataFrame dict
3.4 group_and_hold_data Activity
Scenario 3.4.1: First Run - No Existing Data
Input: No existing data in Redis for key
Expected: Creates new data_hold dict, stores in Redis
Scenario 3.4.2: Subsequent Run - Existing Data
Input: Existing data_hold in Redis
Expected: Merges new data with existing, updates timestamp
Scenario 3.4.3: Removed Tags Cleanup
Input: Tags removed from model_tags
Expected: Removed tags deleted from data_hold
Scenario 3.4.4: Empty Input Data
Input: Empty DataFrame Expected: Returns empty dict, warning logged
Scenario 3.4.5: Redis Get Error
Input: Redis get operation fails Expected: Notification sent, exception raised
Scenario 3.4.6: Redis Set Error
Input: Redis set operation fails Expected: Notification sent, exception raised
3.5 export_data_to_postgres Activity
Scenario 3.5.1: Successful Insert
Input: Valid data, no conflicts
Expected: Data inserted, affected_rows > 0
Scenario 3.5.2: Conflict with Ignore Policy
Input: Duplicate data, on_conflict: 'ignore'
Expected: Duplicates ignored, affected_rows may be less than total
Scenario 3.5.3: Conflict with Replace Policy
Input: Duplicate data, on_conflict: 'replace'
Expected: Duplicates updated, affected_rows includes updates
Scenario 3.5.4: Timestamp Conversion
Input: String timestamps in data Expected: Timestamps converted to datetime format
Scenario 3.5.5: Database Connection Error
Input: Database unavailable Expected: Exception raised, notification sent
3.6 write_metrics Activity
Scenario 3.6.1: Success with Valid Values
Input: Data with non-None values Expected: Metrics written for all non-None values
Scenario 3.6.2: Some None Values
Input: Some values are None Expected: None values skipped, only non-None values written
Scenario 3.6.3: All None Values
Input: All values are None Expected: No metrics written, activity completes
3.7 store_data_package Activity
Scenario 3.7.1: Success
Input: Valid data and held_data Expected: Package stored in Redis with TTL 120
Scenario 3.7.2: Redis Error
Input: Redis set fails Expected: Notification sent, exception raised
4. Integration Scenarios
4.1 End-to-End Scenarios
Scenario 4.1.1: Complete Happy Path
Description: Full workflow from API to database
Flow:
- PI Web API returns data
- Quality gate passes
- Aggregation succeeds
- Redis storage succeeds
- PostgreSQL export succeeds
- Metrics written
- Debug package stored (if enabled)
Assertions:
- All activities called
- Data in all storage layers
- No errors
Scenario 4.1.2: Partial Failure with Retry
Description: Activity fails, retries succeed
Flow:
- First attempt fails (e.g., Redis timeout)
- Retry policy triggers
- Second attempt succeeds
- Workflow continues
Assertions:
- Retry policy applied
- Workflow eventually succeeds
- Error logged but not fatal
Scenario 4.1.3: Complete Failure After Retries
Description: Activity fails after all retries exhausted
Flow:
- Activity fails repeatedly
- Retry policy exhausted
- Workflow fails
Assertions:
- All retries attempted
- Workflow fails with error
- Error notification sent
5. Edge Cases and Boundary Conditions
5.1 Data Edge Cases
Scenario 5.1.1: Very Large Dataset
Input: Thousands of data points Expected: Handles efficiently, all processed
Scenario 5.1.2: Single Data Point
Input: One tag, one data point Expected: Processes correctly
Scenario 5.1.3: Extreme Values
Input: Very large or very small numeric values Expected: Handled correctly, no overflow
Scenario 5.1.4: Special Characters in Tag Names
Input: Tag names with special characters Expected: Handled correctly
5.2 Configuration Edge Cases
Scenario 5.2.1: Very Short Retention Time
Input: retention_time: 1 (1 second)
Expected: Data expires quickly but workflow completes
Scenario 5.2.2: Very Long Retention Time
Input: retention_time: 86400 (1 day)
Expected: Data persists for full duration
Scenario 5.2.3: Max Count = 1
Input: max_count: 1
Expected: Only latest value retrieved
Scenario 5.2.4: Max Count = Large Number
Input: max_count: 10000
Expected: Many values retrieved and processed
5.3 Concurrent Execution Scenarios
Scenario 5.3.1: Multiple Workflows Same Schedule
Input: Two workflows with same schedule_name running concurrently
Expected: Both complete, data merged correctly in Redis
Scenario 5.3.2: Multiple Workflows Different Schedules
Input: Multiple workflows with different schedule_name
Expected: Each uses separate Redis keys, no interference
6. Performance Scenarios
6.1 Load Scenarios
Scenario 6.1.1: High Throughput
Input: Many tags, frequent execution Expected: Handles load efficiently
Scenario 6.1.2: Large Payload
Input: Large amount of data per tag Expected: Processes within timeout limits
7. Test Data Requirements
7.1 Valid Test Data Structure
{
'model_name': 'test_model',
'model_id': 'test_model_id',
'schedule_name': 'test_schedule',
'model_tags': {
'tag1': {
'webid': 'webid1',
'aggr_function': 'avg',
'data_range': [0, 100],
'frequency': 60000,
},
},
'trigger_laborious': False,
'filters': {},
'schema': 'test_schema',
'table_name': 'test_table',
'retention_time': 3600,
'fill_missing_tags': False,
'debug_data_package': False,
'pi_web_api_query': {
'endpoint': '/streamsets/recorded',
'period': '*-1d',
'max_count': 10,
'api_timeout': 30,
},
}
7.2 Mock PI Web API Response
DataFrame({
'timestamp': ['2024-01-01 12:00:00+0000', ...],
'name': ['tag1', 'tag2', ...],
'value': [10.5, 20.3, ...],
'tag': ['webid1', 'webid2', ...],
})
8. Test Implementation Notes
8.1 Test Organization
- Group tests by scenario category
- Use descriptive test names matching scenario IDs
- Share fixtures for common setup
- Use parametrized tests for similar scenarios
8.2 Assertions Checklist
For each scenario, verify:
- Correct activities called
- Correct parameters passed
- Expected data in storage (SQLite/Redis)
- Expected notifications sent
- Expected metrics written
- No unexpected errors
- Workflow state correct
8.3 Mock Configuration
- Mock PI Web API client responses
- Use fake Redis (fakeredis)
- Use fake MongoDB (mongomock)
- Use SQLite for PostgreSQL
- Mock notification handler
- Mock metrics controller
9. Priority Scenarios
High Priority (Must Test)
- Scenario 1.1.1: Happy Path
- Scenario 1.2.1: Empty Data
- Scenario 1.3.1: API Connection Error
- Scenario 2.1.1: Complete Processing
- Scenario 2.2.1: Empty After Grouping
- Scenario 2.3.4: PostgreSQL Error
Medium Priority (Should Test)
- Scenario 1.1.3: Debug Package
- Scenario 2.1.2: Quality Filters
- Scenario 2.1.3: Different Aggregations
- Scenario 3.3.5: Invalid Aggregation
- Scenario 3.5.2: Conflict Ignore
Low Priority (Nice to Have)
- Scenario 4.1.2: Retry Success
- Scenario 5.1.1: Large Dataset
- Scenario 5.3.1: Concurrent Execution