Files
sientia-dataops-scouter_tem…/e2e/scenarios.md
vitor-aignosi b17eaf972c SIENTIAPDE-1445
Update requirements-dev.txt to add E2E testing dependencies: fakeredis and mongomock for in-memory testing, and include testcontainers for PostgreSQL support.
2025-12-30 08:40:07 -03:00

20 KiB

Test Scenarios for PI Web API Scouter Workflow

This document describes all possible test scenarios for the pi_web_api_scouter workflow and its child workflow core_scouter.

Workflow Overview

The pi_web_api_scouter workflow:

  1. Retrieves tag values from PI Web API
  2. Delegates processing to core_scouter child workflow which:
    • Applies data quality gates
    • Aggregates data
    • Groups and holds data in Redis
    • Exports to PostgreSQL
    • Writes metrics
    • Optionally stores debug data package

1. PI Web API Scouter - Main Workflow Scenarios

1.1 Success Scenarios

Scenario 1.1.1: Happy Path - Complete Success

Description: Workflow completes successfully with valid data from PI Web API

Input:

  • Valid model_name, model_id, schedule_name
  • Valid pi_web_api_query with endpoint, period, max_count, api_timeout
  • Valid model_tags with webids and configurations
  • Valid filters, schema, table_name, retention_time

Expected Behavior:

  • get_tag_values returns non-empty list of records
  • Workflow proceeds to core_scouter
  • All activities execute successfully
  • Data is stored in PostgreSQL
  • Metrics are written
  • Workflow completes without errors

Assertions:

  • PI Web API client called once with correct parameters
  • Data exists in SQLite (PostgreSQL substitute)
  • Data cached in Redis
  • Metrics written
  • No errors raised

Scenario 1.1.2: Success with Multiple Tags

Description: Workflow processes multiple tags successfully

Input:

  • Multiple tags in model_tags (3+ tags)
  • Each tag has valid webid, aggr_function, data_range, frequency

Expected Behavior:

  • All tags retrieved from PI Web API
  • All tags processed through quality gates
  • All tags aggregated correctly
  • All tags stored in database

Assertions:

  • Number of records matches number of tags
  • All tags present in final data
  • Aggregation applied per tag configuration

Scenario 1.1.3: Success with Debug Data Package Enabled

Description: Workflow completes with debug_data_package=True

Input:

  • All standard input
  • debug_data_package: True

Expected Behavior:

  • Normal workflow execution
  • store_data_package activity called
  • Data package stored in Redis

Assertions:

  • store_data_package called once
  • Data package key exists in Redis
  • Package contains both data and held_data

1.2 Early Exit Scenarios

Scenario 1.2.1: Empty Data from PI Web API

Description: PI Web API returns empty data

Input:

  • Valid configuration
  • PI Web API returns empty DataFrame or empty list

Expected Behavior:

  • get_tag_values returns empty list []
  • Workflow checks if not data: and returns early
  • core_scouter is NOT called
  • Workflow completes without error

Assertions:

  • PI Web API called once
  • core_scouter NOT called
  • No data in PostgreSQL
  • No data in Redis (except possibly from previous runs)

Scenario 1.2.2: None Returned from PI Web API

Description: PI Web API returns None

Input:

  • Valid configuration
  • PI Web API returns None

Expected Behavior:

  • get_tag_values returns None
  • Workflow checks if not data: and returns early
  • core_scouter is NOT called

Assertions:

  • PI Web API called once
  • core_scouter NOT called
  • Workflow completes without error

1.3 Error Scenarios

Scenario 1.3.1: PI Web API Connection Error

Description: PI Web API client raises connection error

Input:

  • Valid configuration
  • PI Web API client raises PIMSRequestError or connection exception

Expected Behavior:

  • get_tag_values catches exception
  • Sends notification with PI_WEB_API_REQUEST_ERROR
  • Raises exception (workflow fails after retries)

Assertions:

  • Notification sent with correct error details
  • Exception propagated to workflow
  • Workflow fails (after retry policy exhausted)
  • core_scouter NOT called

Scenario 1.3.2: PI Web API Timeout

Description: PI Web API request times out

Input:

  • Valid configuration
  • api_timeout set to low value
  • PI Web API takes longer than timeout

Expected Behavior:

  • Request times out
  • Exception raised
  • Notification sent
  • Workflow fails after retries

Assertions:

  • Timeout exception caught
  • Notification sent
  • Workflow fails

Scenario 1.3.3: Invalid Endpoint

Description: Invalid PI Web API endpoint provided

Input:

  • Invalid endpoint path in pi_web_api_query

Expected Behavior:

  • PI Web API client raises error
  • Notification sent
  • Workflow fails

Assertions:

  • Error notification sent
  • Workflow fails

Scenario 1.3.4: Missing Required Input Fields

Description: Missing required input fields

Input:

  • Missing model_id, model_name, schedule_name, or pi_web_api_query

Expected Behavior:

  • KeyError raised when accessing missing fields
  • Workflow fails immediately

Assertions:

  • KeyError or similar exception
  • Workflow fails before any activity execution

2. CoreScouter - Child Workflow Scenarios

2.1 Success Scenarios

Scenario 2.1.1: Complete Processing Success

Description: All stages complete successfully

Input:

  • Valid data from parent workflow
  • Valid filters, model_tags, schema, table_name
  • fill_missing_tags: False
  • debug_data_package: False

Expected Behavior:

  • data_quality_gate filters data
  • aggregate_data aggregates by tag
  • group_and_hold_data stores in Redis
  • export_data_to_postgres writes to database
  • write_metrics records metrics
  • Workflow completes

Assertions:

  • All activities called in correct order
  • Data in PostgreSQL
  • Data in Redis
  • Metrics written
  • store_data_package NOT called

Scenario 2.1.2: Success with Data Quality Filters

Description: Data quality filters applied successfully

Input:

  • Data with some quality issues
  • Filters configured with NULL_VALUES_FILTER or OUT_OF_BOUNDS_FILTER
  • Policy set to DISCARD or WARN

Expected Behavior:

  • Quality gate identifies issues
  • Notification sent (WARNING level)
  • If policy is DISCARD, bad rows removed
  • Remaining data processed normally

Assertions:

  • Quality issues detected
  • Notification sent
  • Bad data discarded if policy is DISCARD
  • Good data processed

Scenario 2.1.3: Success with Different Aggregation Functions

Description: Different aggregation functions applied correctly

Input:

  • Multiple tags with different aggr_function: avg, mdn, max, min, lts
  • Time-series data with multiple points per tag

Expected Behavior:

  • Each tag aggregated with its configured function
  • Aggregated values correct for each function type

Assertions:

  • avg calculates mean correctly
  • mdn calculates median correctly
  • max returns maximum value
  • min returns minimum value
  • lts returns latest value

Scenario 2.1.4: Success with Fill Missing Tags

Description: Missing tags filled with None

Input:

  • fill_missing_tags: True
  • Some tags missing from data

Expected Behavior:

  • Missing tags added to data_hold with value None
  • All expected tags present in final data

Assertions:

  • Missing tags present with None value
  • All model_tags represented in output

2.2 Early Exit Scenarios

Scenario 2.2.1: Empty Data After Grouping

Description: group_and_hold_data returns empty dict

Input:

  • Data that results in empty held_data after grouping

Expected Behavior:

  • group_and_hold_data returns {}
  • Workflow checks if held_data == {}: and returns early
  • export_data_to_postgres NOT called
  • write_metrics NOT called
  • store_data_package NOT called

Assertions:

  • Early return after grouping
  • No database export
  • No metrics written
  • Workflow completes without error

Scenario 2.2.2: Zero Affected Rows After Export

Description: PostgreSQL export returns zero affected rows

Input:

  • Data that results in affected_rows: 0 from export

Expected Behavior:

  • export_data_to_postgres returns {'affected_rows': 0}
  • Workflow checks if data_exported.get('affected_rows', 0) <= 0: and returns early
  • write_metrics NOT called
  • store_data_package NOT called

Assertions:

  • Early return after export
  • No metrics written
  • Workflow completes without error

2.3 Error Scenarios

Scenario 2.3.1: Data Quality Gate Error

Description: Error during quality gate processing

Input:

  • Invalid filter configuration
  • Filter function raises exception

Expected Behavior:

  • Exception caught in quality gate
  • Notification sent with DATA_QUALITY_GATE_ISSUES
  • Exception propagated (workflow fails after retries)

Assertions:

  • Error notification sent
  • Workflow fails

Scenario 2.3.2: Aggregation Error

Description: Error during data aggregation

Input:

  • Invalid aggregation function
  • Data format issues

Expected Behavior:

  • Invalid function sends notification with AGGREGATION_ISSUES
  • Returns 'continue' for invalid function (skips that tag)
  • Other errors raise exception

Assertions:

  • Invalid function handled gracefully
  • Other errors cause workflow failure

Scenario 2.3.3: Redis Connection Error

Description: Redis unavailable during group_and_hold_data

Input:

  • Valid data
  • Redis connection fails

Expected Behavior:

  • redis_repository.get() or redis_repository.set() raises exception
  • Notification sent with REDIS_GET_ERROR or REDIS_SET_ERROR
  • Exception propagated (workflow fails after retries)

Assertions:

  • Error notification sent
  • Workflow fails

Scenario 2.3.4: PostgreSQL Connection Error

Description: PostgreSQL unavailable during export

Input:

  • Valid data
  • PostgreSQL connection fails

Expected Behavior:

  • export_data_to_postgres raises exception
  • Notification sent with ERROR_EXPORTING_DATA_TO_POSTGRES
  • Exception propagated (workflow fails after retries)

Assertions:

  • Error notification sent
  • Workflow fails

Scenario 2.3.5: PostgreSQL Unique Constraint Violation

Description: Duplicate data violates unique constraint

Input:

  • Data with duplicate model_id, timestamp, variable combination
  • on_conflict: 'ignore' configured

Expected Behavior:

  • PostgreSQL handles conflict with ON CONFLICT DO NOTHING
  • affected_rows may be 0 for duplicates
  • Workflow continues normally

Assertions:

  • No exception raised
  • Duplicates ignored
  • Workflow continues

3. Activity-Specific Scenarios

3.1 get_tag_values Activity

Scenario 3.1.1: Success with Valid WebIds

Input: All webids valid and present Expected: Returns list of records with timestamp, name, value, tag

Scenario 3.1.2: Some WebIds are None

Input: Some webids in model_tags are None Expected: None webids filtered out, only valid webids queried

Scenario 3.1.3: DataFrame with NaN Values

Input: PI Web API returns DataFrame with NaN values Expected: NaN values handled, data normalized correctly

Scenario 3.1.4: Timestamp Normalization

Input: Multiple timestamps in response Expected: All timestamps normalized to max timestamp value


3.2 data_quality_gate Activity

Scenario 3.2.1: No Filters Configured

Input: Empty filters: {} Expected: Data passes through unchanged, filtered by model_tags only

Scenario 3.2.2: NULL_VALUES_FILTER with DISCARD Policy

Input: Data with null values, policy DISCARD Expected: Null rows removed, notification sent

Scenario 3.2.3: OUT_OF_BOUNDS_FILTER with WARN Policy

Input: Data outside range, policy WARN Expected: Notification sent, data kept

Scenario 3.2.4: Unknown Filter Type

Input: Filter name not in quality_gate_filters Expected: Warning logged, filter skipped, processing continues

Scenario 3.2.5: Filter Removes All Data

Input: Filter that removes all rows Expected: Empty DataFrame returned, processing continues


3.3 aggregate_data Activity

Scenario 3.3.1: Single Value Per Tag

Input: One data point per tag Expected: Fast path returns value directly

Scenario 3.3.2: Multiple Values - Latest (lts)

Input: Multiple points, aggr_function: 'lts' Expected: Returns last value in sorted order

Scenario 3.3.3: Multiple Values with NaN

Input: Some NaN values in series Expected: NaN values dropped before aggregation

Scenario 3.3.4: All NaN Values

Input: All values are NaN Expected: Returns None, tag skipped

Scenario 3.3.5: Invalid Aggregation Function

Input: Unknown aggr_function Expected: Notification sent, returns 'continue', tag skipped

Scenario 3.3.6: Empty DataFrame After Filtering

Input: No data after quality gate Expected: Returns empty DataFrame dict


3.4 group_and_hold_data Activity

Scenario 3.4.1: First Run - No Existing Data

Input: No existing data in Redis for key Expected: Creates new data_hold dict, stores in Redis

Scenario 3.4.2: Subsequent Run - Existing Data

Input: Existing data_hold in Redis Expected: Merges new data with existing, updates timestamp

Scenario 3.4.3: Removed Tags Cleanup

Input: Tags removed from model_tags Expected: Removed tags deleted from data_hold

Scenario 3.4.4: Empty Input Data

Input: Empty DataFrame Expected: Returns empty dict, warning logged

Scenario 3.4.5: Redis Get Error

Input: Redis get operation fails Expected: Notification sent, exception raised

Scenario 3.4.6: Redis Set Error

Input: Redis set operation fails Expected: Notification sent, exception raised


3.5 export_data_to_postgres Activity

Scenario 3.5.1: Successful Insert

Input: Valid data, no conflicts Expected: Data inserted, affected_rows > 0

Scenario 3.5.2: Conflict with Ignore Policy

Input: Duplicate data, on_conflict: 'ignore' Expected: Duplicates ignored, affected_rows may be less than total

Scenario 3.5.3: Conflict with Replace Policy

Input: Duplicate data, on_conflict: 'replace' Expected: Duplicates updated, affected_rows includes updates

Scenario 3.5.4: Timestamp Conversion

Input: String timestamps in data Expected: Timestamps converted to datetime format

Scenario 3.5.5: Database Connection Error

Input: Database unavailable Expected: Exception raised, notification sent


3.6 write_metrics Activity

Scenario 3.6.1: Success with Valid Values

Input: Data with non-None values Expected: Metrics written for all non-None values

Scenario 3.6.2: Some None Values

Input: Some values are None Expected: None values skipped, only non-None values written

Scenario 3.6.3: All None Values

Input: All values are None Expected: No metrics written, activity completes


3.7 store_data_package Activity

Scenario 3.7.1: Success

Input: Valid data and held_data Expected: Package stored in Redis with TTL 120

Scenario 3.7.2: Redis Error

Input: Redis set fails Expected: Notification sent, exception raised


4. Integration Scenarios

4.1 End-to-End Scenarios

Scenario 4.1.1: Complete Happy Path

Description: Full workflow from API to database

Flow:

  1. PI Web API returns data
  2. Quality gate passes
  3. Aggregation succeeds
  4. Redis storage succeeds
  5. PostgreSQL export succeeds
  6. Metrics written
  7. Debug package stored (if enabled)

Assertions:

  • All activities called
  • Data in all storage layers
  • No errors

Scenario 4.1.2: Partial Failure with Retry

Description: Activity fails, retries succeed

Flow:

  1. First attempt fails (e.g., Redis timeout)
  2. Retry policy triggers
  3. Second attempt succeeds
  4. Workflow continues

Assertions:

  • Retry policy applied
  • Workflow eventually succeeds
  • Error logged but not fatal

Scenario 4.1.3: Complete Failure After Retries

Description: Activity fails after all retries exhausted

Flow:

  1. Activity fails repeatedly
  2. Retry policy exhausted
  3. Workflow fails

Assertions:

  • All retries attempted
  • Workflow fails with error
  • Error notification sent

5. Edge Cases and Boundary Conditions

5.1 Data Edge Cases

Scenario 5.1.1: Very Large Dataset

Input: Thousands of data points Expected: Handles efficiently, all processed

Scenario 5.1.2: Single Data Point

Input: One tag, one data point Expected: Processes correctly

Scenario 5.1.3: Extreme Values

Input: Very large or very small numeric values Expected: Handled correctly, no overflow

Scenario 5.1.4: Special Characters in Tag Names

Input: Tag names with special characters Expected: Handled correctly


5.2 Configuration Edge Cases

Scenario 5.2.1: Very Short Retention Time

Input: retention_time: 1 (1 second) Expected: Data expires quickly but workflow completes

Scenario 5.2.2: Very Long Retention Time

Input: retention_time: 86400 (1 day) Expected: Data persists for full duration

Scenario 5.2.3: Max Count = 1

Input: max_count: 1 Expected: Only latest value retrieved

Scenario 5.2.4: Max Count = Large Number

Input: max_count: 10000 Expected: Many values retrieved and processed


5.3 Concurrent Execution Scenarios

Scenario 5.3.1: Multiple Workflows Same Schedule

Input: Two workflows with same schedule_name running concurrently Expected: Both complete, data merged correctly in Redis

Scenario 5.3.2: Multiple Workflows Different Schedules

Input: Multiple workflows with different schedule_name Expected: Each uses separate Redis keys, no interference


6. Performance Scenarios

6.1 Load Scenarios

Scenario 6.1.1: High Throughput

Input: Many tags, frequent execution Expected: Handles load efficiently

Scenario 6.1.2: Large Payload

Input: Large amount of data per tag Expected: Processes within timeout limits


7. Test Data Requirements

7.1 Valid Test Data Structure

{
    'model_name': 'test_model',
    'model_id': 'test_model_id',
    'schedule_name': 'test_schedule',
    'model_tags': {
        'tag1': {
            'webid': 'webid1',
            'aggr_function': 'avg',
            'data_range': [0, 100],
            'frequency': 60000,
        },
    },
    'trigger_laborious': False,
    'filters': {},
    'schema': 'test_schema',
    'table_name': 'test_table',
    'retention_time': 3600,
    'fill_missing_tags': False,
    'debug_data_package': False,
    'pi_web_api_query': {
        'endpoint': '/streamsets/recorded',
        'period': '*-1d',
        'max_count': 10,
        'api_timeout': 30,
    },
}

7.2 Mock PI Web API Response

DataFrame({
    'timestamp': ['2024-01-01 12:00:00+0000', ...],
    'name': ['tag1', 'tag2', ...],
    'value': [10.5, 20.3, ...],
    'tag': ['webid1', 'webid2', ...],
})

8. Test Implementation Notes

8.1 Test Organization

  • Group tests by scenario category
  • Use descriptive test names matching scenario IDs
  • Share fixtures for common setup
  • Use parametrized tests for similar scenarios

8.2 Assertions Checklist

For each scenario, verify:

  • Correct activities called
  • Correct parameters passed
  • Expected data in storage (SQLite/Redis)
  • Expected notifications sent
  • Expected metrics written
  • No unexpected errors
  • Workflow state correct

8.3 Mock Configuration

  • Mock PI Web API client responses
  • Use fake Redis (fakeredis)
  • Use fake MongoDB (mongomock)
  • Use SQLite for PostgreSQL
  • Mock notification handler
  • Mock metrics controller

9. Priority Scenarios

High Priority (Must Test)

  1. Scenario 1.1.1: Happy Path
  2. Scenario 1.2.1: Empty Data
  3. Scenario 1.3.1: API Connection Error
  4. Scenario 2.1.1: Complete Processing
  5. Scenario 2.2.1: Empty After Grouping
  6. Scenario 2.3.4: PostgreSQL Error

Medium Priority (Should Test)

  1. Scenario 1.1.3: Debug Package
  2. Scenario 2.1.2: Quality Filters
  3. Scenario 2.1.3: Different Aggregations
  4. Scenario 3.3.5: Invalid Aggregation
  5. Scenario 3.5.2: Conflict Ignore

Low Priority (Nice to Have)

  1. Scenario 4.1.2: Retry Success
  2. Scenario 5.1.1: Large Dataset
  3. Scenario 5.3.1: Concurrent Execution