24 Commits

Author SHA1 Message Date
Eduardo Rios
695e4c07a6 SIENTIAPDE-2072: bump sientia_do pin to 1.12.2
Picks up the notification timestamp -> native datetime fix.
2026-08-17 16:42:31 -03:00
Bruno Domingues
6a8c41b328 fix(pytest): Workaround unraisableexception plugin crash
Disables the unraisableexception plugin in pytest due to a known bug in
pytest>=9.1 where it crashes with tracemalloc errors when multiple
unraisable exceptions occur close together.
2026-08-04 15:21:14 -03:00
Bruno Domingues
773c980fc3 fix(storage): Prevent AttributeError when closing uninitialized Postgres engine 2026-08-04 14:57:52 -03:00
Bruno Domingues
35efb67c89 fix(test_activities): Use aclose for async mock shutdown 2026-08-04 14:43:15 -03:00
Bruno Domingues
acd925cf2a refactor(opc): Rename async close method to aclose
Renamed the OPC.close asynchronous method to OPC.aclose to align with common Python conventions for asynchronous context managers and methods, improving clarity. All call sites and tests have been updated accordingly.
2026-08-04 12:01:06 -03:00
Bruno Domingues
310ceea0d8 ci(quality-gate): Configure push triggers, concurrency, and granular permissions 2026-08-04 11:41:23 -03:00
Bruno Domingues
2d73ec9ec2 Merge pull request #42 from Aignosi/feature/SIENTIAPDE-1945
SIENTIAPDE-1945: Remove values.yaml from sientia-module Helm chart
2026-07-10 11:49:36 -03:00
Bruno Domingues
2142143ab9 SIENTIAPDE-1945: Delete values.yaml configuration file for sientia-module Helm chart. 2026-07-08 22:12:50 -03:00
Bruno Domingues
c07457bfbe chore(sonar): update project key 2026-07-01 20:59:29 -03:00
vitor-aignosi
086b12492e Merge pull request #41 from Aignosi/feature/SIENTIAPDE-1646-legacy-laborious-worker
SIENTIAPDE-1646: Refactor Worker Task Queue Management and Update Dependencies
2026-05-21 15:30:09 -03:00
vitor-aignosi
4ea0754f0c SIENTIAPDE-1646
Enhance MLFlow run ID resolution with error handling for missing and invalid source URIs

- Added checks in `get_model_run_id` method to raise exceptions for models with missing or invalid source URIs.
- Introduced new test cases to validate error handling for these scenarios.
- Updated `requirements-light.txt` to include `mlflow` as a dependency.
2026-05-20 09:53:53 -03:00
vitor-aignosi
ddb1618209 SIENTIAPDE-1646
Update quality-gate workflow to use python-quality-gate template
2026-05-20 09:30:05 -03:00
vitor-aignosi
90f8bdda61 SIENTIAPDE-1646
Update requirements.txt to align with recent dependency changes and ensure compatibility across the project.
2026-05-20 09:26:35 -03:00
vitor-aignosi
2ccda3e440 SIENTIAPDE-1646
Refactor worker task queue management and update README

- Introduced runtime-scoped task queues for workflows, replacing legacy queue names.
- Updated worker implementation to utilize `sientia_do.temporal.worker.prepare_worker`.
- Added `RUNTIME` environment variable to configure task queue suffixes.
- Enhanced README documentation to reflect changes in task queue structure and worker setup.
2026-05-19 17:07:18 -03:00
vitor-aignosi
d856150e24 Update requirements.txt 2026-05-19 14:45:26 -03:00
vitor-aignosi
6569810756 Update requirements.txt 2026-05-19 14:41:58 -03:00
vitor-aignosi
7153f1da0d SIENTIAPDE-1646
Remove requirements-light.txt and update requirements.txt to specify versions for asyncua and new sientia dependencies.
2026-05-19 14:13:23 -03:00
vitor-aignosi
fcc8920a8b Merge pull request #40 from Aignosi/fix/SIENTIAPDE-1811-fix
SIENTIAPDE-1811: Update dependencies, gitignore, and strengthen OPC UA error handling
2026-05-19 14:01:37 -03:00
vitor-aignosi
cd1be2430a SIENTIAPDE-1811
Update .gitignore, requirements, and enhance OPC UA error handling

- Added new entries to .gitignore for openspec and cursor directories.
- Updated sientia-dataops-library dependency version in requirements-light.txt from 1.10.4 to 1.12.0.
- Enhanced OPC UA communication by refining reconnect logic and error handling in opc_repository.py, including the introduction of a reconnect flag and improved session management.
- Updated tests to cover new reconnect scenarios and ensure robust error handling for protocol states.
2026-05-19 11:45:08 -03:00
vitor-aignosi
7c1dae8ef6 Merge pull request #39 from Aignosi/fix/SIENTIAPDE-1811
Enhance OPC UA Communication and Metrics Tracking
2026-05-18 13:32:33 -03:00
vitor-aignosi
2a6def4056 SIENTIAPDE-1811
SIENTIAPDE-1811 Implement OPC write error handling and refactor tag writing logic

- Introduced a new function `_apply_opc_write_error` to manage session and reconnect flags based on OPC write error responses.
- Refactored the `_write_tags_from_config` method to streamline the writing of OPC tags for both prediction and confidence data.
- Enhanced unit tests to cover various scenarios for OPC write errors, including session bad and reconnect in progress cases.
- Updated existing tests to validate the new logic and ensure robust error handling.
2026-05-18 10:36:57 -03:00
vitor-aignosi
7d59fd7c8c SIENTIAPDE-1811
Update values.yaml to rename worker references and adjust GitHub branch for SIENTIAPDE-1811. Changed nameOverride, fullnameOverride, and service account name to "sientia-laborious-legacy-worker" and updated the GITHUB_BRANCH value to "fix/SIENTIAPDE-1811".
2026-05-18 10:17:52 -03:00
vitor-aignosi
e3636d4b88 SIENTIAPDE-1811
Enhance OPC UA testing framework and documentation

- Added a new marker in pyproject.toml for tests using the in-process OPC UA server.
- Updated opc-communication.md to clarify E2E test scenarios involving the real OPC server and mock server.
- Introduced an in-process asyncua OPC UA server fixture in conftest.py for E2E tests.
- Created a new fixture for activities using the real OpcRepository connected to the in-process server.
- Updated scenarios.md to include instructions for running OPC real-server tests.
2026-05-15 15:46:39 -03:00
vitor-aignosi
638d5b70b4 SIENTIAPDE-1811
Enhance OPC UA communication and metrics tracking

- Updated README.md to include new OPC UA Communication section and detailed metrics for session and write diagnostics.
- Added new metrics in laborious/metrics.py for tracking OPC UA session states and write attempts.
- Refactored OPC activity in laborious/activities/opc.py to handle session errors and improve error reporting.
- Updated e2e tests to cover new scenarios for OPC session/channel errors and reconnect handling.
- Modified .gitignore to include relatorio files and mlruns directory.
- Added ipykernel to requirements-dev.txt for Jupyter notebook support.
2026-05-15 15:28:28 -03:00
31 changed files with 2496 additions and 934 deletions

View File

@@ -1,15 +1,29 @@
name: Quality gate name: Quality gate
on: on:
push:
branches:
- main
- 'release/**'
- 'feature/**'
pull_request: pull_request:
branches: branches:
- main - main
- 'release/**'
- 'feature/**'
types: [ opened, synchronize, reopened ] types: [ opened, synchronize, reopened ]
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
jobs: jobs:
quality-gate: quality-gate:
uses: Aignosi/github_workflow_templates/.github/workflows/dataops-module-quality-gate.yml@main uses: Aignosi/github_workflow_templates/.github/workflows/python-quality-gate.yml@main
permissions: write-all permissions:
contents: read
pull-requests: write
issues: write
with: with:
project_name: 'laborious' project_name: 'laborious'
repositories: 'sientia-dataops-library, sientia-mlops-library' repositories: 'sientia-dataops-library, sientia-mlops-library'

4
.gitignore vendored
View File

@@ -52,3 +52,7 @@ catboost_info/
.ruff_cache/ .ruff_cache/
.mypy_cache/ .mypy_cache/
mlruns/ mlruns/
relatorio*
openspec/
.cursor/

View File

@@ -48,6 +48,7 @@ A comprehensive, Temporal-based ML orchestration system for industrial data proc
- [Prediction Operation Metrics](#prediction-operation-metrics) - [Prediction Operation Metrics](#prediction-operation-metrics)
- [OPC Export Metrics](#opc-export-metrics) - [OPC Export Metrics](#opc-export-metrics)
- [Data Quality Metrics](#data-quality-metrics) - [Data Quality Metrics](#data-quality-metrics)
- [OPC UA Communication](#opc-ua-communication)
- [Configuration](#configuration-1) - [Configuration](#configuration-1)
- [Environment Variables](#environment-variables) - [Environment Variables](#environment-variables)
- [OPC Configuration](#opc-configuration) - [OPC Configuration](#opc-configuration)
@@ -125,10 +126,15 @@ Laborious uses a Temporal-based architecture with strong separation of concerns
### Key Components ### Key Components
#### **Worker (`laborious/worker/worker.py`)** #### **Worker (`laborious/worker/worker.py`)**
- Temporal client setup, worker lifecycle, task queues - Temporal client setup, four workers via `sientia_do.temporal.worker.prepare_worker`
- Runtime-scoped task queues: `{workflow}-{RUNTIME}-queue` for all workflows
- Metrics server initialization, notification handler setup - Metrics server initialization, notification handler setup
- Graceful shutdown and autoscaling-friendly behavior - Graceful shutdown and autoscaling-friendly behavior
**Breaking (schedulers):** drift and simple_metrics queues are no longer `drift-queue` /
`simple_metrics-queue`. Use `drift-{RUNTIME}-queue` and `simple_metrics-{RUNTIME}-queue`
matching the worker pod `RUNTIME` env (same as `predictions_batch` / `minimal_retrain`).
#### **Workflows (`laborious/workflows/`)** #### **Workflows (`laborious/workflows/`)**
- `predictions_batch.py`: Batch prediction entry point - `predictions_batch.py`: Batch prediction entry point
- `sub_workflows/prediction_process.py`: Core prediction pipeline - `sub_workflows/prediction_process.py`: Core prediction pipeline
@@ -164,7 +170,7 @@ Laborious uses a Temporal-based architecture with strong separation of concerns
- `connectors_config.py`: Env-driven configuration builders - `connectors_config.py`: Env-driven configuration builders
- `models/minio_dataframe_payload.py`: MinIO-offloaded DataFrame payload model - `models/minio_dataframe_payload.py`: MinIO-offloaded DataFrame payload model
- `repository/model_repository.py`: MLFlow operations and retraining - `repository/model_repository.py`: MLFlow operations and retraining
- `repository/opc_repository.py`: OPC communication and writes - `repository/opc_repository.py`: OPC UA client, writes, session recovery (see [OPC UA Communication](#opc-ua-communication))
- `repository/minio_manager.py`: MinIO object storage operations - `repository/minio_manager.py`: MinIO object storage operations
- `filters/conditional_filters.py` and `filters/mlflow_filters.py` - `filters/conditional_filters.py` and `filters/mlflow_filters.py`
@@ -779,6 +785,15 @@ The Laborious system exposes comprehensive Prometheus metrics for operational vi
- Labels: `pod_id`, `model_name`, `workflow_name`, `opc_server_id` - Labels: `pod_id`, `model_name`, `workflow_name`, `opc_server_id`
- Buckets: [0.01, 0.05, 0.1, 0.2, 0.5, 1.0, 2.0, 5.0, 10.0] - Buckets: [0.01, 0.05, 0.1, 0.2, 0.5, 1.0, 2.0, 5.0, 10.0]
OPC UA session and write diagnostics (Prometheus, `laborious/metrics.py`):
- `opc_connections_initiated_total`, `opc_connections_failed_total`, `opc_connection_status`
- `opc_session_created_total`, `opc_session_closed_total`, `opc_session_revised_timeout_milliseconds`
- `opc_write_attempts_total` (label `result`: `OK` or exception name, e.g. `BadSessionIdInvalid`)
- `opc_write_inter_arrival_over_session_timeout_total`
See [OPC UA Communication](#opc-ua-communication) for semantics, concurrency, and confidence codes **12** / **14**.
### Data Quality Metrics ### Data Quality Metrics
- Filter pass/fail rates through notification system - Filter pass/fail rates through notification system
- MLFlow API response validation metrics - MLFlow API response validation metrics
@@ -792,6 +807,7 @@ The Laborious system exposes comprehensive Prometheus metrics for operational vi
|----------|-------------|---------|----------| |----------|-------------|---------|----------|
| `TEMPORAL_HOST` | Temporal server address | `localhost:7233` | Yes | | `TEMPORAL_HOST` | Temporal server address | `localhost:7233` | Yes |
| `TEMPORAL_NAMESPACE` | Temporal namespace | `laborious` | No | | `TEMPORAL_NAMESPACE` | Temporal namespace | `laborious` | No |
| `RUNTIME` | Task queue suffix for all workflows (`{workflow}-{RUNTIME}-queue`) | _(none)_ | Yes |
| `POSTGRES_HOST` | PostgreSQL hostname | `localhost` | Yes | | `POSTGRES_HOST` | PostgreSQL hostname | `localhost` | Yes |
| `POSTGRES_PORT` | PostgreSQL port | `5432` | Yes | | `POSTGRES_PORT` | PostgreSQL port | `5432` | Yes |
| `POSTGRES_USER` | PostgreSQL username | `sientia` | Yes | | `POSTGRES_USER` | PostgreSQL username | `sientia` | Yes |
@@ -810,7 +826,7 @@ The Laborious system exposes comprehensive Prometheus metrics for operational vi
| `OPC_CERT_PATH` | OPC client certificate path | `None` | No | | `OPC_CERT_PATH` | OPC client certificate path | `None` | No |
| `OPC_PRIVATE_KEY_PATH` | OPC private key path | `None` | No | | `OPC_PRIVATE_KEY_PATH` | OPC private key path | `None` | No |
| `OPC_SERVER_CERT_PATH` | OPC server certificate path | `None` | No | | `OPC_SERVER_CERT_PATH` | OPC server certificate path | `None` | No |
| `OPC_RECONNECTION_INTERVAL` | OPC reconnection interval (ms) | `120` | No | | `OPC_RECONNECTION_INTERVAL` | Minimum seconds between OPC reconnects | `120` | No |
| `PI_WEB_API_BASE_URL` | PI Web API server base URL | `None` | No | | `PI_WEB_API_BASE_URL` | PI Web API server base URL | `None` | No |
| `PI_WEB_API_AUTH_TYPE` | PI Web API authentication type (basic/bearer) | `None` | No | | `PI_WEB_API_AUTH_TYPE` | PI Web API authentication type (basic/bearer) | `None` | No |
| `PI_WEB_API_AUTH_TOKEN` | PI Web API authentication token | `None` | No | | `PI_WEB_API_AUTH_TOKEN` | PI Web API authentication token | `None` | No |
@@ -903,6 +919,12 @@ Legacy MinIO object layout (relative key):
`training_datasets/{model_name}/{object_prefix}_{timestamp}.parquet` where `object_prefix` is sanitized `training_datasets/{model_name}/{object_prefix}_{timestamp}.parquet` where `object_prefix` is sanitized
(slashes replaced by underscores) to keep a stable model-level directory. (slashes replaced by underscores) to keep a stable model-level directory.
## OPC UA Communication
Full reference: **[docs/opc-communication.md](docs/opc-communication.md)** (connection lifecycle, Tier-1 `Bad*` reconnect, connection lock / session readiness, metrics, PostgreSQL confidence **12** vs **14**, tests).
Implementation plan: [`.cursor/plans/opc_bad_reconnect_ac4c6045.plan.md`](.cursor/plans/opc_bad_reconnect_ac4c6045.plan.md).
### OPC Configuration ### OPC Configuration
For multiple OPC servers, use the `OPC_CONFIG` environment variable: For multiple OPC servers, use the `OPC_CONFIG` environment variable:
@@ -1081,8 +1103,7 @@ laborious/
│ ├── prediction_process.py # Core prediction workflow │ ├── prediction_process.py # Core prediction workflow
│ └── format_and_export_prediction.py # Export workflow │ └── format_and_export_prediction.py # Export workflow
├── worker/ # Worker implementation ├── worker/ # Worker implementation
── worker.py # Main worker orchestrator ── worker.py # Main worker orchestrator (uses sientia_do prepare_worker)
│ └── prepare_worker.py # Worker factory with autoscaling config
├── utils/ # Utility functions ├── utils/ # Utility functions
│ ├── connectors_config.py # Environment-driven config builders │ ├── connectors_config.py # Environment-driven config builders
│ ├── models/ # Data models │ ├── models/ # Data models
@@ -1126,8 +1147,10 @@ laborious/
- Ensure proper connection pool configuration - Ensure proper connection pool configuration
4. **OPC Connection Failures** 4. **OPC Connection Failures**
- Verify OPC server is accessible - See [docs/opc-communication.md](docs/opc-communication.md)
- Verify OPC server is accessible and `OPC_RECONNECTION_INTERVAL` is appropriate
- Check certificate and key file paths - Check certificate and key file paths
- Correlate `opc_write_attempts_total` with `opc_session_*` metrics; count session errors via `prediction_confidence = 14`
- Review OPC server logs for connection issues - Review OPC server logs for connection issues
5. **PI Web API Connection Failures** 5. **PI Web API Connection Failures**
@@ -1161,7 +1184,9 @@ export LOG_LEVEL=DEBUG
### Scaling Considerations ### Scaling Considerations
- **Horizontal Scaling**: Deploy multiple worker instances - **Horizontal Scaling**: Deploy multiple worker instances
- **Task Queue Distribution**: Use multiple task queues for different workflow types - **Task Queue Distribution**: One worker pod per `RUNTIME`; queues are
`predictions_batch-{RUNTIME}-queue`, `minimal_retrain-{RUNTIME}-queue`,
`drift-{RUNTIME}-queue`, `simple_metrics-{RUNTIME}-queue`
- **Database Performance**: Optimize indexes and connection pooling - **Database Performance**: Optimize indexes and connection pooling
- **MLFlow Performance**: Configure appropriate model serving resources - **MLFlow Performance**: Configure appropriate model serving resources

170
docs/opc-communication.md Normal file
View File

@@ -0,0 +1,170 @@
# OPC UA communication (Laborious)
Laborious exports predictions to OPC UA servers through `OpcRepository` ([`laborious/utils/repository/opc_repository.py`](../laborious/utils/repository/opc_repository.py)) and the Temporal activity layer in [`laborious/activities/opc.py`](../laborious/activities/opc.py).
Implementation plan for session/channel recovery on Tier-1 `Bad*` errors: [`.cursor/plans/opc_bad_reconnect_ac4c6045.plan.md`](../.cursor/plans/opc_bad_reconnect_ac4c6045.plan.md).
## Architecture
```text
Worker (long-lived)
└── OpcRepository per OPC server id (from OPC_CONFIG / env)
├── connect / disconnect / validate_connection (read-only)
├── _connect_locked / _reconnect_locked (under _connection_lock)
├── write_data (single attempt per call)
└── background reconnect on Tier-1 Bad*, protocol closed, or stale session
Temporal activity write_opc_data
└── OPC.manage_output_tags → write_data per tag (sequential per activity)
```
One worker process holds one `OpcRepository` instance per configured server. Multiple Temporal activities can call `write_data` concurrently on the same repository.
## Connection lifecycle
| Phase | Behavior |
|-------|----------|
| Startup | `init_opc()` creates repositories and calls `connect()``_connect_locked()` |
| Steady state | `validate_connection()` is read-only (`protocol.state` only); `_session_ready` is checked in `write_data` |
| Tier-1 Bad* / protocol closed | `_start_reconnect``_run_reconnect``_reconnect_locked()` (respects `reconnection_interval`) |
| Write | `write_data()` checks reconnect task, `_session_ready`, validates protocol, then one `get_node` + `write_value` |
| Shutdown | `close()` disconnects all repositories |
### Session and channel timeouts
Requested session and secure-channel lifetime: **10 minutes** (`OPC_UA_SESSION_AND_CHANNEL_TIMEOUT_MS` in `opc_repository.py`). The server may revise these values; negotiated values are logged after connect and exposed as `opc_session_revised_timeout_milliseconds`.
### Reconnection interval
`OPC_RECONNECTION_INTERVAL` is in **seconds** (default `120`). It gates **background** reconnect after Tier-1 `Bad*`, closed protocol, or stale session (`last_reconnection_time` is updated only in `_reconnect_locked()`). It limits load on the OPC server when many workflows fail at once.
## Concurrency: connection lock and session readiness
To allow **multiple concurrent writes** when the session is healthy, but **block all writes** while the connection is being torn down or re-established:
| Primitive | Role |
|-----------|------|
| `_connection_lock` (`asyncio.Lock`) | Held for the entire `disconnect``connect` path. Only one connection-maintenance task at a time. |
| `_session_ready` (`asyncio.Event`) | Set when a session is ready for writes; cleared before reconnect starts and set again after a successful connect. |
**Connection methods (caller holds `_connection_lock` for `_*_locked` helpers):**
| Method | Role |
|--------|------|
| `_create_client()` | Create asyncua `Client` + optional `set_security`; raises if `client` already exists |
| `_open_session()` | `client.connect()` + metrics; raises if session already open or client missing |
| `_connect_locked()` | `_create_client()` (when needed) + `_open_session()`; raises if already connected |
| `_disconnect_locked()` | Teardown session and clear `client` |
| `_reconnect_locked()` | `_disconnect_locked()` + `_connect_locked()`; sets `last_reconnection_time` |
Public `connect()` / `disconnect()` acquire the lock and call `_connect_locked()` / `_disconnect_locked()`.
**Write path (`write_data`):**
1. If a reconnect task is **in flight****fail immediately** (`opc_error_kind=reconnect_in_progress`).
2. If `_session_ready` is cleared and no task is running → schedule reconnect (`SessionNotReady`); fail with `connection_lost` or `reconnect_in_progress` if a task started.
3. `validate_connection()` checks `protocol.state` only (read-only). If closed → schedule reconnect (`ProtocolClosed`) and fail with `opc_error_kind=connection_lost`.
4. Single `get_node` + `write_value` (no retry). Tier-1 `Bad*` on write also schedules reconnect.
**Reconnect path (`_run_reconnect`):**
1. `_start_reconnect` clears `_session_ready` and schedules the task when the interval allows and `_allow_reconnect` is true.
2. `async with _connection_lock:``_reconnect_locked()`.
3. `_session_ready` is set on successful `_open_session()`.
4. `disconnect()` sets `_allow_reconnect=False` so shutdown does not respawn sessions.
A second `_connect_locked()` while a session is already open raises `OpcSessionAlreadyConnectedError` (disconnect first).
**asyncua note:** Concurrent `write_value` on the same session is only safe if the stack tolerates it. If production shows issues, serialize writes with an optional `asyncio.Semaphore(1)` while keeping the connection lock semantics above.
**Future threads:** replace `asyncio.Lock` / `Event` with `threading` primitives or route all OPC I/O through one dedicated loop.
## Reconnect triggers
Background reconnect is scheduled when:
- `validate_connection()` sees a closed or missing protocol (`ProtocolClosed`).
- `_session_ready` is clear after a failed reconnect (`SessionNotReady`).
- A write raises a Tier-1 `UaStatusCodeError` in `RECONNECTABLE_OPC_BAD_NAMES`.
For Tier-1 `Bad*` when the server invalidates the session (e.g. `BadSessionIdInvalid`) but the client still sees transport as open, `write_data` fails once, records the OPC status in metrics, and **schedules** reconnect if:
- The exception is a `UaStatusCodeError` whose name is in `RECONNECTABLE_OPC_BAD_NAMES` (see plan), and
- `reconnection_interval` has elapsed since `last_reconnection_time`, and
- No reconnect task is already running.
There is **no write retry**: the failed export is not sent again in the same activity.
## Prediction confidence and PostgreSQL comments
| `prediction_confidence` | Meaning |
|-------------------------|---------|
| (unchanged) | Successful OPC export |
| **12** | Generic OPC write failure (`OPC_WRITTING_ERROR_CONFIDENCE`) |
| **14** | Tier-1 session/channel `Bad*` on export (`OPC_SESSION_BAD_CONFIDENCE`) |
| **14** | Write while reconnect in progress (`OPC_SESSION_BAD_CONFIDENCE`, comment `OPC UA reconnect in progress`) |
| **13** | PI Web API write failure (separate path) |
Session/channel errors use a stable comment for counting:
```text
OPC UA session/channel error: BadSessionIdInvalid
```
Reconnect-in-progress exports use:
```text
OPC UA reconnect in progress
```
Example SQL:
```sql
SELECT count(*) FROM predictions WHERE prediction_confidence = 14;
SELECT count(*) FROM predictions WHERE comments LIKE 'OPC UA session/channel error:%';
```
## Prometheus metrics (`opc_*`)
Defined in [`laborious/metrics.py`](../laborious/metrics.py). Do not rename in production without a dashboard migration.
| Metric | Purpose |
|--------|---------|
| `opc_connections_initiated_total` | Connection attempts |
| `opc_connections_failed_total` | Failed connects |
| `opc_connection_status` | Gauge 1=connected, 0=disconnected |
| `opc_session_created_total` | Session established after connect |
| `opc_session_closed_total` | Disconnect initiated |
| `opc_session_revised_timeout_milliseconds` | Negotiated session timeout (ms) |
| `opc_write_attempts_total` | Per write; label `result` = `OK` or exception name |
| `opc_write_inter_arrival_over_session_timeout_total` | Successful writes spaced longer than revised session timeout |
Legacy activity metrics: `laborious_prediction_opc_writing_count`, `laborious_prediction_opc_writing_response_time_monitor`.
## Environment variables
| Variable | Default | Description |
|----------|---------|-------------|
| `OPC_CONFIG` | — | JSON map of server configs (overrides single-server env) |
| `OPC_ID` | `1` | Server id |
| `OPC_URL` | `opc.tcp://localhost:4840` | Endpoint |
| `OPC_SERVER_NAME` | `default_server` | Label for metrics/logs |
| `OPC_SERVER_URI` | same as URL | Application URI / cert SAN |
| `OPC_CERT_PATH` | — | Client certificate (secure mode) |
| `OPC_PRIVATE_KEY_PATH` | — | Client private key |
| `OPC_SERVER_CERT_PATH` | — | Server certificate |
| `OPC_RECONNECTION_INTERVAL` | `120` | Minimum seconds between reconnects |
## Operations checklist
- Correlate `BadSessionIdInvalid` in `opc_write_attempts_total` with `opc_session_closed_total` / `opc_session_created_total` (reconnect may finish after the row is stored with confidence 14).
- Use confidence **14** and comment prefix for session invalidation rates; use **12** for other OPC failures.
- Respect `OPC_RECONNECTION_INTERVAL` under parallel load; bursts of confidence 14 are expected until the next successful cycle.
## Related tests
- Unit: [`tests/laborious/utils/repository/test_opc_repository.py`](../tests/laborious/utils/repository/test_opc_repository.py)
- Unit: [`tests/laborious/activities/test_opc.py`](../tests/laborious/activities/test_opc.py)
- E2E (mock OPC): [`e2e/test_predictions_batch_format_export.py`](../e2e/test_predictions_batch_format_export.py)
- E2E (in-process asyncua server + real `OpcRepository`): [`e2e/test_opc_real_server.py`](../e2e/test_opc_real_server.py) — scenarios 3.1.2, 3.2.2, 3.2.4, 3.2.5
- Scenarios: [`e2e/scenarios.md`](../e2e/scenarios.md)

View File

@@ -2,7 +2,14 @@
Pytest configuration and fixtures for E2E tests. Pytest configuration and fixtures for E2E tests.
""" """
import sys
from unittest.mock import AsyncMock, MagicMock, patch from unittest.mock import AsyncMock, MagicMock, patch
# E2E workflows under test do not run ModelAnalysis; stub before Activities import.
_model_analysis_module = MagicMock()
_model_analysis_module.ModelAnalysis = MagicMock
sys.modules.setdefault('sientia', MagicMock())
sys.modules.setdefault('sientia.ModelAnalysis', _model_analysis_module)
from io import BytesIO from io import BytesIO
import pandas as pd import pandas as pd
@@ -14,6 +21,7 @@ from testcontainers.postgres import PostgresContainer
from temporalio.testing import WorkflowEnvironment from temporalio.testing import WorkflowEnvironment
from temporalio.worker import Worker from temporalio.worker import Worker
from e2e.opc_test_server import OpcE2ETestServer
from laborious.activities.activities import Activities from laborious.activities.activities import Activities
from laborious.workflows.predictions_batch import PredictionsBatch from laborious.workflows.predictions_batch import PredictionsBatch
from laborious.workflows.sub_workflows.prediction_process import PredictionProcess from laborious.workflows.sub_workflows.prediction_process import PredictionProcess
@@ -297,6 +305,19 @@ def mock_opc_repository():
mock_repo.disconnect = AsyncMock() mock_repo.disconnect = AsyncMock()
return mock_repo return mock_repo
@pytest_asyncio.fixture
async def opc_e2e_server():
"""
In-process asyncua OPC UA server for E2E tests against OpcRepository.
"""
server = OpcE2ETestServer()
await server.start()
try:
yield server
finally:
await server.stop()
@pytest_asyncio.fixture @pytest_asyncio.fixture
def patch_create_engine(postgres_engine): def patch_create_engine(postgres_engine):
"""Patch create_engine to return test postgres_engine.""" """Patch create_engine to return test postgres_engine."""
@@ -577,3 +598,84 @@ async def temporal_worker_real_minio(temporal_test_env, test_activities_real_min
activities=_worker_activity_list(test_activities_real_minio), activities=_worker_activity_list(test_activities_real_minio),
) as worker: ) as worker:
yield worker yield worker
@pytest_asyncio.fixture(scope='function')
async def test_activities_real_opc(
postgres_engine,
postgres_container,
opc_e2e_server: OpcE2ETestServer,
mock_logger,
notification_handler,
metrics_controller,
patch_create_engine,
patch_minio_repository,
patch_mlflow,
patch_pi_web_api_repository,
):
"""
Activities with a real OpcRepository connected to the in-process OPC UA server.
"""
activities = Activities(
postgres_config={
'host': 'localhost',
'port': postgres_container.get_exposed_port(5432),
'user': 'test',
'password': 'test',
'dbname': 'test',
'min_connections': 1,
'max_connections': 5,
},
mlflow_config={
'host': 'http://localhost',
'port': '5000',
'username': 'test',
'password': 'test',
},
minio_config={
'endpoint_url': 'localhost:9000',
'access_key': 'test',
'secret_key': 'test',
'default_bucket': 'test-bucket',
'retention_hours': 24,
'secure': False,
},
opc_config={
'1': {
'id': '1',
'server_name': 'e2e-opc',
'url': opc_e2e_server.url,
'server_uri': opc_e2e_server.url,
'cert_path': None,
'private_key_path': None,
'server_cert_path': None,
'reconnection_interval': 0,
}
},
pi_web_api_config={
'base_url': 'http://localhost:8080',
'auth_type': 'bearer',
'auth_token': 'test_token',
},
logger=mock_logger,
notification_handler=notification_handler,
)
await activities.init_opc()
repo = activities.opc_repository['1']
assert repo._session_ready.is_set(), 'OPC E2E server connection failed during init_opc'
try:
yield activities
finally:
await activities.shutdown()
@pytest_asyncio.fixture(scope='function')
async def temporal_worker_real_opc(temporal_test_env, test_activities_real_opc):
"""Temporal worker backed by Activities using the in-process OPC UA server."""
async with Worker(
temporal_test_env.client,
task_queue='test-queue',
workflows=[PredictionsBatch, PredictionProcess, FormatAndExportPrediction],
activities=_worker_activity_list(test_activities_real_opc),
) as worker:
yield worker

View File

@@ -64,7 +64,8 @@ def assert_prediction(
prediction: float = 0.5, prediction: float = 0.5,
prediction_confidence: int | Decimal = 0, prediction_confidence: int | Decimal = 0,
prediction_status: str = 'Good', prediction_status: str = 'Good',
comments: str = '', comments: str | None = None,
comments_contains: str | None = None,
) -> None: ) -> None:
""" """
Assert exactly one prediction row exists for model_id with expected columns. Assert exactly one prediction row exists for model_id with expected columns.
@@ -75,7 +76,8 @@ def assert_prediction(
prediction: Expected prediction value. prediction: Expected prediction value.
prediction_confidence: Expected confidence (int or Decimal for numeric column). prediction_confidence: Expected confidence (int or Decimal for numeric column).
prediction_status: Expected status string. prediction_status: Expected status string.
comments: Expected comments string. comments: Expected exact comments string (optional).
comments_contains: Substring expected in comments when queued (optional).
""" """
import pytest import pytest
@@ -98,7 +100,12 @@ def assert_prediction(
str(prediction_confidence) str(prediction_confidence)
), f'Expected prediction_confidence={prediction_confidence}, got {row[2]}' ), f'Expected prediction_confidence={prediction_confidence}, got {row[2]}'
assert row[3] == prediction_status, f"Expected prediction_status='{prediction_status}', got {row[3]}" assert row[3] == prediction_status, f"Expected prediction_status='{prediction_status}', got {row[3]}"
assert row[4] == comments, f"Expected comments='{comments}', got {row[4]}" if comments is not None:
assert row[4] == comments, f"Expected comments='{comments}', got {row[4]}"
if comments_contains is not None:
assert comments_contains in row[4], (
f"Expected comments to contain '{comments_contains}', got {row[4]}"
)
def assert_continue( def assert_continue(

189
e2e/opc_test_server.py Normal file
View File

@@ -0,0 +1,189 @@
"""
In-process OPC UA server for E2E tests (asyncua).
Provides writable prediction/confidence nodes and optional write faults
(Tier-1 BadSessionIdInvalid via PreWrite callback).
"""
from __future__ import annotations
import socket
from dataclasses import dataclass
from typing import TYPE_CHECKING
from asyncua import Server, ua
from asyncua.common.callback import CallbackType
from asyncua.common.utils import ServiceError
if TYPE_CHECKING:
from asyncua.common.node import Node
UNKNOWN_NODE_ID = 'ns=99;i=9999'
@dataclass(frozen=True)
class OpcE2ENodeIds:
"""NodeId strings used in opc_output_config for E2E workflows."""
prediction: str
confidence: str
unknown: str = UNKNOWN_NODE_ID
class OpcE2ETestServer:
"""
Ephemeral asyncua server with Laborious E2E variables and controllable faults.
Args:
host: Bind address (default 127.0.0.1).
"""
def __init__(self, host: str = '127.0.0.1') -> None:
self._host = host
self._server: Server | None = None
self._prediction_node: Node | None = None
self._confidence_node: Node | None = None
self._session_bad_on_write = False
self._url: str | None = None
self._node_ids: OpcE2ENodeIds | None = None
@property
def url(self) -> str:
if self._url is None:
raise RuntimeError('OPC E2E server is not started')
return self._url
@property
def node_ids(self) -> OpcE2ENodeIds:
if self._node_ids is None:
raise RuntimeError('OPC E2E server is not started')
return self._node_ids
def set_session_bad_on_write(self, enabled: bool) -> None:
"""
When enabled, every client Write is rejected with BadSessionIdInvalid.
Args:
enabled (bool): Turn Tier-1 session fault injection on or off.
"""
self._session_bad_on_write = enabled
async def start(self) -> OpcE2ENodeIds:
"""
Start the OPC UA server on a free TCP port.
Return:
OpcE2ENodeIds: NodeId strings for prediction and confidence tags.
"""
port = _free_port(self._host)
self._url = f'opc.tcp://{self._host}:{port}/freeopcua/server/'
server = Server()
server.set_endpoint(self._url)
await server.init()
server.iserver.callback_service.addListener(
CallbackType.PreWrite,
self._pre_write_callback,
)
idx = await server.register_namespace('http://sientia.test/laborious-e2e')
e2e_object = await server.nodes.objects.add_object(idx, 'LaboriousE2E')
prediction = await e2e_object.add_variable(
idx,
'Prediction',
ua.Variant(0.0, ua.VariantType.Float),
)
confidence = await e2e_object.add_variable(
idx,
'Confidence',
ua.Variant(0.0, ua.VariantType.Float),
)
await prediction.set_writable()
await confidence.set_writable()
await server.start()
self._server = server
self._prediction_node = prediction
self._confidence_node = confidence
self._node_ids = OpcE2ENodeIds(
prediction=prediction.nodeid.to_string(),
confidence=confidence.nodeid.to_string(),
)
return self._node_ids
async def stop(self) -> None:
"""Stop the OPC UA server and release the listening port."""
if self._server is not None:
await self._server.stop()
self._server = None
self._prediction_node = None
self._confidence_node = None
self._url = None
self._node_ids = None
self._session_bad_on_write = False
async def read_prediction(self) -> float:
"""
Read the current prediction variable value from the address space.
Return:
float: Stored prediction value.
"""
if self._prediction_node is None:
raise RuntimeError('OPC E2E server is not started')
value = await self._prediction_node.read_value()
return float(value)
async def read_confidence(self) -> float:
"""
Read the current confidence variable value from the address space.
Return:
float: Stored confidence value.
"""
if self._confidence_node is None:
raise RuntimeError('OPC E2E server is not started')
value = await self._confidence_node.read_value()
return float(value)
async def _pre_write_callback(self, _event, _service) -> None:
if self._session_bad_on_write:
raise ServiceError(ua.StatusCodes.BadSessionIdInvalid)
def _free_port(host: str) -> int:
with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as sock:
sock.bind((host, 0))
return int(sock.getsockname()[1])
def build_opc_output_config(
node_ids: OpcE2ENodeIds,
*,
prediction_tag: str | None = None,
confidence_tag: str | None = None,
prediction_only: bool = False,
server_key: str = '1',
) -> dict[str, dict]:
"""
Build opc_output_config for PredictionsBatch using real server NodeIds.
Args:
node_ids (OpcE2ENodeIds): Node ids from OpcE2ETestServer.
prediction_tag (str | None): Override prediction NodeId (default: node_ids.prediction).
confidence_tag (str | None): Override confidence NodeId (default: node_ids.confidence).
prediction_only (bool): When True, omit confidence_tags (single write per activity).
server_key (str): OPC server id key in opc_output_config.
Return:
dict: opc_output_config payload for workflow input.
"""
pred = prediction_tag if prediction_tag is not None else node_ids.prediction
conf = confidence_tag if confidence_tag is not None else node_ids.confidence
server_config: dict = {
'prediction_tags': {pred: {'data_type': 'float'}},
}
if not prediction_only:
server_config['confidence_tags'] = {conf: {'data_type': 'float'}}
return {server_key: server_config}

View File

@@ -8,6 +8,8 @@ This document describes all possible test scenarios for the `predictions_batch`
- **Dependencies**: install dev requirements (includes `testcontainers[postgres,minio]`). - **Dependencies**: install dev requirements (includes `testcontainers[postgres,minio]`).
- **Invocation**: run only integration-marked tests, for example: `pytest e2e/ -m integration`. - **Invocation**: run only integration-marked tests, for example: `pytest e2e/ -m integration`.
- **MinIO tests**: `e2e/test_minio_offload.py` exercises real S3 uploads; other E2E modules continue to mock MinIO on the worker used by most scenarios. - **MinIO tests**: `e2e/test_minio_offload.py` exercises real S3 uploads; other E2E modules continue to mock MinIO on the worker used by most scenarios.
- **OPC tests (real server)**: `e2e/test_opc_real_server.py` uses an in-process **asyncua** server and real `OpcRepository` (`test_activities_real_opc`). Scenarios 3.1.2, 3.2.2, 3.2.4, and 3.2.5 are covered there. Other E2E modules keep the OPC mock.
- Run only OPC real-server tests: `pytest e2e/test_opc_real_server.py -m "integration and opc"`.
## Workflow Overview ## Workflow Overview
@@ -495,6 +497,48 @@ These paths do **not** rely on Temporal activity retries for export failures: th
--- ---
#### Scenario 3.2.4: OPC Session / Channel Bad* (Tier-1)
**Description**: OPC write fails with a Tier-1 session or channel status (e.g. `BadSessionIdInvalid`) while transport may still appear open on the client
**Input**:
- Valid prediction and OPC output config
- Mock or server returning Tier-1 `UaStatusCodeError` on write (no write retry in the same activity)
**Expected Behavior**:
- `write_opc_data` fails forward for affected tags; background reconnect may be scheduled if `OPC_RECONNECTION_INTERVAL` allows
- Workflow **completes**
- PostgreSQL row uses **`prediction_confidence` 14** and comment prefix `OPC UA session/channel error:` (including OPC status name)
- `opc_write_attempts_total` records `result=BadSessionIdInvalid` (or matching status); no second write attempt in the same activity
**Assertions**:
- Workflow completes
- `prediction_confidence = 14`
- `comments` matches `OPC UA session/channel error:%`
- Generic OPC error confidence **12** is not used for this case
**Reference**: [docs/opc-communication.md](../docs/opc-communication.md), plan `.cursor/plans/opc_bad_reconnect_ac4c6045.plan.md`
---
#### Scenario 3.2.5: OPC Write Blocked During Reconnect
**Description**: A write is attempted while the repository is reconnecting (session not ready)
**Input**:
- Valid prediction
- Simulated slow reconnect (e.g. delayed `connect`) or concurrent writes where the first triggers reconnect
**Expected Behavior**:
- Second write (or parallel write) is rejected **immediately** when reconnect is in progress or `_session_ready` is cleared — **without** calling `write_value`
- No wait/sleep on the write path; no duplicate `connect` from parallel writers (connection lock)
- `prediction_confidence = 14`, `comments = OPC UA reconnect in progress` (distinguish from Tier-1 `Bad*` via comment prefix in SQL)
**Assertions**:
- At most one reconnect sequence (`disconnect` + `connect`) for the overlapping window
- No write retry after failure
- Tests in `test_opc_repository` (unit) and optional e2e in `test_predictions_batch_format_export.py`
---
#### Scenario 3.2.3: PI Web API Partial Write Error #### Scenario 3.2.3: PI Web API Partial Write Error
**Description**: Two prediction tags attempt to be written to PI Web API, but only one succeeds **Description**: Two prediction tags attempt to be written to PI Web API, but only one succeeds

194
e2e/test_opc_real_server.py Normal file
View File

@@ -0,0 +1,194 @@
"""
E2E tests for OPC export using an in-process asyncua server and real OpcRepository.
Covers scenarios 3.1.2, 3.2.2, 3.2.4, and 3.2.5 from e2e/scenarios.md.
Mock-based OPC tests remain in test_predictions_batch_format_export.py.
"""
import asyncio
import pytest
from temporalio.testing import WorkflowEnvironment
from temporalio.worker import Worker
from e2e.helpers import assert_prediction, insert_sample_data, make_workflow_id, start_and_await_workflow
from e2e.opc_test_server import UNKNOWN_NODE_ID, OpcE2ETestServer, build_opc_output_config
from e2e.test_predictions_batch_format_export import get_base_input_data
from laborious.activities.activities import Activities
from laborious.activities.opc import OPC_RECONNECT_IN_PROGRESS_COMMENT
from laborious.utils.repository.opc_repository import OpcRepository
from laborious.workflows.predictions_batch import PredictionsBatch
async def _slow_reconnect_under_lock(repo: OpcRepository, hold_seconds: float = 0.75) -> None:
"""
Hold the connection lock briefly so concurrent writes see reconnect_in_progress.
Args:
repo (OpcRepository): Connected repository.
hold_seconds (float): Time to keep the lock before reconnecting.
"""
async with repo._connection_lock:
await asyncio.sleep(hold_seconds)
await repo._reconnect_locked()
@pytest.mark.asyncio
@pytest.mark.integration
@pytest.mark.opc
async def test_scenario_3_1_2_export_with_opc_only_real_server(
temporal_test_env: WorkflowEnvironment,
temporal_worker_real_opc: Worker,
test_activities_real_opc: Activities,
opc_e2e_server: OpcE2ETestServer,
postgres_engine,
):
"""
Scenario 3.1.2 (real OPC): connect, write prediction and confidence, verify server values.
"""
client = temporal_test_env.client
model_id = 412
insert_sample_data(postgres_engine, model_id, [23.5, 78.2])
input_data = get_base_input_data(model_id)
input_data['opc_output_config'] = build_opc_output_config(opc_e2e_server.node_ids)
input_data['pi_web_api_output_config'] = None
await start_and_await_workflow(
client,
PredictionsBatch.run,
input_data,
make_workflow_id('test-opc-real-happy'),
)
test_activities_real_opc.pi_web_api_client.write_value.assert_not_called()
assert await opc_e2e_server.read_prediction() == pytest.approx(0.5)
assert await opc_e2e_server.read_confidence() == pytest.approx(0.0)
assert_prediction(postgres_engine, model_id)
@pytest.mark.asyncio
@pytest.mark.integration
@pytest.mark.opc
async def test_scenario_3_2_2_opc_write_error_real_server(
temporal_test_env: WorkflowEnvironment,
temporal_worker_real_opc: Worker,
test_activities_real_opc: Activities,
opc_e2e_server: OpcE2ETestServer,
postgres_engine,
):
"""
Scenario 3.2.2 (real OPC): unknown NodeId yields generic write failure (confidence 12).
"""
client = temporal_test_env.client
model_id = 422
insert_sample_data(postgres_engine, model_id, [23.5, 78.2])
node_ids = opc_e2e_server.node_ids
input_data = get_base_input_data(model_id)
input_data['opc_output_config'] = build_opc_output_config(
node_ids,
prediction_tag=UNKNOWN_NODE_ID,
confidence_tag=UNKNOWN_NODE_ID,
)
input_data['pi_web_api_output_config'] = None
await start_and_await_workflow(
client,
PredictionsBatch.run,
input_data,
make_workflow_id('test-opc-real-bad-node'),
)
assert_prediction(
postgres_engine,
model_id,
prediction_confidence=12,
comments='Some data could not be written to OPC servers',
)
@pytest.mark.asyncio
@pytest.mark.integration
@pytest.mark.opc
async def test_scenario_3_2_4_opc_session_bad_real_server(
temporal_test_env: WorkflowEnvironment,
temporal_worker_real_opc: Worker,
test_activities_real_opc: Activities,
opc_e2e_server: OpcE2ETestServer,
postgres_engine,
):
"""
Scenario 3.2.4 (real OPC): server PreWrite fault injects BadSessionIdInvalid (confidence 14).
"""
client = temporal_test_env.client
model_id = 424
insert_sample_data(postgres_engine, model_id, [23.5, 78.2])
opc_e2e_server.set_session_bad_on_write(True)
try:
input_data = get_base_input_data(model_id)
input_data['opc_output_config'] = build_opc_output_config(
opc_e2e_server.node_ids,
prediction_only=True,
)
input_data['pi_web_api_output_config'] = None
await start_and_await_workflow(
client,
PredictionsBatch.run,
input_data,
make_workflow_id('test-opc-real-session-bad'),
)
finally:
opc_e2e_server.set_session_bad_on_write(False)
assert_prediction(
postgres_engine,
model_id,
prediction_confidence=14,
comments_contains='OPC UA session/channel error: BadSessionIdInvalid',
)
@pytest.mark.asyncio
@pytest.mark.integration
@pytest.mark.opc
async def test_scenario_3_2_5_opc_write_blocked_during_reconnect_real_server(
temporal_test_env: WorkflowEnvironment,
temporal_worker_real_opc: Worker,
test_activities_real_opc: Activities,
opc_e2e_server: OpcE2ETestServer,
postgres_engine,
):
"""
Scenario 3.2.5 (real OPC): writes rejected while reconnect holds the connection lock.
"""
client = temporal_test_env.client
model_id = 425
insert_sample_data(postgres_engine, model_id, [23.5, 78.2])
repo = test_activities_real_opc.opc_repository['1']
repo._session_ready.clear()
reconnect_task = asyncio.create_task(_slow_reconnect_under_lock(repo))
input_data = get_base_input_data(model_id)
input_data['opc_output_config'] = build_opc_output_config(opc_e2e_server.node_ids)
input_data['pi_web_api_output_config'] = None
try:
await start_and_await_workflow(
client,
PredictionsBatch.run,
input_data,
make_workflow_id('test-opc-real-reconnect-block'),
)
finally:
await reconnect_task
assert_prediction(
postgres_engine,
model_id,
prediction_confidence=14,
comments_contains=OPC_RECONNECT_IN_PROGRESS_COMMENT,
)

View File

@@ -153,7 +153,6 @@ async def test_scenario_3_1_1_default_prediction_export(
'addr_1', 'addr_1',
0, 0,
'float', 'float',
ANY,
{ {
'model_id': 311, 'model_id': 311,
'model_name': 'test_model', 'model_name': 'test_model',
@@ -165,7 +164,6 @@ async def test_scenario_3_1_1_default_prediction_export(
'addr_2', 'addr_2',
2, 2,
'float', 'float',
ANY,
{ {
'model_id': 311, 'model_id': 311,
'model_name': 'test_model', 'model_name': 'test_model',
@@ -252,20 +250,28 @@ async def test_scenario_3_1_2_export_with_opc_only(
opc_write_data = cast(Any, test_activities.opc_repository['1'].write_data) opc_write_data = cast(Any, test_activities.opc_repository['1'].write_data)
opc_write_data.assert_has_calls( opc_write_data.assert_has_calls(
[ [
call('addr_1', 0.5, 'float', ANY, call(
{ 'addr_1',
'model_id': 312, 0.5,
'model_name': 'test_model', 'float',
'schedule_name': 'test-schedule', {
'workflow_name': 'predictions_batch', 'model_id': 312,
}), 'model_name': 'test_model',
call('addr_2', 0, 'float', ANY, 'schedule_name': 'test-schedule',
{ 'workflow_name': 'predictions_batch',
'model_id': 312, },
'model_name': 'test_model', ),
'schedule_name': 'test-schedule', call(
'workflow_name': 'predictions_batch', 'addr_2',
}), 0,
'float',
{
'model_id': 312,
'model_name': 'test_model',
'schedule_name': 'test-schedule',
'workflow_name': 'predictions_batch',
},
),
] ]
) )
@@ -500,20 +506,28 @@ async def test_scenario_3_1_5_export_without_transformed_data(
opc_write_data = cast(Any, test_activities.opc_repository['1'].write_data) opc_write_data = cast(Any, test_activities.opc_repository['1'].write_data)
opc_write_data.assert_has_calls( opc_write_data.assert_has_calls(
[ [
call('addr_1', 0.5, 'float', ANY, call(
{ 'addr_1',
'model_id': 315, 0.5,
'model_name': 'test_model', 'float',
'schedule_name': 'test-schedule', {
'workflow_name': 'predictions_batch', 'model_id': 315,
}), 'model_name': 'test_model',
call('addr_2', 0, 'float', ANY, 'schedule_name': 'test-schedule',
{ 'workflow_name': 'predictions_batch',
'model_id': 315, },
'model_name': 'test_model', ),
'schedule_name': 'test-schedule', call(
'workflow_name': 'predictions_batch', 'addr_2',
}), 0,
'float',
{
'model_id': 315,
'model_name': 'test_model',
'schedule_name': 'test-schedule',
'workflow_name': 'predictions_batch',
},
),
] ]
) )
@@ -648,6 +662,67 @@ async def test_scenario_3_2_2_opc_write_error(
) )
@pytest.mark.asyncio
@pytest.mark.integration
async def test_scenario_3_2_4_opc_session_bad_error(
temporal_test_env: WorkflowEnvironment,
temporal_worker: Worker,
test_activities: Activities,
postgres_engine,
):
"""
Scenario 3.2.4: OPC session/channel Tier-1 Bad* (e.g. BadSessionIdInvalid).
PostgreSQL stores prediction_confidence 14 and a stable session error comment.
"""
client = temporal_test_env.client
model_id = 324
insert_sample_data(postgres_engine, model_id, [23.5, 78.2])
opc_write_data = cast(Any, test_activities.opc_repository['1'].write_data)
opc_write_data.return_value = (
False,
{
'notification_id': 'OPC_WRITE_DATA_ERROR_1',
'message': 'BadSessionIdInvalid',
'block': 'opc_repository',
'level': NotificationLevel.ERROR,
'attachment_content': 'BadSessionIdInvalid',
'opc_error_kind': 'session_bad',
'opc_status': 'BadSessionIdInvalid',
},
)
input_data = get_base_input_data(model_id)
input_data['opc_output_config'] = {
'1': {
'prediction_tags': {
'addr_1': {
'data_type': 'float',
}
},
'confidence_tags': {
'addr_2': {
'data_type': 'float',
}
},
}
}
input_data['pi_web_api_output_config'] = None
await start_and_await_workflow(
client, PredictionsBatch.run, input_data, make_workflow_id('test-opc-session-bad')
)
assert_prediction(
postgres_engine,
model_id,
prediction_confidence=14,
comments='OPC UA session/channel error: BadSessionIdInvalid',
)
@pytest.mark.asyncio @pytest.mark.asyncio
@pytest.mark.integration @pytest.mark.integration

25
inter_arrival.py Normal file
View File

@@ -0,0 +1,25 @@
# %%
# Load logs.txt
with open('logs.txt', 'r') as file:
lines = file.readlines()
# %%
import re
# Grep "inter-arrival_s=number" with regex
intervals = []
for line in lines:
match = re.search(r'inter-arrival_s=([0-9.]+)', line)
if match:
intervals.append(float(match.group(1)))
# %%
print(intervals)
# %%
import matplotlib.pyplot as plt
plt.plot(intervals)
plt.ylabel('Inter-arrival time (s)')
plt.xlabel('Sample')
plt.title('Inter-arrival time distribution')
plt.show()
# %%

View File

@@ -164,6 +164,6 @@ class Activities(Storage, MLFlow, Gates, OPC, ModelMetrics, API):
Storage.close(self) Storage.close(self)
MLFlow.close(self) MLFlow.close(self)
Gates.close(self) Gates.close(self)
await OPC.close(self) await OPC.aclose(self)
ModelMetrics.close(self) ModelMetrics.close(self)
API.close(self) API.close(self)

View File

@@ -15,6 +15,45 @@ with workflow.unsafe.imports_passed_through():
from laborious.utils.repository.opc_repository import OpcRepository from laborious.utils.repository.opc_repository import OpcRepository
OPC_WRITTING_ERROR_CONFIDENCE = 12 OPC_WRITTING_ERROR_CONFIDENCE = 12
OPC_SESSION_BAD_CONFIDENCE = 14
OPC_SESSION_BAD_COMMENT_PREFIX = 'OPC UA session/channel error:'
OPC_WRITTING_ERROR_MESSAGE = 'Some data could not be written to OPC servers'
OPC_RECONNECT_IN_PROGRESS_COMMENT = 'OPC UA reconnect in progress'
OPC_COMMENT_SEPARATOR = ' | '
def _opc_session_bad_comment(opc_status: str | None) -> str:
status = opc_status or 'Unknown'
return f'{OPC_SESSION_BAD_COMMENT_PREFIX} {status}'
def _apply_opc_write_error(
error_info: dict[str, Any] | None,
session_bad_seen: bool,
session_bad_status: str | None,
reconnect_in_progress_seen: bool,
) -> tuple[bool, str | None, bool]:
"""
Update session/reconnect flags from an OPC write error payload.
Args:
error_info: Repository error details, or None when the write succeeded.
session_bad_seen: Whether a session_bad error was seen so far.
session_bad_status: Last known OPC status for session errors.
reconnect_in_progress_seen: Whether reconnect_in_progress was seen so far.
Return:
Updated (session_bad_seen, session_bad_status, reconnect_in_progress_seen).
"""
if not error_info:
return session_bad_seen, session_bad_status, reconnect_in_progress_seen
kind = error_info.get('opc_error_kind')
if kind == 'session_bad':
return True, error_info.get('opc_status', session_bad_status), reconnect_in_progress_seen
if kind == 'reconnect_in_progress':
return session_bad_seen, session_bad_status, True
return session_bad_seen, session_bad_status, reconnect_in_progress_seen
class OPC(SientiaMonitoring): class OPC(SientiaMonitoring):
@@ -118,29 +157,18 @@ class OPC(SientiaMonitoring):
data_type: str, data_type: str,
tag_type: str, tag_type: str,
metadata: dict[str, Any], metadata: dict[str, Any],
) -> float | None: ) -> tuple[float | None, dict[str, Any] | None]:
""" """
Write data to a specific OPC server tag with comprehensive error handling. Write data to a specific OPC server tag with comprehensive error handling.
This method provides a secure and reliable way to write data to OPC servers Return:
with automatic error handling, notification integration, and detailed logging. tuple[float | None, dict[str, Any] | None]: Response time on success, or
It validates server availability before attempting write operations and (None, error info_data) on repository failure.
provides comprehensive error reporting for operational monitoring.
Args:
- server_id (str): The id of the OPC server.
- tag (str): The tag to write to.
- data (Any): The data to write.
- data_type (str): The data type.
- tag_type (str): The tag type.
Returns:
- bool: True if the data was written successfully, False otherwise.
""" """
try: try:
is_success, info_data = await self.opc_repository[server_id].write_data( is_success, info_data = await self.opc_repository[server_id].write_data(
tag, data, data_type, self.logger, metadata tag, data, data_type, metadata
) )
if not is_success: if not is_success:
await self.send_notification_async( await self.send_notification_async(
@@ -151,8 +179,8 @@ class OPC(SientiaMonitoring):
level=info_data.get('level', NotificationLevel.ERROR), level=info_data.get('level', NotificationLevel.ERROR),
attachment_content=info_data.get('attachment_content', None), attachment_content=info_data.get('attachment_content', None),
) )
return None return None, info_data
return info_data['response_time'] return info_data['response_time'], None
except Exception as e: except Exception as e:
trace = traceback.format_exc() trace = traceback.format_exc()
await self.send_notification_async( await self.send_notification_async(
@@ -199,13 +227,69 @@ class OPC(SientiaMonitoring):
return False return False
return True return True
async def _write_tags_from_config(
self,
server_id: str,
tags_config: dict[str, dict[str, Any]],
data: DataFrame,
data_column: str,
tag_type: str,
log_label: str,
metadata: dict[str, Any],
) -> tuple[dict[str, float | None], bool, str | None, bool]:
"""
Write a group of OPC tags and collect response times and error flags.
Args:
server_id: Target OPC server identifier.
tags_config: Tag name to configuration mapping.
data: DataFrame with prediction/confidence columns.
data_column: Column name whose first row value is written.
tag_type: Tag category passed to write_data ('prediction' or 'confidence').
log_label: Human-readable label for success logs.
metadata: Context metadata for logging and notifications.
Return:
(response_times, session_bad_seen, session_bad_status, reconnect_in_progress_seen)
"""
response_times: dict[str, float | None] = {}
session_bad_seen = False
session_bad_status: str | None = None
reconnect_in_progress_seen = False
for tag, tag_config in tags_config.items():
response_time, error_info = await self.write_data(
server_id=server_id,
tag=tag,
data=data.head(1)[data_column].values[0],
data_type=tag_config['data_type'],
tag_type=tag_type,
metadata=metadata,
)
session_bad_seen, session_bad_status, reconnect_in_progress_seen = (
_apply_opc_write_error(
error_info,
session_bad_seen,
session_bad_status,
reconnect_in_progress_seen,
)
)
if response_time is not None:
self.info(
f'{log_label} written to OPC server {server_id} for tag {tag}.',
metadata,
)
response_times[tag] = response_time
return response_times, session_bad_seen, session_bad_status, reconnect_in_progress_seen
async def manage_output_tags( async def manage_output_tags(
self, self,
server_id: str, server_id: str,
config: dict[str, Any], config: dict[str, Any],
data: DataFrame, data: DataFrame,
metadata: dict[str, Any], metadata: dict[str, Any],
) -> tuple[bool, dict[str, float | None]]: ) -> tuple[bool, dict[str, float | None], bool, str | None, bool]:
""" """
Manage the writing of prediction and confidence data to OPC server tags. Manage the writing of prediction and confidence data to OPC server tags.
@@ -232,46 +316,47 @@ class OPC(SientiaMonitoring):
- overall_success: True if all configured tags were written successfully - overall_success: True if all configured tags were written successfully
- total_tags_written: Count of successfully written tags - total_tags_written: Count of successfully written tags
""" """
response_times: dict[str, float | None] = {} response_times: dict[str, float | None] = {}
session_bad_seen = False
session_bad_status: str | None = None
reconnect_in_progress_seen = False
if 'prediction_tags' in config: tag_groups = (
for tag, tag_config in config['prediction_tags'].items(): ('prediction_tags', 'prediction', 'prediction', 'Prediction data'),
response_time = await self.write_data( ('confidence_tags', 'prediction_confidence', 'confidence', 'Confidence data'),
server_id=server_id, )
tag=tag, for config_key, data_column, tag_type, log_label in tag_groups:
data=data.head(1)['prediction'].values[0], if config_key not in config:
data_type=tag_config['data_type'], continue
tag_type='prediction', (
metadata=metadata, group_times,
) group_session_bad,
if response_time is not None: group_status,
self.info( group_reconnect,
f'Prediction data written to OPC server {server_id} for tag {tag}.', ) = await self._write_tags_from_config(
metadata, server_id=server_id,
) tags_config=config[config_key],
response_times[tag] = response_time data=data,
data_column=data_column,
if 'confidence_tags' in config: tag_type=tag_type,
for tag, tag_config in config['confidence_tags'].items(): log_label=log_label,
response_time = await self.write_data( metadata=metadata,
server_id=server_id, )
tag=tag, response_times.update(group_times)
data=data.head(1)['prediction_confidence'].values[0], if group_session_bad:
data_type=tag_config['data_type'], session_bad_seen = True
tag_type='confidence', session_bad_status = group_status or session_bad_status
metadata=metadata, if group_reconnect:
) reconnect_in_progress_seen = True
if response_time is not None:
self.info(
f'Confidence data written to OPC server {server_id} for tag {tag}.',
metadata,
)
response_times[tag] = response_time
success = None not in response_times.values() success = None not in response_times.values()
return (
return success, response_times success,
response_times,
session_bad_seen,
session_bad_status,
reconnect_in_progress_seen,
)
@activity.defn(name='write_opc_data') @activity.defn(name='write_opc_data')
async def write_opc_data( async def write_opc_data(
@@ -301,6 +386,9 @@ class OPC(SientiaMonitoring):
self.info(f'Data to write: {data.size} rows', metadata) self.info(f'Data to write: {data.size} rows', metadata)
success = True success = True
session_bad_seen = False
session_bad_status: str | None = None
reconnect_in_progress_seen = False
metrics: dict[str, dict[str, float | None]] = {} metrics: dict[str, dict[str, float | None]] = {}
@@ -309,22 +397,48 @@ class OPC(SientiaMonitoring):
success = False success = False
continue continue
local_success, local_response_times = await self.manage_output_tags( (
server_id, config, data, metadata local_success,
) local_response_times,
local_session_bad,
local_status,
local_reconnect_in_progress,
) = await self.manage_output_tags(server_id, config, data, metadata)
metrics[server_id] = local_response_times metrics[server_id] = local_response_times
local_count = len(local_response_times) local_count = len(local_response_times)
success = success and local_success success = success and local_success
if local_session_bad:
session_bad_seen = True
session_bad_status = local_status or session_bad_status
if local_reconnect_in_progress:
reconnect_in_progress_seen = True
self.info( self.info(
f'Process completed for OPC server {server_id}: {local_count} of {len(config.get("prediction_tags", []))} prediction tags and {len(config.get("confidence_tags", []))} confidence tags', f'Process completed for OPC server {server_id}: {local_count} of {len(config.get("prediction_tags", []))} prediction tags and {len(config.get("confidence_tags", []))} confidence tags',
metadata, metadata,
) )
return self.process_confidence(data, success, metadata), metrics return (
self.process_confidence(
data,
success,
metadata,
session_bad=session_bad_seen,
opc_status=session_bad_status,
reconnect_in_progress=reconnect_in_progress_seen,
),
metrics,
)
def process_confidence( def process_confidence(
self, data: DataFrame, success: bool, metadata: dict[str, Any] self,
data: DataFrame,
success: bool,
metadata: dict[str, Any],
*,
session_bad: bool = False,
opc_status: str | None = None,
reconnect_in_progress: bool = False,
) -> dict[Hashable, Any]: ) -> dict[Hashable, Any]:
""" """
Process prediction confidence based on OPC write operation success. Process prediction confidence based on OPC write operation success.
@@ -352,22 +466,32 @@ class OPC(SientiaMonitoring):
This allows downstream systems to handle data quality appropriately. This allows downstream systems to handle data quality appropriately.
""" """
message = 'Some data could not be written to OPC servers'
if not success: if not success:
data['prediction_confidence'] = OPC_WRITTING_ERROR_CONFIDENCE comment_parts: list[str] = []
data['comments'] = message confidence = OPC_WRITTING_ERROR_CONFIDENCE
if session_bad:
comment_parts.append(_opc_session_bad_comment(opc_status))
confidence = OPC_SESSION_BAD_CONFIDENCE
if reconnect_in_progress:
comment_parts.append(OPC_RECONNECT_IN_PROGRESS_COMMENT)
confidence = OPC_SESSION_BAD_CONFIDENCE
if not comment_parts:
comment_parts.append(OPC_WRITTING_ERROR_MESSAGE)
comments = OPC_COMMENT_SEPARATOR.join(comment_parts)
data['prediction_confidence'] = confidence
data['comments'] = comments
self.debug( self.debug(
f'{message}, setting confidence to {OPC_WRITTING_ERROR_CONFIDENCE}.', f'OPC write issues, confidence={confidence}, comments={comments}',
metadata, metadata,
) )
else: else:
self.debug('Data written to OPC servers successfully.', metadata) self.debug('Data written to OPC servers successfully.', metadata)
return data.to_dict() return data.to_dict()
async def close(self): async def aclose(self):
""" """
Gracefully shutdown all OPC server connections and cleanup resources. Gracefully shutdown all OPC server connections and cleanup resources.

View File

@@ -202,7 +202,8 @@ class Storage(Postgres, MinioManager):
def close(self) -> None: def close(self) -> None:
"""Close Storage resources (MinIO client and Postgres engine).""" """Close Storage resources (MinIO client and Postgres engine)."""
Postgres.close(self) if hasattr(self, 'engine'):
Postgres.close(self)
MinioManager.close(self) MinioManager.close(self)
def __del__(self): def __del__(self):

View File

@@ -89,6 +89,40 @@ OPC_CONNECTION_STATUS = Gauge(
['pod_id', 'server_name', 'server_url'], ['pod_id', 'server_name', 'server_url'],
) )
_OPC_SESSION_DEBUG_LABELS = ['pod_id', 'server_name', 'runtime', 'opc_server_id', 'session_id']
OPC_SESSION_CREATED_TOTAL = Counter(
'opc_session_created_total',
'OPC UA sessions established (after successful connect)',
_OPC_SESSION_DEBUG_LABELS,
)
OPC_SESSION_CLOSED_TOTAL = Counter(
'opc_session_closed_total',
'OPC UA client disconnects completed (session tear-down initiated)',
_OPC_SESSION_DEBUG_LABELS,
)
OPC_SESSION_REVISED_TIMEOUT_MS = Gauge(
'opc_session_revised_timeout_milliseconds',
'Server-revised OPC UA session timeout (RevisedSessionTimeout) in ms after connect',
_OPC_SESSION_DEBUG_LABELS,
)
OPC_WRITE_ATTEMPT_LABELS = [*_OPC_SESSION_DEBUG_LABELS, 'model_id', 'model_name', 'result']
OPC_WRITE_ATTEMPTS_TOTAL = Counter(
'opc_write_attempts_total',
'OPC UA write attempts with session and outcome (result=OK or exception class name)',
OPC_WRITE_ATTEMPT_LABELS,
)
OPC_WRITE_INTER_ARRIVAL_OVER_SESSION_TIMEOUT_TOTAL = Counter(
'opc_write_inter_arrival_over_session_timeout_total',
'Successful writes where seconds since the previous successful write exceeded RevisedSessionTimeout (ms)',
_OPC_SESSION_DEBUG_LABELS,
)
# ================== Model metrics ================== # ================== Model metrics ==================
MODEL_READ_LAG = Histogram( MODEL_READ_LAG = Histogram(

View File

@@ -170,8 +170,19 @@ class MLFlowRepository(SientiaMonitoring):
# Sort by version number to get the latest # Sort by version number to get the latest
latest_version = max(stage_versions, key=lambda v: int(v.version)) latest_version = max(stage_versions, key=lambda v: int(v.version))
run_id = latest_version.source.split('/') source = latest_version.source
return run_id[2] if source is None:
raise mlflow.exceptions.MlflowException(
f"Model '{model_name}' version '{latest_version.version}' in stage '{stage}' "
'has no source URI to resolve run ID.'
)
parts = source.split('/')
if len(parts) <= 2 or not parts[2]:
raise mlflow.exceptions.MlflowException(
f"Model '{model_name}' version '{latest_version.version}' in stage '{stage}' "
f"has invalid source URI '{source}' for run ID resolution."
)
return parts[2]
def get_next_run_name(self, model_name: str) -> str: def get_next_run_name(self, model_name: str) -> str:
""" """

View File

@@ -9,6 +9,7 @@ from typing import Any
from asyncua import Client from asyncua import Client
from asyncua.crypto.security_policies import SecurityPolicyBasic256 from asyncua.crypto.security_policies import SecurityPolicyBasic256
from asyncua.ua import DataValue, Variant, VariantType from asyncua.ua import DataValue, Variant, VariantType
from asyncua.ua.uaerrors import UaStatusCodeError
from sientia_do.notifications.handlers import CoreNotificationHandler as NotificationHandler from sientia_do.notifications.handlers import CoreNotificationHandler as NotificationHandler
from sientia_do.notifications.models import NotificationLevel from sientia_do.notifications.models import NotificationLevel
from sientia_do.observability.logger import Logger from sientia_do.observability.logger import Logger
@@ -17,6 +18,119 @@ from sientia_do.observability.sientia_monitoring import SientiaMonitoring
from laborious import metrics from laborious import metrics
# Requested session and secure channel lifetime (ms) before server revision; 10 minutes.
OPC_UA_SESSION_AND_CHANNEL_TIMEOUT_MS = 10 * 60 * 1000
class OpcClientAlreadyExistsError(RuntimeError):
"""Raised when _create_client is called while self.client is already set."""
class OpcSessionAlreadyConnectedError(RuntimeError):
"""Raised when _open_session is called while a UA session is already open."""
class OpcClientNotInitializedError(RuntimeError):
"""Raised when _open_session is called before _create_client."""
RECONNECTABLE_OPC_BAD_NAMES: frozenset[str] = frozenset(
{
'BadSessionIdInvalid',
'BadSessionClosed',
'BadSessionNotActivated',
'BadSecureChannelIdInvalid',
'BadSecureChannelClosed',
'BadSecureChannelTokenUnknown',
'BadTcpSecureChannelUnknown',
'BadServerNotConnected',
'BadConnectionClosed',
'BadDisconnect',
'BadConnectionRejected',
'BadCommunicationError',
'BadRequestInterrupted',
'BadUnknownResponse',
'BadTimeout',
'BadRequestTimeout',
'BadSequenceNumberInvalid',
'BadSequenceNumberUnknown',
'BadSecurityModeInsufficient',
'BadRequestHeaderInvalid',
'BadInvalidState',
}
)
def _opc_authentication_token_str(client: Client | None) -> str:
"""
Serialize the current OPC UA authentication token (session handle) for logging and metrics.
Return:
str: Token string, or "unknown" if unavailable.
"""
if client is None:
return 'unknown'
try:
proto = client.uaclient.protocol
if proto is None:
return 'unknown'
tok = getattr(proto, 'authentication_token', None)
if tok is None:
return 'unknown'
return str(tok)
except Exception:
return 'unknown'
def _opc_status_from_exception(exc: BaseException) -> str:
"""
Resolve OPC UA status name from an exception, including chained UaStatusCodeError causes.
Args:
exc (BaseException): Raised error from asyncua.
Return:
str: Status class name or generic Python exception name.
"""
current: BaseException | None = exc
while current is not None:
if isinstance(current, UaStatusCodeError):
return type(current).__name__
current = current.__cause__
return type(exc).__name__
def is_reconnectable_opcua_bad(exc: BaseException) -> bool:
"""
Return whether the exception is a Tier-1 OPC UA Bad* that should trigger reconnect.
Args:
exc (BaseException): Raised error from get_node or write_value.
Return:
bool: True if reconnect should be scheduled.
"""
return _opc_status_from_exception(exc) in RECONNECTABLE_OPC_BAD_NAMES
def _model_labels_from_write_metadata(metadata: dict[str, Any] | None) -> dict[str, str]:
"""
Extract model_id and model_name from write metadata for Prometheus labels.
Args:
metadata (dict[str, Any] | None): Context passed into write_data; may omit keys.
Return:
dict[str, str]: Labels model_id and model_name, defaulting to "unknown".
"""
if not metadata:
return {'model_id': 'unknown', 'model_name': 'unknown'}
return {
'model_id': str(metadata.get('model_id', 'unknown')),
'model_name': str(metadata.get('model_name', 'unknown')),
}
data_type_map = { data_type_map = {
'float': { 'float': {
'converter': float, 'converter': float,
@@ -63,8 +177,6 @@ class OpcRepository(SientiaMonitoring):
self.cert_path = cert_path self.cert_path = cert_path
self.private_key_path = private_key_path self.private_key_path = private_key_path
self.server_cert_path = server_cert_path self.server_cert_path = server_cert_path
self.logger = logger
self.error_count = 0
self.reconnection_interval = reconnection_interval self.reconnection_interval = reconnection_interval
self.last_reconnection_time: None | datetime = None self.last_reconnection_time: None | datetime = None
self.disconnection_interval = 10.0 self.disconnection_interval = 10.0
@@ -79,27 +191,79 @@ class OpcRepository(SientiaMonitoring):
'workflow_name': 'opc_repository', 'workflow_name': 'opc_repository',
'schedule_name': '-', 'schedule_name': '-',
} }
self._last_write_mono: float | None = None
self._connection_lock = asyncio.Lock()
self._session_ready = asyncio.Event()
self._reconnect_task: asyncio.Task[None] | None = None
self._allow_reconnect = True
async def set_security(self): def _opc_debug_tags(self, session_id: str) -> dict[str, str]:
""" """
Configures the security settings for the OPC UA client. Build Prometheus/log label tags for OPC session-scoped metrics.
This method sets up the security policy, certificates, and timeouts
required for establishing a secure connection with the OPC UA server. Args:
session_id (str): OPC UA session token string.
Return:
dict[str, str]: Labels pod_id, server_name, runtime, opc_server_id, session_id.
"""
return {
'pod_id': str(getattr(self, 'pod_id', 'unknown')),
'server_name': self.server_name,
'runtime': str(getattr(self, 'runtime', 'unknown')),
'opc_server_id': self.id,
'session_id': session_id,
}
def _is_session_open(self) -> bool:
"""
Return whether the asyncua client has an open transport session.
Return:
bool: True when protocol exists and is not closed.
"""
if self.client is None:
return False
try:
proto = self.client.uaclient.protocol
return proto is not None and proto.state != 'closed'
except Exception:
return False
def _reconnection_window_elapsed(self) -> bool:
"""
Return whether enough time has passed since the last reconnect attempt.
Return:
bool: True if a new reconnect is allowed.
"""
if self.last_reconnection_time is None:
return True
return (
datetime.now() - self.last_reconnection_time
).total_seconds() > self.reconnection_interval
def _not_connected_error(self) -> dict[str, Any]:
"""
Build the standard error payload when validate_connection finds no open protocol.
Return:
dict[str, Any]: Notification fields for OPC_CONNECTION_NOT_READY.
"""
return {
'notification_id': f'OPC_CONNECTION_NOT_READY_{self.id}',
'message': f'OPC server {self.id} is not connected',
'block': 'opc_repository',
'level': NotificationLevel.WARNING,
}
async def set_security(self) -> None:
"""
Configure certificates and timeouts on the asyncua client.
Raises: Raises:
ValueError: If either the certificate path or private key path is not provided. ValueError: If cert paths or client are missing.
Attributes:
- cert_path (str): Path to the client's certificate file.
- private_key_path (str): Path to the client's private key file.
- server_cert_path (str, optional): Path to the server's certificate file.
- server_uri (str): The URI of the server to be used as the application URI.
- client (opcua.Client): The OPC UA client instance.
- logger (logging.Logger): Logger instance for logging information.
Security Settings:
- Security Policy: Basic256
- Secure Channel Timeout: 10,000,000 ms
- Session Timeout: 10,000,000 ms
""" """
if self.cert_path is None or self.private_key_path is None: if self.cert_path is None or self.private_key_path is None:
raise ValueError( raise ValueError(
'Certificate and private key paths must be provided for secure connection.' 'Certificate and private key paths must be provided for secure connection.'
@@ -113,91 +277,107 @@ class OpcRepository(SientiaMonitoring):
raise ValueError('Client must be initialized before setting security') raise ValueError('Client must be initialized before setting security')
self.client.application_uri = self.server_uri self.client.application_uri = self.server_uri
self.logger.custom_info('Setting security...', self.metadata) self.info('Setting security...', self.metadata)
await self.client.set_security( await self.client.set_security(
SecurityPolicyBasic256, SecurityPolicyBasic256,
certificate=str(cert), certificate=str(cert),
private_key=str(private_key), private_key=str(private_key),
server_certificate=str(server_cert) if server_cert else None, server_certificate=str(server_cert) if server_cert else None,
) )
self.client.secure_channel_timeout = 10000000 self.client.secure_channel_timeout = OPC_UA_SESSION_AND_CHANNEL_TIMEOUT_MS
self.client.session_timeout = 10000000 self.client.session_timeout = OPC_UA_SESSION_AND_CHANNEL_TIMEOUT_MS
async def connect(self) -> tuple[bool, dict[str, Any]]: async def _create_client(self) -> None:
""" """
Establishes a connection to the OPC server. Instantiate the asyncua Client and apply security when configured.
This method initializes the OPC client using the provided URL and
sets up security if a certificate path is specified. It then Caller must hold _connection_lock. Does not open a UA session.
attempts to connect to the server and logs the connection status.
Raises: Raises:
Exception: If the connection to the OPC server fails. OpcClientAlreadyExistsError: If self.client is already set.
""" """
if self.client is not None:
raise OpcClientAlreadyExistsError(
f'OPC client already exists for server {self.id}; '
'call disconnect() before creating a new client'
)
self.client = Client(self.url, timeout=10, watchdog_intervall=3600000) # type: ignore[attr-defined] self.client = Client(self.url, timeout=10, watchdog_intervall=50) # type: ignore[attr-defined]
self.client.name = self.pod_id self.client.name = self.pod_id
self.client.application_name = self.pod_id self.client.application_name = self.pod_id
pod_uri = self.pod_id.replace('-', ':') pod_uri = self.pod_id.replace('-', ':')
self.client.application_uri = pod_uri self.client.application_uri = pod_uri
self.client.product_uri = pod_uri self.client.product_uri = pod_uri
if self.cert_path: if self.cert_path:
await self.set_security() await self.set_security()
self.logger.custom_info(
f'Starting connection to OPC server {self.id}:{self.server_name}...', self.metadata
)
return await self.try_connect()
async def try_connect(self) -> tuple[bool, dict[str, Any]]: async def _open_session(self) -> tuple[bool, dict[str, Any]]:
""" """
Attempt to establish connection to the OPC server. Open the OPC UA session on the existing client.
This method performs the actual connection attempt to the OPC server Caller must hold _connection_lock.
and handles connection failures with comprehensive error reporting.
It updates reconnection timing and provides detailed error information
for operational monitoring and debugging.
Returns: Raises:
tuple[bool, dict[str, Any]]: Connection result OpcClientNotInitializedError: If self.client is None.
- bool: True if connection successful, False otherwise OpcSessionAlreadyConnectedError: If a session is already open.
- dict: Error information if connection failed
Return:
tuple[bool, dict[str, Any]]: Success flag and error payload on connect failure.
""" """
if self.client is None:
raise OpcClientNotInitializedError(
f'OPC client is not initialized for server {self.id}; '
'call _create_client() before opening a session'
)
if self._is_session_open():
raise OpcSessionAlreadyConnectedError(
f'OPC session already connected for server {self.id}; '
'call disconnect() before connecting again'
)
tags = { tags = {
'pod_id': self.pod_id, 'pod_id': self.pod_id,
'server_name': self.server_name, 'server_name': self.server_name,
} }
await self.emit_metric(metrics.OPC_CONNECTIONS_TOTAL, tags) await self.emit_metric(metrics.OPC_CONNECTIONS_TOTAL, tags)
try: try:
self.last_reconnection_time = datetime.now()
if self.client is None:
return False, {
'notification_id': f'OPC_CONNECTION_ERROR_{self.id}',
'message': 'Client is not initialized',
'block': 'opc_repository',
'level': NotificationLevel.ERROR,
}
await self.client.connect() await self.client.connect()
session_id = _opc_authentication_token_str(self.client)
revised_session_timeout_ms = int(self.client.session_timeout)
revised_secure_channel_timeout_ms = int(self.client.secure_channel_timeout)
self.info(
f'OPC new session connected opc_server_id={self.id} session_id={session_id} '
f'revised_session_timeout_ms={revised_session_timeout_ms} '
f'revised_secure_channel_timeout_ms={revised_secure_channel_timeout_ms}',
self.metadata,
)
await self.emit_metric(
metrics.OPC_SESSION_CREATED_TOTAL, self._opc_debug_tags(session_id)
)
await self.emit_metric(
metric_object=metrics.OPC_SESSION_REVISED_TIMEOUT_MS,
method='set',
tags=self._opc_debug_tags(session_id),
value=revised_session_timeout_ms,
)
await self.emit_metric( await self.emit_metric(
metric_object=metrics.OPC_CONNECTION_STATUS, metric_object=metrics.OPC_CONNECTION_STATUS,
method='set', method='set',
tags={ tags={**tags, 'server_url': self.url},
**tags,
'server_url': self.url,
},
value=1, value=1,
) )
self._last_write_mono = None
self._session_ready.set()
return True, {} return True, {}
except Exception as e: except Exception as e:
await self.disconnect() await self._disconnect_locked()
trace = traceback.format_exc() trace = traceback.format_exc()
self.logger.custom_error(trace, self.metadata) self.error(trace, self.metadata)
await self.emit_metric(metrics.OPC_CONNECTIONS_FAILED, tags) await self.emit_metric(metrics.OPC_CONNECTIONS_FAILED, tags)
return False, { return False, {
'notification_id': f'OPC_CONNECTION_ERROR_{self.id}', 'notification_id': f'OPC_CONNECTION_ERROR_{self.id}',
'message': f'Failed to connect to OPC server: {e}', 'message': f'Failed to connect to OPC server: {e}',
@@ -206,21 +386,45 @@ class OpcRepository(SientiaMonitoring):
'attachment_content': trace, 'attachment_content': trace,
} }
async def disconnection_fallback(self) -> list: async def _connect_locked(self) -> tuple[bool, dict[str, Any]]:
"""
Tries 5 times to disconnect from the OPC UA server, with a delay of 100ms x try.
""" """
Create the client when absent, then open a UA session.
Caller must hold _connection_lock.
Raises:
OpcSessionAlreadyConnectedError: If a session is already open.
Return:
tuple[bool, dict[str, Any]]: Result from _open_session on connect failure.
"""
if self._is_session_open():
raise OpcSessionAlreadyConnectedError(
f'OPC session already connected for server {self.id}; '
'call disconnect() before connecting again'
)
if self.client is None:
await self._create_client()
return await self._open_session()
async def _disconnection_fallback(self) -> list[dict[str, Any]]:
"""
Try up to five times to disconnect from the OPC UA server.
"""
assert self.client is not None assert self.client is not None
error_stack = [] error_stack: list[dict[str, Any]] = []
for i in range(5): for i in range(5):
try: try:
self.logger.info(f'Disconnecting from OPC UA server, attempt {i + 1} of 5') self.info(
f'Disconnecting from OPC UA server, attempt {i + 1} of 5',
self.metadata,
)
await self.client.disconnect() await self.client.disconnect()
return [] return []
except Exception as e: except Exception as e:
self.logger.error( self.error(
f'Failed to disconnect from OPC UA server in attempt {i + 1} of 5: {e}' f'Failed to disconnect from OPC UA server in attempt {i + 1} of 5: {e}',
self.metadata,
) )
error_stack.append( error_stack.append(
{ {
@@ -232,18 +436,26 @@ class OpcRepository(SientiaMonitoring):
await asyncio.sleep(self.disconnection_interval * i) await asyncio.sleep(self.disconnection_interval * i)
return error_stack return error_stack
async def disconnect(self): async def _disconnect_locked(self) -> None:
""" """
Gracefully disconnect from the OPC server. Tear down the current session and client.
This method safely terminates the connection to the OPC server Caller must hold _connection_lock.
and cleans up client resources. It handles disconnection errors
gracefully and ensures proper resource cleanup.
""" """
self._last_write_mono = None
self._session_ready.clear()
if self.client is None: if self.client is None:
return return
errors = await self.disconnection_fallback() session_id = _opc_authentication_token_str(self.client)
self.info(
f'OPC disconnecting opc_server_id={self.id} session_id={session_id}',
self.metadata,
)
await self.emit_metric(metrics.OPC_SESSION_CLOSED_TOTAL, self._opc_debug_tags(session_id))
errors = await self._disconnection_fallback()
if errors: if errors:
await self.send_notification_async( await self.send_notification_async(
metadata=self.metadata, metadata=self.metadata,
@@ -254,7 +466,8 @@ class OpcRepository(SientiaMonitoring):
attachment_content=json.dumps(errors, indent=4), attachment_content=json.dumps(errors, indent=4),
) )
else: else:
self.logger.warning(f'Disconnected from OPC server {self.id} successfully') self.warning(f'Disconnected from OPC server {self.id} successfully', self.metadata)
await self.emit_metric( await self.emit_metric(
metric_object=metrics.OPC_CONNECTION_STATUS, metric_object=metrics.OPC_CONNECTION_STATUS,
method='set', method='set',
@@ -265,186 +478,388 @@ class OpcRepository(SientiaMonitoring):
}, },
value=0, value=0,
) )
self.client = None self.client = None
async def _reconnect_locked(self) -> tuple[bool, dict[str, Any]]:
"""
Close the current session and open a new one.
Caller must hold _connection_lock. Records last_reconnection_time for interval gating.
Return:
tuple[bool, dict[str, Any]]: Result from _connect_locked after teardown.
"""
self.last_reconnection_time = datetime.now()
await self._disconnect_locked()
return await self._connect_locked()
async def connect(self) -> tuple[bool, dict[str, Any]]:
"""
Open an OPC UA session under the connection lock (worker initialization).
"""
async with self._connection_lock:
self.info(
f'Starting connection to OPC server {self.id}:{self.server_name}...',
self.metadata,
)
return await self._connect_locked()
async def disconnect(self) -> None:
"""
Gracefully disconnect from the OPC server under the connection lock.
Disables background reconnect so late writes during worker shutdown do not
respawn sessions.
"""
async with self._connection_lock:
self._allow_reconnect = False
await self._disconnect_locked()
async def validate_connection(self) -> tuple[bool, dict[str, Any]]: async def validate_connection(self) -> tuple[bool, dict[str, Any]]:
""" """
Validate and maintain OPC server connection health. Read-only check that the asyncua protocol is open.
This method performs comprehensive connection validation and Caller must ensure _session_ready before writing. Does not connect or reconnect.
implements automatic reconnection logic for production reliability.
It handles various connection states and implements intelligent
reconnection strategies with error counting and timing controls.
Connection Validation: Return:
1. Checks client existence and connection state tuple[bool, dict[str, Any]]: (True, {}) when open, otherwise (False, error).
2. Implements error counting with automatic disconnection """
3. Enforces reconnection timing windows if self._is_session_open():
4. Provides detailed error reporting and notifications return True, {}
self.error(f'OPC server {self.id} is not connected', self.metadata)
return False, self._not_connected_error()
Reconnection Strategy: def _reconnect_task_in_progress(self) -> bool:
- Error Count Threshold: Disconnects after 5 consecutive errors """
- Reconnection Window: Enforces minimum intervals between attempts Return whether a background reconnect task is currently running.
- Automatic Recovery: Attempts reconnection when conditions allow
- State Monitoring: Continuously monitors connection health Return:
bool: True when a reconnect task exists and has not finished.
"""
return self._reconnect_task is not None and not self._reconnect_task.done()
async def _start_reconnect(self, reason: str, session_id: str) -> None:
"""
Schedule a background reconnect when allowed by interval and task state.
Clears _session_ready before starting the task. No-op when _allow_reconnect is
False, the reconnection window has not elapsed, or a reconnect is already running.
Args: Args:
None reason (str): Trigger for reconnect (OPC status name or synthetic reason).
session_id (str): Session token before failure.
Returns:
tuple[bool, dict[str, Any]]: Connection validation result
- bool: True if connection is healthy, False otherwise
- dict: Error information if validation fails
""" """
if self.client is None: if not self._allow_reconnect:
return await self.connect() return
if not self._reconnection_window_elapsed():
self.warning(
f'OPC reconnect skipped reason=reconnection_window opc_server_id={self.id} '
f'reconnect_reason={reason}',
self.metadata,
)
return
if self._reconnect_task_in_progress():
self.warning(
f'OPC reconnect skipped reason=in_progress opc_server_id={self.id} '
f'reconnect_reason={reason}',
self.metadata,
)
return
# if self.error_count > 5: # NOSONAR self._session_ready.clear()
# self.logger.custom_warning( self.info(
# f'OPC server {self.id} will be disconnected due to multiple errors', self.metadata f'OPC reconnect scheduled reconnect_reason={reason} opc_server_id={self.id} '
# ) f'old_session_id={session_id}',
# try: self.metadata,
# await self.disconnect() )
# except Exception as e: self._reconnect_task = asyncio.create_task(self._run_reconnect(reason, session_id))
# trace = traceback.format_exc()
# self.logger.custom_error(
# f'Failed to disconnect from OPC server: {e}', self.metadata
# )
# self.logger.custom_error(trace, self.metadata)
# self.logger.custom_info(
# f'Attempting to reconnect to OPC server {self.id}...', self.metadata
# )
# return await self.connect()
# Check if client is connected using asyncua's connection state async def _run_reconnect(self, reason: str, session_id: str) -> None:
"""
Background task that tears down and re-establishes the OPC UA session.
Args:
reason (str): Trigger for reconnect (OPC status or ProtocolClosed).
session_id (str): Previous session token string for logging.
"""
try: try:
if ( async with self._connection_lock:
self.client.uaclient.protocol is None self.info(
or self.client.uaclient.protocol.state == 'closed' f'OPC reconnect started reconnect_reason={reason} opc_server_id={self.id} '
): f'old_session_id={session_id}',
# OPC server is not connected self.metadata,
self.logger.custom_error(f'OPC server {self.id} is not connected', self.metadata) )
if ( success, error = await self._reconnect_locked()
self.last_reconnection_time is None if not success:
or (datetime.now() - self.last_reconnection_time).total_seconds() self.error(
> self.reconnection_interval f'OPC reconnect failed reconnect_reason={reason} opc_server_id={self.id}',
): self.metadata,
await self.disconnect() )
self.logger.custom_info( if error:
f'Trying to reconnect to OPC server {self.id}...', self.metadata self.error(error.get('message', ''), self.metadata)
except Exception:
self.error(
f'OPC reconnect task failed opc_server_id={self.id} reconnect_reason={reason}',
self.metadata,
)
self.error(traceback.format_exc(), self.metadata)
async def _log_write_inter_arrival(self, session_id: str, node: str) -> None:
"""
Log elapsed wall time since the previous successful OPC write on this repository.
Args:
session_id (str): Current OPC UA session token string.
node (str): Node id written in this operation.
"""
now = time.monotonic()
if self._last_write_mono is not None:
delta_s = now - self._last_write_mono
self.info(
f'OPC write inter-arrival_s={delta_s:.6f} opc_server_id={self.id} '
f'session_id={session_id} node={node}',
self.metadata,
)
if self.client is not None:
session_timeout_ms = float(self.client.session_timeout)
if session_timeout_ms > 0 and delta_s > (session_timeout_ms / 1000.0):
await self.emit_metric(
metrics.OPC_WRITE_INTER_ARRIVAL_OVER_SESSION_TIMEOUT_TOTAL,
self._opc_debug_tags(session_id),
) )
return await self.connect() self._last_write_mono = now
return False, { async def _emit_opc_write_metric(
'notification_id': f'OPC_CONNECTION_AWAITING_RECONNECTION_WINDOW_{self.id}', self, session_id: str, result: str, metadata: dict[str, Any] | None
'message': f'OPC server {self.id} is not connected, waiting for next reconnection window...', ) -> None:
'block': 'opc_repository', """
'level': NotificationLevel.WARNING, Emit opc_write_attempts_total for a single write attempt outcome.
}
return True, {}
except Exception as e:
trace = traceback.format_exc()
message = f'Failed to validate connection to OPC server: {e}'
self.logger.custom_error(message, self.metadata)
return False, {
'notification_id': f'OPC_CONNECTION_CHECK_ERROR_{self.id}',
'message': message,
'block': 'opc_repository',
'level': NotificationLevel.ERROR,
'attachment_content': trace,
}
async def write_data( Args:
self, node: str, value: Any, data_type: str, logger: Logger, metadata: dict[str, Any] session_id (str): OPC UA session token string, or "unknown".
result (str): Outcome label (OK, OPC status name, ProtocolClosed, etc.).
metadata (dict[str, Any] | None): Write context for model_id/model_name labels.
"""
await self.emit_metric(
metrics.OPC_WRITE_ATTEMPTS_TOTAL,
{
**self._opc_debug_tags(session_id),
**_model_labels_from_write_metadata(metadata),
'result': result,
},
)
def _write_failure_payload(
self,
notification_id: str,
message: str,
level: NotificationLevel = NotificationLevel.ERROR,
attachment_content: str | None = None,
opc_error_kind: str | None = None,
opc_status: str | None = None,
) -> dict[str, Any]:
"""
Build a structured error dict returned from failed write_data paths.
Args:
notification_id (str): Stable notification identifier.
message (str): Human-readable failure message.
level (NotificationLevel): Severity for downstream notifications.
attachment_content (str | None): Optional traceback or diagnostic text.
opc_error_kind (str | None): Classifier (session_bad, connection_lost, etc.).
opc_status (str | None): OPC UA status name or synthetic reason.
Return:
dict[str, Any]: Error payload consumed by the OPC activity layer.
"""
payload: dict[str, Any] = {
'notification_id': notification_id,
'message': message,
'block': 'opc_repository',
'level': level,
}
if attachment_content is not None:
payload['attachment_content'] = attachment_content
if opc_error_kind is not None:
payload['opc_error_kind'] = opc_error_kind
if opc_status is not None:
payload['opc_status'] = opc_status
return payload
async def _handle_tier1_bad(
self,
exc: BaseException,
session_id: str,
node: str,
metadata: dict[str, Any],
phase: str,
) -> tuple[bool, dict[str, Any]]: ) -> tuple[bool, dict[str, Any]]:
""" """
Write data to OPC server with comprehensive validation and monitoring. Record metrics/logs and schedule reconnect after a Tier-1 Bad* error.
This method provides secure and reliable data writing to OPC servers
with automatic connection validation, data type conversion, and
comprehensive error handling. It implements performance monitoring
and metrics collection for operational visibility.
Data Writing Process:
1. Connection validation and automatic reconnection
2. Node validation and error handling
3. Data type conversion and validation
4. OPC data writing with timestamp
5. Performance metrics collection
6. Error handling and notification
Args: Args:
node (str): OPC node identifier to write data to exc (BaseException): Tier-1 OPC UA error.
value (Any): Data value to write to the OPC node session_id (str): Session token at failure time.
data_type (str): Data type for OPC conversion node (str): Node id being written.
logger (Logger): Logger instance for operation logging metadata (dict[str, Any]): Write context.
metadata (dict[str, Any]): Context metadata for logging and metrics phase (str): get_node or write_value.
Returns: Return:
tuple[bool, dict[str, Any]]: Write operation result tuple[bool, dict[str, Any]]: Always (False, error payload).
- bool: True if write successful, False otherwise
- dict: Error information if write failed
""" """
opc_status = _opc_status_from_exception(exc)
trace = traceback.format_exc()
self.error(trace, metadata)
await self._emit_opc_write_metric(session_id, opc_status, metadata)
self.error(
f'OPC write failed opc_status={opc_status} opc_server_id={self.id} '
f'session_id={session_id} model_id={metadata.get("model_id", "unknown")} '
f'model_name={metadata.get("model_name", "unknown")} node={node} phase={phase}',
metadata,
)
await self._start_reconnect(opc_status, session_id)
return False, self._write_failure_payload(
notification_id=f'OPC_WRITE_DATA_ERROR_{self.id}',
message=f'Failed to {phase} on OPC server: {exc} | metadata: {metadata}',
attachment_content=trace,
opc_error_kind='session_bad',
opc_status=opc_status,
)
is_connected, error = await self.validate_connection() async def _write_reconnect_in_progress(
self, metadata: dict[str, Any]
) -> tuple[bool, dict[str, Any]]:
"""
Fail a write because a background reconnect task is already running.
Args:
metadata (dict[str, Any]): Write context passed through to the activity.
Return:
tuple[bool, dict[str, Any]]: (False, error info with opc_error_kind reconnect_in_progress).
"""
await self._emit_opc_write_metric('unknown', 'ReconnectInProgress', metadata)
self.warning(
f'OPC write rejected reconnect_in_progress opc_server_id={self.id} '
f'model_id={metadata.get("model_id", "unknown")} '
f'model_name={metadata.get("model_name", "unknown")}',
metadata,
)
return False, {
'notification_id': f'OPC_WRITE_RECONNECT_IN_PROGRESS_{self.id}',
'message': f'OPC write skipped: reconnect in progress | metadata: {metadata}',
'block': 'opc_repository',
'level': NotificationLevel.WARNING,
'opc_error_kind': 'reconnect_in_progress',
}
async def _write_connection_lost(
self, metadata: dict[str, Any], opc_status: str
) -> tuple[bool, dict[str, Any]]:
"""
Fail a write after scheduling reconnect for a closed or stale session.
Args:
metadata (dict[str, Any]): Write context passed through to the activity.
opc_status (str): Synthetic reason (ProtocolClosed, SessionNotReady).
Return:
tuple[bool, dict[str, Any]]: (False, error info with opc_error_kind connection_lost).
"""
await self._emit_opc_write_metric('unknown', opc_status, metadata)
return False, self._write_failure_payload(
notification_id=f'OPC_WRITE_CONNECTION_LOST_{self.id}',
message=f'OPC write skipped: connection lost ({opc_status}) | metadata: {metadata}',
level=NotificationLevel.WARNING,
opc_error_kind='connection_lost',
opc_status=opc_status,
)
async def write_data(
self, node: str, value: Any, data_type: str, metadata: dict[str, Any]
) -> tuple[bool, dict[str, Any]]:
"""
Write data to OPC server with a single attempt and background reconnect scheduling.
Reconnect is scheduled on Tier-1 Bad*, closed protocol, or stale session readiness.
There is no retry within the same call.
Args:
node (str): OPC UA node id to write.
value (Any): Value to convert and send.
data_type (str): Logical type key (float, int, bool, str, double).
metadata (dict[str, Any]): Activity context (model_id, model_name, etc.).
Return:
tuple[bool, dict[str, Any]]: (True, {response_time}) on success, or
(False, structured error info) on failure.
"""
if self._reconnect_task_in_progress():
return await self._write_reconnect_in_progress(metadata)
if not self._session_ready.is_set():
session_id = _opc_authentication_token_str(self.client)
await self._start_reconnect('SessionNotReady', session_id)
if self._reconnect_task_in_progress():
return await self._write_reconnect_in_progress(metadata)
return await self._write_connection_lost(metadata, 'SessionNotReady')
is_connected, _error = await self.validate_connection()
if not is_connected: if not is_connected:
return False, error session_id = _opc_authentication_token_str(self.client)
await self._start_reconnect('ProtocolClosed', session_id)
return await self._write_connection_lost(metadata, 'ProtocolClosed')
start_time = time.time() start_time = time.time()
session_id = _opc_authentication_token_str(self.client)
try: try:
# ignored because self.validate_connection is called before, so we know self.client is not None
node_obj = self.client.get_node(node) # type: ignore[union-attr] node_obj = self.client.get_node(node) # type: ignore[union-attr]
except Exception as e: except Exception as e:
if is_reconnectable_opcua_bad(e):
return await self._handle_tier1_bad(e, session_id, node, metadata, 'get_node')
trace = traceback.format_exc() trace = traceback.format_exc()
logger.custom_error(trace, metadata.get('schedule_name', 'N/A')) self.error(trace, metadata)
self.error_count += 1 await self._emit_opc_write_metric(
return False, { session_id, f'GetNodeError:{type(e).__name__}', metadata
'notification_id': f'OPC_WRITE_GET_NODE_ERROR_{self.id}', )
'message': f'Failed to get node from OPC server: {e} | metadata: {metadata}', return False, self._write_failure_payload(
'block': 'opc_repository', notification_id=f'OPC_WRITE_GET_NODE_ERROR_{self.id}',
'level': NotificationLevel.ERROR, message=f'Failed to get node from OPC server: {e} | metadata: {metadata}',
'attachment_content': trace, attachment_content=trace,
} )
if data_type not in data_type_map: if data_type not in data_type_map:
return False, { await self._emit_opc_write_metric(session_id, 'UnsupportedDataType', metadata)
'notification_id': f'OPC_WRITE_DATA_TYPE_ERROR_{self.id}', return False, self._write_failure_payload(
'message': f'Unsupported data type: {data_type} | metadata: {metadata}', notification_id=f'OPC_WRITE_DATA_TYPE_ERROR_{self.id}',
'block': 'opc_repository', message=f'Unsupported data type: {data_type} | metadata: {metadata}',
'level': NotificationLevel.ERROR, )
}
data = data_type_map[data_type]['converter'](value) data = data_type_map[data_type]['converter'](value)
logger.custom_info(f'Writing {data} - {type(data)} to {node}', metadata) self.info(f'Writing {data} - {type(data)} to {node}', metadata)
# now = datetime.now() # NOSONAR
ua_data = DataValue( ua_data = DataValue(
Variant(data, data_type_map[data_type]['opc_type']), Variant(data, data_type_map[data_type]['opc_type']),
# SourceTimestamp=DateTime( # NOSONAR
# now.year, now.month, now.day, now.hour, now.minute, now.second, now.microsecond # NOSONAR
# ), # NOSONAR
) )
try: try:
await node_obj.write_value(ua_data) await node_obj.write_value(ua_data)
end_time = time.time() end_time = time.time()
response_time = end_time - start_time response_time = end_time - start_time
except Exception as e: except Exception as e:
if is_reconnectable_opcua_bad(e):
return await self._handle_tier1_bad(e, session_id, node, metadata, 'write_value')
trace = traceback.format_exc() trace = traceback.format_exc()
logger.custom_error(trace, metadata) self.error(trace, metadata)
self.error_count += 1 await self._emit_opc_write_metric(session_id, type(e).__name__, metadata)
return False, { return False, self._write_failure_payload(
'notification_id': f'OPC_WRITE_DATA_ERROR_{self.id}', notification_id=f'OPC_WRITE_DATA_ERROR_{self.id}',
'message': f'Failed to write data to OPC server: {e} | metadata: {metadata}', message=f'Failed to write data to OPC server: {e} | metadata: {metadata}',
'block': 'opc_repository', attachment_content=trace,
'level': NotificationLevel.ERROR, )
'attachment_content': trace,
} await self._emit_opc_write_metric(session_id, 'OK', metadata)
self.error_count = 0 await self._log_write_inter_arrival(session_id, node)
return True, { return True, {
'response_time': response_time, 'response_time': response_time,

View File

@@ -1,73 +0,0 @@
import os
import re
from collections.abc import Sequence
from typing import Any
from sientia_do.observability.logger import Logger
from temporalio.client import Client
from temporalio.worker import PollerBehaviorAutoscaling, Worker
parameters = [
('MAX_CONCURRENT_WORKFLOW_TASKS', '200'),
('MAX_CONCURRENT_ACTIVITIES', '200'),
('MAX_CONCURRENT_LOCAL_ACTIVITIES', '200'),
('MAX_CACHED_WORKFLOWS', '200'),
('WORKFLOW_POLLER_BEHAVIOUR_MINIMUM', '10'),
('WORKFLOW_POLLER_BEHAVIOUR_INITIAL', '100'),
('WORKFLOW_POLLER_BEHAVIOUR_MAXIMUM', '200'),
('ACTIVITY_POLLER_BEHAVIOUR_MINIMUM', '10'),
('ACTIVITY_POLLER_BEHAVIOUR_INITIAL', '100'),
('ACTIVITY_POLLER_BEHAVIOUR_MAXIMUM', '200'),
]
def camel_to_snake(text: str) -> str:
"""Convert camelCase or PascalCase to snake_case."""
text = re.sub('(.)([A-Z][a-z]+)', r'\1_\2', text)
text = re.sub('([a-z0-9])([A-Z])', r'\1_\2', text)
return text.lower()
def prepare_worker(
main_workflow: type,
other_workflows: Sequence[type],
activities: Sequence[Any],
temporal_client: Client,
logger: Logger,
) -> Worker:
main_workflow_name = main_workflow.__name__.upper()
queue_name = f'{camel_to_snake(main_workflow.__name__)}-queue'
local_workflow_parameters = {}
for parameter in parameters:
local_workflow_parameters[parameter[0]] = int(
os.getenv(main_workflow_name + '_' + parameter[0], parameter[1])
)
logger.info(f'Preparing worker for {main_workflow_name} with queue {queue_name}')
logger.info(f'Worker runtime config: {local_workflow_parameters}')
return Worker(
temporal_client,
task_queue=queue_name,
workflows=[main_workflow, *other_workflows],
activities=[*activities],
max_concurrent_workflow_tasks=local_workflow_parameters['MAX_CONCURRENT_WORKFLOW_TASKS'],
max_concurrent_activities=local_workflow_parameters['MAX_CONCURRENT_ACTIVITIES'],
max_concurrent_local_activities=local_workflow_parameters[
'MAX_CONCURRENT_LOCAL_ACTIVITIES'
],
max_cached_workflows=local_workflow_parameters['MAX_CACHED_WORKFLOWS'],
workflow_task_poller_behavior=PollerBehaviorAutoscaling(
minimum=local_workflow_parameters['WORKFLOW_POLLER_BEHAVIOUR_MINIMUM'],
initial=local_workflow_parameters['WORKFLOW_POLLER_BEHAVIOUR_INITIAL'],
maximum=local_workflow_parameters['WORKFLOW_POLLER_BEHAVIOUR_MAXIMUM'],
),
activity_task_poller_behavior=PollerBehaviorAutoscaling(
minimum=local_workflow_parameters['ACTIVITY_POLLER_BEHAVIOUR_MINIMUM'],
initial=local_workflow_parameters['ACTIVITY_POLLER_BEHAVIOUR_INITIAL'],
maximum=local_workflow_parameters['ACTIVITY_POLLER_BEHAVIOUR_MAXIMUM'],
),
)

View File

@@ -5,12 +5,14 @@ This module provides the main worker implementation for the Sientia DataOps Labo
It orchestrates Temporal workers, manages task queues, and handles the lifecycle of It orchestrates Temporal workers, manages task queues, and handles the lifecycle of
prediction and retraining workflows. prediction and retraining workflows.
The worker supports multiple task queues: The worker supports multiple runtime-scoped task queues (via ``sientia_do.temporal.worker.prepare_worker``):
- predictions_batch-queue: Handles batch prediction workflows (heavy workload) - predictions_batch-{runtime}-queue: Batch prediction workflows (heavy workload)
Includes activities for MLFlow, data quality gates, OPC export, PI Web API export, and PostgreSQL - minimal_retrain-{runtime}-queue: Model retraining workflows
- minimal_retrain-queue: Handles model retraining workflows - drift-{runtime}-queue: Drift detection workflows
- drift-queue: Handles drift detection workflows - simple_metrics-{runtime}-queue: Simple metrics workflows
- simple_metrics-queue: Handles simple metrics calculation workflows
``RUNTIME`` must be set; it is passed to every ``prepare_worker`` call. Schedulers must use the
same queue names (breaking change vs legacy ``drift-queue`` / ``simple_metrics-queue``).
Key Features: Key Features:
- Resource-based scaling with WorkerTuner (CPU and memory aware) - Resource-based scaling with WorkerTuner (CPU and memory aware)
@@ -21,6 +23,7 @@ Key Features:
- Multiple worker instances for different workflow types - Multiple worker instances for different workflow types
Environment Variables: Environment Variables:
- RUNTIME: Required non-empty string; suffix for all task queue names
- TEMPORAL_HOST: Temporal server address (default: localhost:7233) - TEMPORAL_HOST: Temporal server address (default: localhost:7233)
- TEMPORAL_NAMESPACE: Temporal namespace (default: laborious) - TEMPORAL_NAMESPACE: Temporal namespace (default: laborious)
- POD_ID: Kubernetes pod identifier for metrics - POD_ID: Kubernetes pod identifier for metrics
@@ -40,6 +43,7 @@ with workflow.unsafe.imports_passed_through():
from prometheus_client import start_http_server from prometheus_client import start_http_server
from sientia_do.notifications.handlers import CoreNotificationHandler as NotificationHandler from sientia_do.notifications.handlers import CoreNotificationHandler as NotificationHandler
from sientia_do.observability.logger import get_logger from sientia_do.observability.logger import get_logger
from sientia_do.temporal.worker.prepare_worker import prepare_worker
from sientia_do.utils.connectors_config import ( from sientia_do.utils.connectors_config import (
build_api_config, build_api_config,
build_mongodb_config, build_mongodb_config,
@@ -53,7 +57,6 @@ with workflow.unsafe.imports_passed_through():
build_mlflow_config, build_mlflow_config,
build_opc_config, build_opc_config,
) )
from laborious.worker.prepare_worker import prepare_worker
from laborious.workflows.drift import Drift from laborious.workflows.drift import Drift
from laborious.workflows.minimal_retrain import MinimalRetrain from laborious.workflows.minimal_retrain import MinimalRetrain
from laborious.workflows.predictions_batch import PredictionsBatch from laborious.workflows.predictions_batch import PredictionsBatch
@@ -63,7 +66,7 @@ with workflow.unsafe.imports_passed_through():
) )
from laborious.workflows.sub_workflows.prediction_process import PredictionProcess from laborious.workflows.sub_workflows.prediction_process import PredictionProcess
POD_ID = os.getenv('POD_ID') POD_ID = os.getenv('HOSTNAME')
SDK_METRICS_PORT = int(os.getenv('HTTP_SDK_METRICS_PORT', '9091')) SDK_METRICS_PORT = int(os.getenv('HTTP_SDK_METRICS_PORT', '9091'))
@@ -100,10 +103,21 @@ async def main():
logger.custom_info(f'Starting Worker with POD_ID: {POD_ID}', metadata) logger.custom_info(f'Starting Worker with POD_ID: {POD_ID}', metadata)
logger.custom_info('Starting prometheus client...', metadata) runtime = os.getenv('RUNTIME', '').strip()
if not runtime:
logger.custom_critical(
'RUNTIME environment variable is required and must be non-empty',
metadata,
)
metrics.APP_UP.labels(pod_id=POD_ID).set(0)
sys.exit(1)
metadata_runtime = {**metadata, 'runtime': runtime}
logger.custom_info('Starting prometheus client...', metadata_runtime)
start_prometheus_server() start_prometheus_server()
logger.custom_info('Starting Notification Handler...', metadata) logger.custom_info('Starting Notification Handler...', metadata_runtime)
mongo_config = build_mongodb_config() mongo_config = build_mongodb_config()
notification_handler = NotificationHandler( notification_handler = NotificationHandler(
@@ -113,7 +127,7 @@ async def main():
project_name=os.getenv('PROJECT_NAME', 'laborious'), project_name=os.getenv('PROJECT_NAME', 'laborious'),
) )
logger.custom_info('Starting Activities...', metadata) logger.custom_info('Starting Activities...', metadata_runtime)
activities = Activities( activities = Activities(
postgres_config=build_postgres_config(), postgres_config=build_postgres_config(),
@@ -125,10 +139,13 @@ async def main():
notification_handler=notification_handler, notification_handler=notification_handler,
) )
logger.custom_info('Initializing OPC...', metadata) logger.custom_info('Initializing OPC...', metadata_runtime)
await activities.init_opc() await activities.init_opc()
logger.custom_info(f'Starting SDK Metrics Server on port {SDK_METRICS_PORT}...', metadata) logger.custom_info(
f'Starting SDK Metrics Server on port {SDK_METRICS_PORT}...',
metadata_runtime,
)
new_runtime = Runtime( new_runtime = Runtime(
telemetry=TelemetryConfig( telemetry=TelemetryConfig(
@@ -136,7 +153,7 @@ async def main():
) )
) )
logger.custom_info(f'Starting Temporal Client at {host}...', metadata) logger.custom_info(f'Starting Temporal Client at {host}...', metadata_runtime)
temporal_client = await client.Client.connect( temporal_client = await client.Client.connect(
target_host=host, target_host=host,
@@ -144,7 +161,7 @@ async def main():
runtime=new_runtime, runtime=new_runtime,
) )
logger.custom_info('Starting Workers...', metadata) logger.custom_info(f'Starting Workers (runtime={runtime})...', metadata_runtime)
workers = [ workers = [
prepare_worker( prepare_worker(
@@ -159,6 +176,7 @@ async def main():
activities.export_data_to_postgres, activities.export_data_to_postgres,
], ],
logger=logger, logger=logger,
runtime=runtime,
), ),
prepare_worker( prepare_worker(
temporal_client=temporal_client, temporal_client=temporal_client,
@@ -170,6 +188,7 @@ async def main():
activities.export_data_to_postgres, activities.export_data_to_postgres,
], ],
logger=logger, logger=logger,
runtime=runtime,
), ),
prepare_worker( prepare_worker(
temporal_client=temporal_client, temporal_client=temporal_client,
@@ -182,6 +201,7 @@ async def main():
activities.export_data_to_postgres, activities.export_data_to_postgres,
], ],
logger=logger, logger=logger,
runtime=runtime,
), ),
prepare_worker( prepare_worker(
temporal_client=temporal_client, temporal_client=temporal_client,
@@ -205,11 +225,13 @@ async def main():
activities.cleanup_minio_objects_expired, activities.cleanup_minio_objects_expired,
activities.repeat_last_prediction, activities.repeat_last_prediction,
activities.export_data_to_postgres, activities.export_data_to_postgres,
activities.export_payload_to_postgres,
activities.write_metrics, activities.write_metrics,
# Pi Web API # Pi Web API
activities.write_pi_web_api_data, activities.write_pi_web_api_data,
], ],
logger=logger, logger=logger,
runtime=runtime,
), ),
] ]
@@ -217,7 +239,7 @@ async def main():
for w in workers: for w in workers:
handlers.append(w.run()) handlers.append(w.run())
logger.custom_info('Workers started successfully', metadata) logger.custom_info('Workers started successfully', metadata_runtime)
exit_code = 0 exit_code = 0
try: try:

View File

@@ -116,11 +116,18 @@ python_functions = ["test_*"]
addopts = [ addopts = [
"-v", "-v",
"--strict-markers", "--strict-markers",
# pytest>=9.1 has a known bug where its unraisableexception plugin crashes
# (tracemalloc partially-initialized AttributeError) when 2+ unraisable
# exceptions land close together — e.g. "coroutine was never awaited" from
# AsyncMock-mocked sync methods (metrics_controller, minio_repository) being
# GC'd. Harmless mock artifacts turned into a hard ERROR by the plugin itself.
"-p", "no:unraisableexception",
] ]
markers = [ markers = [
"asyncio: marks tests as async", "asyncio: marks tests as async",
"integration: marks tests as integration tests", "integration: marks tests as integration tests",
"unit: marks tests as unit tests", "unit: marks tests as unit tests",
"opc: marks tests that use the in-process OPC UA server (OpcRepository E2E)",
] ]
[tool.coverage.run] [tool.coverage.run]

View File

@@ -18,3 +18,4 @@ testcontainers[postgres,minio] # PostgreSQL and MinIO containers for E2E tests
# Development Tools # Development Tools
ipython>=8.12.0 # Enhanced Python shell ipython>=8.12.0 # Enhanced Python shell
ipdb>=0.13.13 # IPython debugger ipdb>=0.13.13 # IPython debugger
ipykernel==6.30.1 # IPython kernel for Jupyter notebooks

View File

@@ -1,12 +1,18 @@
temporalio temporalio
psycopg2-binary psycopg2-binary
sqlalchemy sqlalchemy
asyncua asyncua==1.0.6
redis redis
git+ssh://git@github.com/Aignosi/sientia-dataops-library.git@1.10.4 sientia_do>=1.12.2
mlflow
prometheus-client prometheus-client
botocore botocore
boto3 boto3
s3fs s3fs
pyarrow pyarrow
mlflow kaleido
hyperopt
shap
pycurl
scipy<1.14.0
scikit-learn==1.5.2

18
requirements-local.txt Normal file
View File

@@ -0,0 +1,18 @@
temporalio
psycopg2-binary
sqlalchemy
asyncua==1.0.6
redis
git+ssh://git@github.com/Aignosi/sientia-dataops-library.git@1.12.2
git+ssh://git@github.com/Aignosi/sientia-model-library.git@0.10.0
prometheus-client
botocore
boto3
s3fs
pyarrow
kaleido
hyperopt
shap
pycurl
scipy<1.14.0
scikit-learn==1.5.2

View File

@@ -1,10 +1,10 @@
temporalio temporalio
psycopg2-binary psycopg2-binary
sqlalchemy sqlalchemy
asyncua asyncua==1.0.6
redis redis
git+ssh://git@github.com/Aignosi/sientia-dataops-library.git@1.10.4 sientia_do>=1.12.2
git+ssh://git@github.com/Aignosi/sientia-mlops-library.git@0.41.0 sientia>0.40.0
prometheus-client prometheus-client
botocore botocore
boto3 boto3

View File

@@ -1,4 +1,4 @@
sonar.projectKey=Aignosi_sientia-dataops-laborious_temporal_beaec423-6c42-4f26-8134-b676287b499d sonar.projectKey=Aignosi_sientia-dataops-laborious_temporal_ca1a7039-6db9-49e5-be78-54d29bc93e4f
sonar.projectName=sientia-dataops-laborious_temporal sonar.projectName=sientia-dataops-laborious_temporal
sonar.sources=laborious sonar.sources=laborious
sonar.tests=tests sonar.tests=tests

View File

@@ -173,7 +173,7 @@ async def test_shutdown(
mock_mlflow_init, mock_mlflow_init,
mock_storage_init, mock_storage_init,
): ):
mock_opc_init.close = AsyncMock() mock_opc_init.aclose = AsyncMock()
postgres_config = { postgres_config = {
'host': 'localhost', 'host': 'localhost',
'port': 5432, 'port': 5432,
@@ -221,7 +221,7 @@ async def test_shutdown(
) )
await activities.shutdown() await activities.shutdown()
mock_opc_init.close.assert_called_once() mock_opc_init.aclose.assert_called_once()
mock_storage_init.close.assert_called_once() mock_storage_init.close.assert_called_once()
mock_mlflow_init.close.assert_called_once() mock_mlflow_init.close.assert_called_once()
mock_gates_init.close.assert_called_once() mock_gates_init.close.assert_called_once()

View File

@@ -5,7 +5,16 @@ from pandas import DataFrame
from pytest import mark from pytest import mark
from sientia_do.notifications.models import NotificationLevel from sientia_do.notifications.models import NotificationLevel
from laborious.activities.opc import OPC from laborious.activities.opc import (
OPC,
OPC_COMMENT_SEPARATOR,
OPC_RECONNECT_IN_PROGRESS_COMMENT,
OPC_SESSION_BAD_COMMENT_PREFIX,
OPC_SESSION_BAD_CONFIDENCE,
OPC_WRITTING_ERROR_CONFIDENCE,
OPC_WRITTING_ERROR_MESSAGE,
_apply_opc_write_error,
)
metadata = { metadata = {
'metadata': { 'metadata': {
@@ -206,7 +215,7 @@ WRITE_DATA_CASES = [
async def test_write_data_success(opc, tag, data_type, data): async def test_write_data_success(opc, tag, data_type, data):
opc.opc_repository['server1'].write_data.return_value = (True, {'response_time': 0.1}) opc.opc_repository['server1'].write_data.return_value = (True, {'response_time': 0.1})
result = await opc.write_data( response_time, error_info = await opc.write_data(
server_id='server1', server_id='server1',
tag=tag, tag=tag,
data=data, data=data,
@@ -214,10 +223,9 @@ async def test_write_data_success(opc, tag, data_type, data):
tag_type='prediction', tag_type='prediction',
metadata=metadata, metadata=metadata,
) )
assert result == 0.1 assert response_time == 0.1
opc.opc_repository['server1'].write_data.assert_called_once_with( assert error_info is None
tag, data, data_type, opc.logger, metadata opc.opc_repository['server1'].write_data.assert_called_once_with(tag, data, data_type, metadata)
)
@mark.asyncio @mark.asyncio
@@ -233,7 +241,7 @@ async def test_write_data_failed(opc):
}, },
) )
result = await opc.write_data( response_time, error_info = await opc.write_data(
server_id='server1', server_id='server1',
tag='tag1', tag='tag1',
data=50, data=50,
@@ -241,7 +249,8 @@ async def test_write_data_failed(opc):
tag_type='prediction', tag_type='prediction',
metadata=metadata, metadata=metadata,
) )
assert result is None assert response_time is None
assert error_info is not None
opc.send_notification_async.assert_called_once_with( opc.send_notification_async.assert_called_once_with(
metadata=metadata, metadata=metadata,
@@ -281,17 +290,211 @@ async def test_write_data_exception(opc):
raise AssertionError('Expected an exception to be raised') raise AssertionError('Expected an exception to be raised')
@mark.parametrize(
'error_info,initial_seen,initial_status,initial_reconnect,expected',
[
(None, False, None, False, (False, None, False)),
({}, False, None, False, (False, None, False)),
(
{'opc_error_kind': 'session_bad', 'opc_status': 'BadSessionIdInvalid'},
False,
None,
False,
(True, 'BadSessionIdInvalid', False),
),
(
{'opc_error_kind': 'session_bad', 'opc_status': 'NewStatus'},
True,
'OldStatus',
False,
(True, 'NewStatus', False),
),
(
{'opc_error_kind': 'session_bad'},
True,
'KeptStatus',
False,
(True, 'KeptStatus', False),
),
(
{'opc_error_kind': 'reconnect_in_progress'},
False,
None,
False,
(False, None, True),
),
(
{'opc_error_kind': 'other'},
True,
'Status',
True,
(True, 'Status', True),
),
],
)
def test_apply_opc_write_error(
error_info, initial_seen, initial_status, initial_reconnect, expected
):
result = _apply_opc_write_error(
error_info,
initial_seen,
initial_status,
initial_reconnect,
)
assert result == expected
@mark.asyncio
async def test_write_tags_from_config_prediction_success(opc):
opc.write_data = AsyncMock(return_value=(0.1, None))
data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]})
tags_config = {'tag1': {'data_type': 'float'}}
response_times, session_bad, opc_status, reconnect = await opc._write_tags_from_config(
server_id='server1',
tags_config=tags_config,
data=data,
data_column='prediction',
tag_type='prediction',
log_label='Prediction data',
metadata=metadata['metadata'],
)
assert response_times == {'tag1': 0.1}
assert session_bad is False
assert opc_status is None
assert reconnect is False
opc.write_data.assert_called_once_with(
server_id='server1',
tag='tag1',
data=0.75,
data_type='float',
tag_type='prediction',
metadata=metadata['metadata'],
)
@mark.asyncio
async def test_write_tags_from_config_confidence_success(opc):
opc.write_data = AsyncMock(return_value=(0.2, None))
data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]})
tags_config = {'tag2': {'data_type': 'float'}}
response_times, session_bad, opc_status, reconnect = await opc._write_tags_from_config(
server_id='server1',
tags_config=tags_config,
data=data,
data_column='prediction_confidence',
tag_type='confidence',
log_label='Confidence data',
metadata=metadata['metadata'],
)
assert response_times == {'tag2': 0.2}
assert session_bad is False
assert opc_status is None
assert reconnect is False
opc.write_data.assert_called_once_with(
server_id='server1',
tag='tag2',
data=0.95,
data_type='float',
tag_type='confidence',
metadata=metadata['metadata'],
)
@mark.asyncio
async def test_write_tags_from_config_write_failure(opc):
opc.write_data = AsyncMock(return_value=(None, {}))
data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]})
response_times, session_bad, opc_status, reconnect = await opc._write_tags_from_config(
server_id='server1',
tags_config={'tag1': {'data_type': 'float'}},
data=data,
data_column='prediction',
tag_type='prediction',
log_label='Prediction data',
metadata=metadata['metadata'],
)
assert response_times == {'tag1': None}
assert session_bad is False
assert opc_status is None
assert reconnect is False
@mark.asyncio
async def test_write_tags_from_config_session_bad(opc):
opc.write_data = AsyncMock(
return_value=(
None,
{
'opc_error_kind': 'session_bad',
'opc_status': 'BadSessionIdInvalid',
},
)
)
data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]})
response_times, session_bad, opc_status, reconnect = await opc._write_tags_from_config(
server_id='server1',
tags_config={'tag1': {'data_type': 'float'}},
data=data,
data_column='prediction',
tag_type='prediction',
log_label='Prediction data',
metadata=metadata['metadata'],
)
assert response_times == {'tag1': None}
assert session_bad is True
assert opc_status == 'BadSessionIdInvalid'
assert reconnect is False
@mark.asyncio
async def test_write_tags_from_config_reconnect_in_progress(opc):
opc.write_data = AsyncMock(
return_value=(
None,
{'opc_error_kind': 'reconnect_in_progress'},
)
)
data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]})
response_times, session_bad, opc_status, reconnect = await opc._write_tags_from_config(
server_id='server1',
tags_config={'tag1': {'data_type': 'float'}},
data=data,
data_column='prediction',
tag_type='prediction',
log_label='Prediction data',
metadata=metadata['metadata'],
)
assert response_times == {'tag1': None}
assert session_bad is False
assert opc_status is None
assert reconnect is True
@mark.asyncio @mark.asyncio
async def test_manage_output_tags_success(opc): async def test_manage_output_tags_success(opc):
opc.write_data = AsyncMock(return_value=0.1) opc._write_tags_from_config = AsyncMock(
side_effect=[
({'tag1': 0.1}, False, None, False),
({'tag2': 0.1}, False, None, False),
]
)
data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]}) data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]})
config = { config = {
'prediction_tags': {'tag1': {'data_type': 'float'}}, 'prediction_tags': {'tag1': {'data_type': 'float'}},
'confidence_tags': {'tag2': {'data_type': 'float'}}, 'confidence_tags': {'tag2': {'data_type': 'float'}},
} }
output_data, opc_metrics = await opc.manage_output_tags( output_data, opc_metrics, session_bad, opc_status, reconnect = await opc.manage_output_tags(
server_id='server1', server_id='server1',
config=config, config=config,
data=data, data=data,
@@ -300,83 +503,53 @@ async def test_manage_output_tags_success(opc):
assert output_data is True assert output_data is True
assert opc_metrics == {'tag1': 0.1, 'tag2': 0.1} assert opc_metrics == {'tag1': 0.1, 'tag2': 0.1}
opc.write_data.assert_has_calls( assert session_bad is False
[ assert opc_status is None
call( assert reconnect is False
server_id='server1', assert opc._write_tags_from_config.await_count == 2
tag='tag1',
data=0.75,
data_type='float',
tag_type='prediction',
metadata=metadata['metadata'],
),
call(
server_id='server1',
tag='tag2',
data=0.95,
data_type='float',
tag_type='confidence',
metadata=metadata['metadata'],
),
]
)
@mark.asyncio @mark.asyncio
@mark.parametrize('side_effect', [[0.1, None], [None, 0.2]]) async def test_manage_output_tags_failed(opc):
async def test_manage_output_tags_failed(opc, side_effect): opc._write_tags_from_config = AsyncMock(
opc.write_data = AsyncMock(side_effect=side_effect) side_effect=[
({'tag1': 0.1}, False, None, False),
({'tag2': None}, False, None, False),
]
)
data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]}) data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]})
config = { config = {
'prediction_tags': {'tag1': {'data_type': 'float'}}, 'prediction_tags': {'tag1': {'data_type': 'float'}},
'confidence_tags': {'tag2': {'data_type': 'float'}}, 'confidence_tags': {'tag2': {'data_type': 'float'}},
} }
output_data, opc_metrics = await opc.manage_output_tags(
output_data, opc_metrics, _, _, _ = await opc.manage_output_tags(
server_id='server1', server_id='server1',
config=config, config=config,
data=data, data=data,
metadata=metadata['metadata'], metadata=metadata['metadata'],
) )
assert output_data is False assert output_data is False
assert opc_metrics == {'tag1': side_effect[0], 'tag2': side_effect[1]} assert opc_metrics == {'tag1': 0.1, 'tag2': None}
opc.write_data.assert_has_calls(
[
call(
server_id='server1',
tag='tag1',
data=0.75,
data_type='float',
tag_type='prediction',
metadata=metadata['metadata'],
),
call(
server_id='server1',
tag='tag2',
data=0.95,
data_type='float',
tag_type='confidence',
metadata=metadata['metadata'],
),
]
)
@mark.asyncio @mark.asyncio
async def test_manage_output_tags_do_nothing(opc): async def test_manage_output_tags_do_nothing(opc):
opc.write_data = AsyncMock(return_value=0.1) opc._write_tags_from_config = AsyncMock()
data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]}) data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]})
config = { config = {'_invalid_key': {'tag1': {'data_type': 'float'}}}
'_invalid_key': {'tag1': {'data_type': 'float'}},
} output_data, opc_metrics, _, _, _ = await opc.manage_output_tags(
output_data, opc_metrics = await opc.manage_output_tags(
server_id='server1', server_id='server1',
config=config, config=config,
data=data, data=data,
metadata=metadata['metadata'], metadata=metadata['metadata'],
) )
assert output_data is True assert output_data is True
assert opc_metrics == {} assert opc_metrics == {}
opc.write_data.assert_not_called() opc._write_tags_from_config.assert_not_called()
@mark.asyncio @mark.asyncio
@@ -395,7 +568,9 @@ async def test_write_opc_data_success(mock_dataframe, opc):
} }
# Act # Act
opc.manage_output_tags = AsyncMock(return_value=(True, {'tag1': 0.1, 'tag2': 0.2})) opc.manage_output_tags = AsyncMock(
return_value=(True, {'tag1': 0.1, 'tag2': 0.2}, False, None, False)
)
opc.process_confidence = MagicMock(return_value={'data': 'data'}) opc.process_confidence = MagicMock(return_value={'data': 'data'})
output_data, opc_metrics = await opc.write_opc_data(input_data) output_data, opc_metrics = await opc.write_opc_data(input_data)
@@ -413,6 +588,9 @@ async def test_write_opc_data_success(mock_dataframe, opc):
mock_dataframe.return_value, mock_dataframe.return_value,
True, True,
metadata['metadata'], metadata['metadata'],
session_bad=False,
opc_status=None,
reconnect_in_progress=False,
) )
@@ -462,13 +640,91 @@ async def test_write_opc_data_no_validate_server(opc):
], ],
) )
def test_process_confidence(opc, data, success, expected): def test_process_confidence(opc, data, success, expected):
# Act result = opc.process_confidence(data, success, metadata['metadata'])
result = opc.process_confidence(data, success, metadata)
# Assert
assert result['prediction_confidence'][0] == expected assert result['prediction_confidence'][0] == expected
def test_process_confidence_session_bad(opc):
data = DataFrame({'prediction_confidence': [0.9]})
result = opc.process_confidence(
data,
False,
metadata['metadata'],
session_bad=True,
opc_status='BadSessionIdInvalid',
)
assert result['prediction_confidence'][0] == OPC_SESSION_BAD_CONFIDENCE
assert result['comments'][0].startswith(OPC_SESSION_BAD_COMMENT_PREFIX)
assert 'BadSessionIdInvalid' in result['comments'][0]
def test_process_confidence_generic_failure(opc):
data = DataFrame({'prediction_confidence': [0.9]})
result = opc.process_confidence(data, False, metadata['metadata'])
assert result['prediction_confidence'][0] == OPC_WRITTING_ERROR_CONFIDENCE
assert result['comments'][0] == OPC_WRITTING_ERROR_MESSAGE
@mark.asyncio
async def test_manage_output_tags_merges_error_flags(opc):
opc._write_tags_from_config = AsyncMock(
side_effect=[
({'tag1': None}, True, 'BadSessionIdInvalid', False),
({'tag2': 0.2}, False, None, True),
]
)
data = DataFrame({'prediction': [0.75], 'prediction_confidence': [0.95]})
config = {
'prediction_tags': {'tag1': {'data_type': 'float'}},
'confidence_tags': {'tag2': {'data_type': 'float'}},
}
(
success,
metrics,
session_bad_seen,
opc_status,
reconnect_in_progress,
) = await opc.manage_output_tags('server1', config, data, metadata['metadata'])
assert success is False
assert session_bad_seen is True
assert reconnect_in_progress is True
assert opc_status == 'BadSessionIdInvalid'
assert metrics == {'tag1': None, 'tag2': 0.2}
def test_process_confidence_reconnect_in_progress(opc):
data = DataFrame({'prediction_confidence': [0.9]})
result = opc.process_confidence(
data,
False,
metadata['metadata'],
reconnect_in_progress=True,
)
assert result['prediction_confidence'][0] == OPC_SESSION_BAD_CONFIDENCE
assert result['comments'][0] == OPC_RECONNECT_IN_PROGRESS_COMMENT
def test_process_confidence_concatenates_multiple_comments(opc):
data = DataFrame({'prediction_confidence': [0.9]})
session_comment = f'{OPC_SESSION_BAD_COMMENT_PREFIX} BadSessionIdInvalid'
result = opc.process_confidence(
data,
False,
metadata['metadata'],
session_bad=True,
opc_status='BadSessionIdInvalid',
reconnect_in_progress=True,
)
assert result['prediction_confidence'][0] == OPC_SESSION_BAD_CONFIDENCE
assert result['comments'][0] == OPC_COMMENT_SEPARATOR.join(
[session_comment, OPC_RECONNECT_IN_PROGRESS_COMMENT]
)
@mark.asyncio @mark.asyncio
async def test_validate_server(opc): async def test_validate_server(opc):
assert await opc.validate_server('server1', metadata) is True assert await opc.validate_server('server1', metadata) is True
@@ -478,5 +734,5 @@ async def test_validate_server(opc):
@mark.asyncio @mark.asyncio
async def test_close(opc): async def test_close(opc):
opc.opc_repository['server1'].disconnect = AsyncMock(return_value=True) opc.opc_repository['server1'].disconnect = AsyncMock(return_value=True)
await opc.close() await opc.aclose()
opc.opc_repository['server1'].disconnect.assert_called_once() opc.opc_repository['server1'].disconnect.assert_called_once()

View File

@@ -142,6 +142,41 @@ def test_get_model_run_id_success(mlflow_repository):
assert output == '1' assert output == '1'
def test_get_model_run_id_missing_source(mlflow_repository):
mlflow_repository.client.search_registered_models.return_value = [MagicMock(name='test')]
mlflow_repository.client.search_model_versions.return_value = [
MagicMock(current_stage='Production', version='1', source='runs/test/0'),
MagicMock(current_stage='Production', version='2', source=None),
]
with pytest.raises(mlflow_lib.exceptions.MlflowException) as exc_info:
mlflow_repository.get_model_run_id('test')
assert (
str(exc_info.value)
== "Model 'test' version '2' in stage 'Production' has no source URI to resolve run ID."
)
def test_get_model_run_id_invalid_source(mlflow_repository):
mlflow_repository.client.search_registered_models.return_value = [MagicMock(name='test')]
mlflow_repository.client.search_model_versions.return_value = [
MagicMock(current_stage='Production', version='1', source='runs/test/0'),
MagicMock(current_stage='Production', version='2', source='runs/test'),
]
with pytest.raises(mlflow_lib.exceptions.MlflowException) as exc_info:
mlflow_repository.get_model_run_id('test')
assert (
str(exc_info.value)
== "Model 'test' version '2' in stage 'Production' has invalid source URI "
"'runs/test' for run ID resolution."
)
def test_get_next_run_name(mlflow, mlflow_repository): def test_get_next_run_name(mlflow, mlflow_repository):
mlflow.search_runs.return_value = [1, 2, 3] mlflow.search_runs.return_value = [1, 2, 3]
output = mlflow_repository.get_next_run_name('run') output = mlflow_repository.get_next_run_name('run')

View File

@@ -1,12 +1,20 @@
import asyncio
import json import json
from datetime import datetime from datetime import datetime
from unittest.mock import ANY, AsyncMock, MagicMock, Mock, patch from unittest.mock import ANY, AsyncMock, MagicMock, Mock, patch
import pytest import pytest
from asyncua.crypto.security_policies import SecurityPolicyBasic256 from asyncua.crypto.security_policies import SecurityPolicyBasic256
from asyncua.ua.uaerrors import BadNodeIdUnknown, BadSessionIdInvalid
from sientia_do.notifications.models import NotificationLevel from sientia_do.notifications.models import NotificationLevel
from laborious.utils.repository.opc_repository import OpcRepository from laborious.utils.repository.opc_repository import (
OpcClientAlreadyExistsError,
OpcClientNotInitializedError,
OpcRepository,
OpcSessionAlreadyConnectedError,
is_reconnectable_opcua_bad,
)
@pytest.fixture @pytest.fixture
@@ -33,6 +41,11 @@ def opc_repository(mock_logger):
repository.send_notification = MagicMock() repository.send_notification = MagicMock()
repository.send_notification_async = AsyncMock() repository.send_notification_async = AsyncMock()
repository.emit_metric = AsyncMock() repository.emit_metric = AsyncMock()
repository.info = MagicMock()
repository.error = MagicMock()
repository.warning = MagicMock()
repository.debug = MagicMock()
repository._session_ready.set()
return repository return repository
@@ -65,7 +78,6 @@ def test_init(opc_repository):
assert opc_repository.reconnection_interval == 60 assert opc_repository.reconnection_interval == 60
assert opc_repository.client is None assert opc_repository.client is None
assert opc_repository.last_reconnection_time is None assert opc_repository.last_reconnection_time is None
assert opc_repository.error_count == 0
@pytest.mark.asyncio @pytest.mark.asyncio
@@ -80,8 +92,8 @@ async def test_set_security(opc_repository, mock_client):
private_key='/path/to/key.pem', private_key='/path/to/key.pem',
server_certificate='/path/to/server_cert.pem', server_certificate='/path/to/server_cert.pem',
) )
assert mock_client.secure_channel_timeout == 10000000 assert mock_client.secure_channel_timeout == 600_000
assert mock_client.session_timeout == 10000000 assert mock_client.session_timeout == 600_000
@pytest.mark.asyncio @pytest.mark.asyncio
@@ -106,48 +118,95 @@ async def test_set_security_missing_client(opc_repository):
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_connect_with_security(opc_repository, mock_client): async def test_connect_with_security(opc_repository, mock_client):
opc_repository.try_connect = AsyncMock(return_value=(True, {})) opc_repository._create_client = AsyncMock()
opc_repository._open_session = AsyncMock(return_value=(True, {}))
result = await opc_repository.connect() result = await opc_repository.connect()
opc_repository.try_connect.assert_called_once() opc_repository._create_client.assert_called_once()
assert opc_repository.client == mock_client opc_repository._open_session.assert_called_once()
assert result == (True, {}) assert result == (True, {})
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_connect_without_security(opc_repository, mock_client): async def test_connect_without_security(opc_repository, mock_client):
opc_repository.cert_path = None opc_repository.cert_path = None
opc_repository.try_connect = AsyncMock(return_value=(True, {})) opc_repository._create_client = AsyncMock()
opc_repository._open_session = AsyncMock(return_value=(True, {}))
opc_repository.set_security = AsyncMock() opc_repository.set_security = AsyncMock()
result = await opc_repository.connect() result = await opc_repository.connect()
opc_repository.try_connect.assert_called_once() opc_repository._create_client.assert_called_once()
opc_repository._open_session.assert_called_once()
opc_repository.set_security.assert_not_called() opc_repository.set_security.assert_not_called()
assert opc_repository.client == mock_client
assert result == (True, {}) assert result == (True, {})
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_try_connect_success(opc_repository): async def test_connect_raises_when_session_already_open(opc_repository, mock_client):
opc_repository.last_reconnection_time = None opc_repository.client = mock_client
proto = MagicMock()
proto.state = 'open'
mock_client.uaclient = MagicMock(protocol=proto)
with pytest.raises(OpcSessionAlreadyConnectedError, match='disconnect'):
await opc_repository.connect()
@pytest.mark.asyncio
async def test_create_client_raises_when_client_exists(opc_repository, mock_client):
opc_repository.client = mock_client
with pytest.raises(OpcClientAlreadyExistsError, match='already exists'):
await opc_repository._create_client()
@pytest.mark.asyncio
async def test_open_session_success(opc_repository):
closed_proto = MagicMock()
closed_proto.state = 'closed'
opc_repository.client = AsyncMock() opc_repository.client = AsyncMock()
result = await opc_repository.try_connect() opc_repository.client.uaclient = MagicMock(protocol=closed_proto)
opc_repository.client.session_timeout = 600_000
opc_repository.client.secure_channel_timeout = 600_000
open_proto = MagicMock()
open_proto.state = 'open'
open_proto.authentication_token = 'tok'
async def connect_side_effect():
opc_repository.client.uaclient.protocol = open_proto
opc_repository.client.connect = AsyncMock(side_effect=connect_side_effect)
result = await opc_repository._open_session()
opc_repository.client.connect.assert_called_once() opc_repository.client.connect.assert_called_once()
assert opc_repository.last_reconnection_time is not None assert opc_repository.last_reconnection_time is None
assert result == (True, {}) assert result == (True, {})
assert opc_repository._session_ready.is_set()
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_try_connect_fail(opc_repository): async def test_open_session_raises_when_already_connected(opc_repository, mock_client):
opc_repository.last_reconnection_time = None opc_repository.client = mock_client
opc_repository.disconnect = AsyncMock() proto = MagicMock()
proto.state = 'open'
mock_client.uaclient = MagicMock(protocol=proto)
with pytest.raises(OpcSessionAlreadyConnectedError, match='disconnect'):
await opc_repository._open_session()
@pytest.mark.asyncio
async def test_open_session_fail(opc_repository):
opc_repository._disconnect_locked = AsyncMock()
opc_repository.client = MagicMock() opc_repository.client = MagicMock()
opc_repository.client.connect.side_effect = Exception('Test error') opc_repository.client.uaclient = MagicMock(protocol=MagicMock(state='closed'))
opc_repository.client.connect = AsyncMock(side_effect=Exception('Test error'))
is_connected, error_data = await opc_repository.try_connect() is_connected, error_data = await opc_repository._open_session()
opc_repository.disconnect.assert_called_once() opc_repository._disconnect_locked.assert_called_once()
opc_repository.client.connect.assert_called_once() opc_repository.client.connect.assert_called_once()
assert is_connected is False assert is_connected is False
assert error_data['notification_id'] == f'OPC_CONNECTION_ERROR_{opc_repository.id}' assert error_data['notification_id'] == f'OPC_CONNECTION_ERROR_{opc_repository.id}'
@@ -158,25 +217,18 @@ async def test_try_connect_fail(opc_repository):
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_try_connect_no_client(opc_repository): async def test_open_session_raises_when_no_client(opc_repository):
opc_repository.client = None opc_repository.client = None
result = await opc_repository.try_connect()
assert result == ( with pytest.raises(OpcClientNotInitializedError, match='not initialized'):
False, await opc_repository._open_session()
{
'notification_id': f'OPC_CONNECTION_ERROR_{opc_repository.id}',
'message': 'Client is not initialized',
'block': 'opc_repository',
'level': NotificationLevel.ERROR,
},
)
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_disconnection_fallback_success(opc_repository, mock_client): async def test_disconnection_fallback_success(opc_repository, mock_client):
opc_repository.client = mock_client opc_repository.client = mock_client
mock_client.disconnect.return_value = True mock_client.disconnect.return_value = True
result = await opc_repository.disconnection_fallback() result = await opc_repository._disconnection_fallback()
mock_client.disconnect.assert_called_once() mock_client.disconnect.assert_called_once()
assert result == [] assert result == []
@@ -186,7 +238,7 @@ async def test_disconnection_fallback_success(opc_repository, mock_client):
async def test_disconnection_fallback_fail(opc_repository, mock_client): async def test_disconnection_fallback_fail(opc_repository, mock_client):
opc_repository.client = mock_client opc_repository.client = mock_client
mock_client.disconnect.side_effect = Exception('Test error') mock_client.disconnect.side_effect = Exception('Test error')
result = await opc_repository.disconnection_fallback() result = await opc_repository._disconnection_fallback()
assert result == [ assert result == [
{'attempt': 1, 'error': 'Test error', 'traceback': ANY}, {'attempt': 1, 'error': 'Test error', 'traceback': ANY},
{'attempt': 2, 'error': 'Test error', 'traceback': ANY}, {'attempt': 2, 'error': 'Test error', 'traceback': ANY},
@@ -200,11 +252,12 @@ async def test_disconnection_fallback_fail(opc_repository, mock_client):
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_disconnect(opc_repository, mock_client): async def test_disconnect(opc_repository, mock_client):
opc_repository.client = mock_client opc_repository.client = mock_client
opc_repository.disconnection_fallback = AsyncMock(return_value=[]) opc_repository._disconnection_fallback = AsyncMock(return_value=[])
await opc_repository.disconnect() await opc_repository.disconnect()
opc_repository.disconnection_fallback.assert_called_once() opc_repository._disconnection_fallback.assert_called_once()
assert opc_repository.client is None assert opc_repository.client is None
assert opc_repository._allow_reconnect is False
@pytest.mark.asyncio @pytest.mark.asyncio
@@ -216,12 +269,12 @@ async def test_disconnect_no_client(opc_repository):
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_disconnect_error(opc_repository, mock_client): async def test_disconnect_error(opc_repository, mock_client):
opc_repository.client = mock_client opc_repository.client = mock_client
opc_repository.disconnection_fallback = AsyncMock( opc_repository._disconnection_fallback = AsyncMock(
return_value=[{'attempt': 1, 'error': 'Test error', 'traceback': 'text'}] return_value=[{'attempt': 1, 'error': 'Test error', 'traceback': 'text'}]
) )
await opc_repository.disconnect() await opc_repository.disconnect()
opc_repository.disconnection_fallback.assert_called_once() opc_repository._disconnection_fallback.assert_called_once()
opc_repository.send_notification_async.assert_called_once_with( opc_repository.send_notification_async.assert_called_once_with(
metadata=opc_repository.metadata, metadata=opc_repository.metadata,
notification_id=f'OPC_DISCONNECTION_ERROR_{opc_repository.id}', notification_id=f'OPC_DISCONNECTION_ERROR_{opc_repository.id}',
@@ -238,91 +291,24 @@ async def test_disconnect_error(opc_repository, mock_client):
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_validate_connection_none_client(opc_repository): async def test_validate_connection_none_client(opc_repository):
opc_repository.client = None opc_repository.client = None
opc_repository.connect = AsyncMock(return_value=(True, {}))
response = await opc_repository.validate_connection() response = await opc_repository.validate_connection()
assert response == (True, {}) assert response == (False, opc_repository._not_connected_error())
opc_repository.connect.assert_called_once()
# @pytest.mark.asyncio
# async def test_validate_connection_error_count_disconnect_error(opc_repository):
# opc_repository.error_count = 6
# opc_repository.client = AsyncMock()
# opc_repository.disconnect = AsyncMock(side_effect=Exception('Test error'))
# opc_repository.connect = AsyncMock(return_value=(True, {}))
# response = await opc_repository.validate_connection()
# assert response == opc_repository.connect.return_value
# opc_repository.disconnect.assert_called_once()
# opc_repository.connect.assert_called_once()
# opc_repository.logger.custom_error.assert_has_calls(
# [
# call('Failed to disconnect from OPC server: Test error', ANY),
# ]
# )
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_validate_connection_error_validate_connection_error(opc_repository): async def test_validate_connection_session_not_open(opc_repository):
opc_repository.client = MagicMock(uaclient=Exception('Test error'))
opc_repository.error_count = 0
response = await opc_repository.validate_connection()
assert response == (
False,
{
'notification_id': f'OPC_CONNECTION_CHECK_ERROR_{opc_repository.id}',
'message': "Failed to validate connection to OPC server: 'Exception' object has no attribute 'protocol'",
'block': 'opc_repository',
'level': NotificationLevel.ERROR,
'attachment_content': ANY,
},
)
@pytest.mark.asyncio
@patch('laborious.utils.repository.opc_repository.datetime')
async def test_validate_connection_lost_not_time_to_reconnect(_mock_datetime, opc_repository):
_mock_datetime.now = MagicMock(return_value=datetime(2025, 1, 1, 0, 0, 0))
opc_repository.error_count = 0
opc_repository.client = MagicMock() opc_repository.client = MagicMock()
opc_repository.client.uaclient.protocol = None opc_repository.client.uaclient.protocol = None
opc_repository.last_reconnection_time = datetime(2025, 1, 1, 0, 0, 0)
opc_repository.connect = MagicMock(return_value=(True, {}))
response = await opc_repository.validate_connection() response = await opc_repository.validate_connection()
opc_repository.connect.assert_not_called()
assert response == (
False,
{
'notification_id': f'OPC_CONNECTION_AWAITING_RECONNECTION_WINDOW_{opc_repository.id}',
'message': f'OPC server {opc_repository.id} is not connected, waiting for next reconnection window...',
'block': 'opc_repository',
'level': NotificationLevel.WARNING,
},
)
assert response == (False, opc_repository._not_connected_error())
@pytest.mark.asyncio opc_repository.error.assert_called_once()
@patch('laborious.utils.repository.opc_repository.datetime')
async def test_validate_connection_lost_time_to_reconnect(mock_datetime, opc_repository):
mock_datetime.now = MagicMock(return_value=datetime(2025, 1, 1, 1, 0, 0))
opc_repository.error_count = 0
opc_repository.client = AsyncMock()
opc_repository.client.uaclient.protocol = None
opc_repository.last_reconnection_time = datetime(2025, 1, 1, 0, 0, 0)
opc_repository.connect = AsyncMock(return_value=(True, {}))
response = await opc_repository.validate_connection()
opc_repository.connect.assert_called_once()
assert response == opc_repository.connect.return_value
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_validate_connection_success(opc_repository): async def test_validate_connection_success(opc_repository):
opc_repository.client = MagicMock() opc_repository.client = MagicMock()
opc_repository.error_count = 0
opc_repository.client.uaclient.protocol = MagicMock() opc_repository.client.uaclient.protocol = MagicMock()
opc_repository.client.uaclient.protocol.state = 'open' opc_repository.client.uaclient.protocol.state = 'open'
@@ -337,9 +323,7 @@ async def test_write_data_validate_connection_do_nothing(opc_repository):
mock_node = AsyncMock() mock_node = AsyncMock()
opc_repository.client.get_node.return_value = mock_node opc_repository.client.get_node.return_value = mock_node
result = await opc_repository.write_data( result = await opc_repository.write_data('ns=2;s=TestNode', 42.0, 'float', metadata['metadata'])
'ns=2;s=TestNode', 42.0, 'float', opc_repository.logger, metadata['metadata']
)
opc_repository.validate_connection.assert_called_once() opc_repository.validate_connection.assert_called_once()
opc_repository.client.get_node.assert_called_once_with('ns=2;s=TestNode') opc_repository.client.get_node.assert_called_once_with('ns=2;s=TestNode')
@@ -348,28 +332,29 @@ async def test_write_data_validate_connection_do_nothing(opc_repository):
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_write_data_validate_connection_failed(opc_repository): async def test_write_data_validate_connection_failed(opc_repository):
opc_repository.validate_connection = AsyncMock(return_value=(False, {})) opc_repository.client = MagicMock()
opc_repository.client = AsyncMock() opc_repository.client.uaclient.protocol = MagicMock(state='closed')
opc_repository.error_count = 0 opc_repository._start_reconnect = AsyncMock()
result = await opc_repository.write_data( is_success, error_data = await opc_repository.write_data(
'ns=2;s=TestNode', 42.0, 'float', opc_repository.logger, metadata['metadata'] 'ns=2;s=TestNode', 42.0, 'float', metadata['metadata']
) )
opc_repository.validate_connection.assert_called_once() opc_repository._start_reconnect.assert_called_once()
opc_repository.client.get_node.assert_not_called() assert opc_repository._start_reconnect.call_args.args[0] == 'ProtocolClosed'
assert result == (False, {}) assert is_success is False
assert error_data['opc_error_kind'] == 'connection_lost'
assert error_data['opc_status'] == 'ProtocolClosed'
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_write_data_get_node_failed(opc_repository): async def test_write_data_get_node_failed(opc_repository):
opc_repository.validate_connection = AsyncMock(return_value=(True, {})) opc_repository.validate_connection = AsyncMock(return_value=(True, {}))
opc_repository.client = AsyncMock() opc_repository.client = AsyncMock()
opc_repository.error_count = 0
opc_repository.client.get_node = MagicMock(side_effect=Exception('Test error')) opc_repository.client.get_node = MagicMock(side_effect=Exception('Test error'))
is_success, error_data = await opc_repository.write_data( is_success, error_data = await opc_repository.write_data(
'ns=2;s=TestNode', 42.0, 'float', opc_repository.logger, metadata['metadata'] 'ns=2;s=TestNode', 42.0, 'float', metadata['metadata']
) )
opc_repository.validate_connection.assert_called_once() opc_repository.validate_connection.assert_called_once()
@@ -393,7 +378,7 @@ async def test_write_data_invalid_data_type(opc_repository, mock_client):
mock_client.get_node = MagicMock(return_value=mock_node) mock_client.get_node = MagicMock(return_value=mock_node)
is_success, error_data = await opc_repository.write_data( is_success, error_data = await opc_repository.write_data(
'ns=2;s=TestNode', 42.0, 'invalid_type', opc_repository.logger, metadata['metadata'] 'ns=2;s=TestNode', 42.0, 'invalid_type', metadata['metadata']
) )
opc_repository.validate_connection.assert_called_once() opc_repository.validate_connection.assert_called_once()
@@ -417,9 +402,7 @@ async def test_write_data(opc_repository, mock_client):
mock_node = AsyncMock() mock_node = AsyncMock()
mock_client.get_node = MagicMock(return_value=mock_node) mock_client.get_node = MagicMock(return_value=mock_node)
result = await opc_repository.write_data( result = await opc_repository.write_data('ns=2;s=TestNode', 42.0, 'float', metadata['metadata'])
'ns=2;s=TestNode', 42.0, 'float', opc_repository.logger, metadata['metadata']
)
mock_client.get_node.assert_called_once_with('ns=2;s=TestNode') mock_client.get_node.assert_called_once_with('ns=2;s=TestNode')
mock_node.write_value.assert_called_once() mock_node.write_value.assert_called_once()
@@ -431,12 +414,11 @@ async def test_write_data_write_value_failed(opc_repository, mock_client):
opc_repository.validate_connection = AsyncMock(return_value=(True, {})) opc_repository.validate_connection = AsyncMock(return_value=(True, {}))
opc_repository.client = mock_client opc_repository.client = mock_client
mock_node = AsyncMock() mock_node = AsyncMock()
opc_repository.error_count = 0
mock_client.get_node = MagicMock(return_value=mock_node) mock_client.get_node = MagicMock(return_value=mock_node)
mock_node.write_value.side_effect = Exception('Test error') mock_node.write_value.side_effect = Exception('Test error')
is_success, error_data = await opc_repository.write_data( is_success, error_data = await opc_repository.write_data(
'ns=2;s=TestNode', 42.0, 'float', opc_repository.logger, metadata['metadata'] 'ns=2;s=TestNode', 42.0, 'float', metadata['metadata']
) )
opc_repository.validate_connection.assert_called_once() opc_repository.validate_connection.assert_called_once()
@@ -451,3 +433,178 @@ async def test_write_data_write_value_failed(opc_repository, mock_client):
assert error_data['block'] == 'opc_repository' assert error_data['block'] == 'opc_repository'
assert error_data['level'] == NotificationLevel.ERROR assert error_data['level'] == NotificationLevel.ERROR
assert error_data['attachment_content'] is not None assert error_data['attachment_content'] is not None
def test_is_reconnectable_opcua_bad():
assert is_reconnectable_opcua_bad(BadSessionIdInvalid()) is True
assert is_reconnectable_opcua_bad(BadNodeIdUnknown()) is False
assert is_reconnectable_opcua_bad(Exception('other')) is False
@pytest.mark.asyncio
async def test_write_data_bad_session_id_invalid_schedules_reconnect(opc_repository, mock_client):
opc_repository.validate_connection = AsyncMock(return_value=(True, {}))
opc_repository.client = mock_client
opc_repository._start_reconnect = AsyncMock()
mock_node = AsyncMock()
mock_client.get_node = MagicMock(return_value=mock_node)
mock_node.write_value.side_effect = BadSessionIdInvalid()
is_success, error_data = await opc_repository.write_data(
'ns=2;s=TestNode', 42.0, 'float', metadata['metadata']
)
mock_node.write_value.assert_called_once()
opc_repository._start_reconnect.assert_called_once()
assert is_success is False
assert error_data['opc_error_kind'] == 'session_bad'
assert error_data['opc_status'] == 'BadSessionIdInvalid'
@pytest.mark.asyncio
async def test_write_data_reconnect_in_progress_immediate(opc_repository):
opc_repository._session_ready.clear()
opc_repository._reconnect_task = asyncio.create_task(asyncio.sleep(60))
opc_repository.validate_connection = AsyncMock()
is_success, error_data = await opc_repository.write_data(
'ns=2;s=TestNode', 42.0, 'float', metadata['metadata']
)
opc_repository._reconnect_task.cancel()
with pytest.raises(asyncio.CancelledError):
await opc_repository._reconnect_task
opc_repository._reconnect_task = None
opc_repository.validate_connection.assert_not_called()
assert is_success is False
assert error_data['opc_error_kind'] == 'reconnect_in_progress'
@pytest.mark.asyncio
async def test_start_reconnect_skips_within_interval(opc_repository):
opc_repository.last_reconnection_time = datetime.now()
opc_repository.reconnection_interval = 3600
await opc_repository._start_reconnect('BadSessionIdInvalid', 'tok')
assert opc_repository._reconnect_task is None
@pytest.mark.asyncio
async def test_write_data_protocol_closed_schedules_reconnect(opc_repository):
opc_repository.client = MagicMock()
opc_repository.client.uaclient.protocol = MagicMock(state='closed')
opc_repository._start_reconnect = AsyncMock()
is_success, error_data = await opc_repository.write_data(
'ns=2;s=TestNode', 42.0, 'float', metadata['metadata']
)
opc_repository._start_reconnect.assert_called_once()
assert opc_repository._start_reconnect.call_args.args[0] == 'ProtocolClosed'
assert is_success is False
assert error_data['opc_error_kind'] == 'connection_lost'
assert error_data['opc_status'] == 'ProtocolClosed'
@pytest.mark.asyncio
async def test_write_data_protocol_closed_skips_reconnect_within_interval(opc_repository):
opc_repository.client = MagicMock()
opc_repository.client.uaclient.protocol = MagicMock(state='closed')
opc_repository.last_reconnection_time = datetime.now()
opc_repository.reconnection_interval = 3600
is_success, error_data = await opc_repository.write_data(
'ns=2;s=TestNode', 42.0, 'float', metadata['metadata']
)
assert opc_repository._reconnect_task is None
assert is_success is False
assert error_data['opc_error_kind'] == 'connection_lost'
@pytest.mark.asyncio
async def test_write_data_after_failed_reconnect_schedules_again(opc_repository):
opc_repository._session_ready.clear()
opc_repository.reconnection_interval = 0
opc_repository.last_reconnection_time = None
opc_repository._reconnect_locked = AsyncMock(
return_value=(False, {'message': 'connect failed'})
)
await opc_repository.write_data('ns=2;s=TestNode', 42.0, 'float', metadata['metadata'])
await asyncio.sleep(0.1)
assert opc_repository._reconnect_locked.call_count == 1
assert not opc_repository._reconnect_task_in_progress()
await opc_repository.write_data('ns=2;s=TestNode', 42.0, 'float', metadata['metadata'])
await asyncio.sleep(0.1)
assert opc_repository._reconnect_locked.call_count == 2
@pytest.mark.asyncio
async def test_write_data_after_disconnect_does_not_schedule_reconnect(opc_repository, mock_client):
opc_repository.client = mock_client
proto = MagicMock()
proto.state = 'closed'
mock_client.uaclient = MagicMock(protocol=proto)
opc_repository._disconnection_fallback = AsyncMock(return_value=[])
await opc_repository.disconnect()
is_success, error_data = await opc_repository.write_data(
'ns=2;s=TestNode', 42.0, 'float', metadata['metadata']
)
assert opc_repository._reconnect_task is None
assert is_success is False
assert error_data['opc_error_kind'] == 'connection_lost'
@pytest.mark.asyncio
async def test_parallel_bad_writes_single_reconnect_task(opc_repository, mock_client):
opc_repository.validate_connection = AsyncMock(return_value=(True, {}))
opc_repository.client = mock_client
opc_repository.reconnection_interval = 0
opc_repository.last_reconnection_time = None
mock_node = AsyncMock()
mock_client.get_node = MagicMock(return_value=mock_node)
mock_node.write_value.side_effect = BadSessionIdInvalid()
connect_count = 0
async def slow_reconnect():
nonlocal connect_count
connect_count += 1
await asyncio.sleep(0.05)
opc_repository._session_ready.set()
return True, {}
opc_repository._reconnect_locked = slow_reconnect
results = await asyncio.gather(
opc_repository.write_data('ns=2;s=TestNode', 1.0, 'float', metadata['metadata']),
opc_repository.write_data('ns=2;s=TestNode2', 2.0, 'float', metadata['metadata']),
)
await asyncio.sleep(0.15)
assert connect_count <= 1
assert 1 <= mock_node.write_value.call_count <= 2
error_kinds = [r[1].get('opc_error_kind') for r in results]
assert error_kinds.count('session_bad') >= 1
assert all(k in ('session_bad', 'reconnect_in_progress') for k in error_kinds)
@pytest.mark.asyncio
@patch('laborious.utils.repository.opc_repository.datetime')
async def test_reconnect_locked_sets_last_reconnection_time(mock_datetime, opc_repository):
mock_datetime.now = MagicMock(return_value=datetime(2025, 1, 1, 12, 0, 0))
opc_repository._disconnect_locked = AsyncMock()
opc_repository._connect_locked = AsyncMock(return_value=(True, {}))
result = await opc_repository._reconnect_locked()
opc_repository._disconnect_locked.assert_called_once()
opc_repository._connect_locked.assert_called_once()
assert result == (True, {})
assert opc_repository.last_reconnection_time == datetime(2025, 1, 1, 12, 0, 0)

View File

@@ -0,0 +1,14 @@
from sientia_do.temporal.worker.prepare_worker import build_queue_name
from laborious.workflows.drift import Drift
from laborious.workflows.minimal_retrain import MinimalRetrain
from laborious.workflows.predictions_batch import PredictionsBatch
from laborious.workflows.simple_metrics import SimpleMetrics
def test_runtime_scoped_queue_names():
runtime = 'prod-a'
assert build_queue_name(PredictionsBatch.__name__, runtime) == 'predictions_batch-prod-a-queue'
assert build_queue_name(MinimalRetrain.__name__, runtime) == 'minimal_retrain-prod-a-queue'
assert build_queue_name(Drift.__name__, runtime) == 'drift-prod-a-queue'
assert build_queue_name(SimpleMetrics.__name__, runtime) == 'simple_metrics-prod-a-queue'

View File

@@ -1,315 +0,0 @@
# Default values for sientia-module.
# This is a YAML-formatted file.
# Declare variables to be passed into your templates.
# This will set the replicaset count more information can be found here: https://kubernetes.io/docs/concepts/workloads/controllers/replicaset/
replicaCount: 1
# This sets the container image more information can be found here: https://kubernetes.io/docs/concepts/containers/images/
image:
repository: aignosi.azurecr.io/sientia-module
# This sets the pull policy for images.
pullPolicy: Always
# Overrides the image tag whose default is the chart appVersion.
tag: "1.1.2"
# This is for the secrets for pulling an image from a private repository more information can be found here: https://kubernetes.io/docs/tasks/configure-pod-container/pull-image-private-registry/
imagePullSecrets:
- name: docker-hub-secret
# This is to override the chart name.
nameOverride: "sientia-laborious-worker"
fullnameOverride: "sientia-laborious-worker"
namespace: sientia
# This section builds out the service account more information can be found here: https://kubernetes.io/docs/concepts/security/service-accounts/
serviceAccount:
# Specifies whether a service account should be created
create: true
# Automatically mount a ServiceAccount's API credentials?
automount: true
# Annotations to add to the service account
annotations: {}
# The name of the service account to use.
# If not set and create is true, a name is generated using the fullname template
name: "sientia-laborious-worker"
# This is for setting Kubernetes Annotations to a Pod.
# For more information checkout: https://kubernetes.io/docs/concepts/overview/working-with-objects/annotations/
podAnnotations: {}
# This is for setting Kubernetes Labels to a Pod.
# For more information checkout: https://kubernetes.io/docs/concepts/overview/working-with-objects/labels/
podLabels: {}
podSecurityContext: {}
# fsGroup: 2000
securityContext: {}
# capabilities:
# drop:
# - ALL
# readOnlyRootFilesystem: true
# runAsNonRoot: true
# runAsUser: 1000
resources:
# Resource limits and requests are important for ResourceBasedTuner to work correctly.
# The tuner monitors system CPU and memory usage, so proper resource limits must be set.
limits:
cpu: 2000m # 2 CPU cores
memory: 20Gi # 20 GB memory
requests:
cpu: 1000m # 1 CPU core
memory: 2Gi # 2 GB memory
# This is to setup the liveness and readiness probes more information can be found here: https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/
# This is to setup the liveness and readiness probes more information can be found here: https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/
livenessProbe:
exec:
command:
- sh
- -c
- |
curl -sf http://localhost:9090/metrics | grep -q '^app_up{.*} 1'
initialDelaySeconds: 1260
periodSeconds: 15
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
exec:
command:
- sh
- -c
- |
curl -sf http://localhost:9090/metrics | grep -q '^app_up{.*} 1'
initialDelaySeconds: 1200
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 2
# This section is for setting up autoscaling more information can be found here: https://kubernetes.io/docs/concepts/workloads/autoscaling/
autoscaling:
enabled: false
minReplicas: 1
maxReplicas: 100
targetCPUUtilizationPercentage: 80
# targetMemoryUtilizationPercentage: 80
# Additional volumes on the output Deployment definition.
volumes: []
# - name: foo
# secret:
# secretName: mysecret
# optional: false
# Additional volumeMounts on the output Deployment definition.
volumeMounts: []
# - name: foo
# mountPath: "/etc/foo"
# readOnly: true
nodeSelector: {}
tolerations: []
affinity: {}
services:
sdk-metrics:
enabled: true
type: ClusterIP
port: 9091
targetPort: 9091
name: sdk-metrics
metrics:
enabled: true
type: ClusterIP
port: 9090
targetPort: 9090
name: metrics
# Configuração do ServiceMonitor para o Prometheus Operator
# ref: https://github.com/prometheus-operator/prometheus-operator
serviceMonitor:
# Se true, um recurso ServiceMonitor será criado.
enabled: true
# O intervalo no qual as métricas devem ser coletadas (ex: 30s, 1m).
endpoints:
- port: metrics
path: /metrics
interval: 30s
relabelings: []
- port: sdk-metrics
path: /metrics
interval: 30s
relabelings: []
additionalLabels:
release: kube-prometheus-stack
env:
# Entrypoint variables
- name: GITHUB_REPO_URL
value: "git@github.com:Aignosi/sientia-dataops-laborious_temporal.git"
- name: GITHUB_BRANCH
value: "feature/SIENTIAPDE-1712"
- name: PYTHON_APP
value: "laborious.worker.worker"
# Application variables
- name: POSTGRES_HOST
value: "paradedb-rw.paradedb.svc.cluster.local"
- name: POSTGRES_PORT
value: "5432"
- name: POSTGRES_USER
value: "postgres"
- name: POSTGRES_PASSWORD
value: "nFqc81y6kwmr2zuAIx43DhiOosFCVPpeEfTtTWZflkNjB2j1KtEeIANkhFR9mAX3"
- name: POSTGRES_DBNAME
value: "sientia"
- name: POSTGRES_MIN_CONNECTIONS
value: "20"
# max_connections = number_of_workers * max_concurrent_activities * safety_factor
# Example: 4 workers * 50 activities * 0.5 = 100 connections
- name: POSTGRES_MAX_CONNECTIONS
value: "100"
- name: MLFLOW_HOST
value: "http://sientia-tracker-mlflow-tracking.sientia-tracker.svc.cluster.local"
- name: MLFLOW_PORT
value: "80"
- name: MLFLOW_USERNAME
value: "aignosi"
- name: MLFLOW_PASSWORD
value: "1L0FP50j3ncp123"
- name: OPC_ID
value: "1"
- name: OPC_SERVER_NAME
value: "default_server"
- name: OPC_URL
value: "opc.tcp://sientia-opc-simulator-opc.sientia.svc.cluster.local:4840"
- name: LOG_LEVEL
value: "DEBUG"
- name: HTTP_METRICS_PORT
value: "9090"
- name: HTTP_SDK_METRICS_PORT
value: "9091"
- name: PROJECT_NAME
value: "sientia-laborious"
- name: TEMPORAL_HOST
value: "temporal-frontend.temporal.svc.cluster.local:7233"
- name: TEMPORAL_NAMESPACE
value: "laborious"
- name: MONGODB_USERNAME
value: "root"
- name: MONGODB_PASSWORD
value: "wKZDbMNU1c"
- name: MONGODB_URL
value: "my-release-mongodb.mongodb.svc.cluster.local:27017"
- name: MONGODB_DATABASE
value: "sientia"
- name: MONGODB_TTL_INDEX_HOURS
value: "1"
- name: MINIO_ENDPOINT_URL
value: "minio.minio.svc.cluster.local:9000"
- name: MINIO_ACCESS_KEY
value: "admin"
- name: MINIO_SECRET_KEY
value: "LiArt4eNmJ"
- name: MINIO_DEFAULT_BUCKET
value: "sientia"
- name: MINIO_RETENTION_HOURS
value: "24"
- name: SIENTIA_MINIO_OFFLOAD_THRESHOLD_MEGABYTES
value: "0.5"
# Temporal worker tuning for PredictionsBatch.
# IMPORTANT: prefix must be PREDICTIONSBATCH_ (from class name PredictionsBatch).
# Keep workflow-task concurrency moderate to reduce task completion races under load.
- name: PREDICTIONSBATCH_MAX_CONCURRENT_WORKFLOW_TASKS
value: "20"
# Allow higher activity parallelism because most activities are I/O-bound, but keep headroom.
- name: PREDICTIONSBATCH_MAX_CONCURRENT_ACTIVITIES
value: "60"
# Keep local activities controlled so they do not monopolize the event loop.
- name: PREDICTIONSBATCH_MAX_CONCURRENT_LOCAL_ACTIVITIES
value: "20"
# Cache enough workflows for reuse without excessive memory growth.
- name: PREDICTIONSBATCH_MAX_CACHED_WORKFLOWS
value: "200"
# Start with one workflow poller to avoid burst contention at startup.
- name: PREDICTIONSBATCH_WORKFLOW_POLLER_BEHAVIOUR_MINIMUM
value: "3"
# Small initial poller count warms up gradually instead of spiking task fetches.
- name: PREDICTIONSBATCH_WORKFLOW_POLLER_BEHAVIOUR_INITIAL
value: "5"
# Cap workflow pollers to limit scheduling pressure and avoid over-polling.
- name: PREDICTIONSBATCH_WORKFLOW_POLLER_BEHAVIOUR_MAXIMUM
value: "15"
# Keep at least two activity pollers so activity queues do not starve during spikes.
- name: PREDICTIONSBATCH_ACTIVITY_POLLER_BEHAVIOUR_MINIMUM
value: "3"
# Moderate initial activity pollers for faster ramp-up with controlled pressure.
- name: PREDICTIONSBATCH_ACTIVITY_POLLER_BEHAVIOUR_INITIAL
value: "10"
# Limit max activity pollers to preserve CPU for workflow-task completion.
- name: PREDICTIONSBATCH_ACTIVITY_POLLER_BEHAVIOUR_MAXIMUM
value: "30"
- name: MINIMALRETRAIN_MAX_CONCURRENT_ACTIVITIES
value: "1"
- name: MINIMALRETRAIN_MAX_CONCURRENT_LOCAL_ACTIVITIES
value: "1"
- name: MINIMALRETRAIN_MAX_CACHED_WORKFLOWS
value: "1"
- name: MINIMALRETRAIN_WORKFLOW_POLLER_BEHAVIOUR_MINIMUM
value: "1"
- name: MINIMALRETRAIN_WORKFLOW_POLLER_BEHAVIOUR_INITIAL
value: "1"
- name: MINIMALRETRAIN_WORKFLOW_POLLER_BEHAVIOUR_MAXIMUM
value: "1"
- name: MINIMALRETRAIN_ACTIVITY_POLLER_BEHAVIOUR_MINIMUM
value: "1"
- name: MINIMALRETRAIN_ACTIVITY_POLLER_BEHAVIOUR_INITIAL
value: "1"
- name: MINIMALRETRAIN_ACTIVITY_POLLER_BEHAVIOUR_MAXIMUM
value: "1"
- name: PI_WEB_API_BASE_URL
value: "https://pivision.votorantimcimentos.com/piwebapi"
- name: PI_WEB_API_AUTH_TYPE
value: "basic"
- name: PI_WEB_API_AUTH_TOKEN
valueFrom:
secretKeyRef:
name: pi-web-api-auth-token
key: token
- name: PYPI_SERVER
value: "http://library-distribution-server.library.svc.cluster.local:5000"
ssh:
enabled: true
secretName: git-ssh-key-sientia-laborious-worker
sshPath: /mnt/.ssh
knownHostsPath: /mnt/known_hosts
# kubectl create secret docker-registry docker-hub-secret --namespace sientia --docker-server=http://aignosi.azurecr.io --docker-username=aignosi --docker-password=5I5zpQ6sRaHqX1hD3dr+2mo647yO3FRc359/wu6gsP+ACRDRz5mp
# helm upgrade --install sientia-laborious-worker sientia/sientia-module -n sientia --create-namespace -f ./values.yaml --version 0.6.0
# kubectl create secret generic git-ssh-key-sientia-laborious-worker \
# --namespace sientia \
# --from-file=ssh-privatekey=git_key \
# --type=kubernetes.io/ssh-auth