Remove Pytest test execution from CI workflow and add a blank line in MinioRepository class for improved readability.
Sientia DataOps Laborious
A comprehensive, Temporal-based ML orchestration system for industrial data processing and model inference. Laborious delivers enterprise-grade batch prediction, model management, optional real-time export (OPC), and automated retraining with strong data quality validation and observability.
📑 Table of Contents
- Features
- Architecture
- Workflows
- Installation & Setup
- How to Run
- Code Quality & Validation
- Testing
- Monitoring and Metrics
- Configuration
- Development
- Troubleshooting
- Performance Tuning
- Contributing
- License
- Support
Features
Core Functionality
- Batch Prediction Processing: High-throughput ML inference using MLFlow models
- Temporal Workflow Orchestration: Robust workflow management with retries and fault tolerance
- Data Quality Gates: Configurable filtering for input data and MLFlow API responses
- Multi-Model Support: Flexible model management with retention and versioning
- Optional Real-time Export: PostgreSQL persistence and OPC server integration for industrial systems
- Comprehensive Monitoring: Prometheus metrics and structured logging for observability
Advanced Capabilities
- Incremental Data Processing: Timestamp-based loading to avoid reprocessing
- Configurable Data Retention: Model retention policies with automatic cleanup
- Notification System: Integrated alerting via MongoDB
- Scalable Architecture: Kubernetes-ready with horizontal scaling
- Model Retraining: Automated retraining workflows with production model updates
Development & Quality Assurance
- Code Quality Tools: Ruff (lint/format), mypy (types), Bandit (security)
- Automated Validation:
validate.shand CI quality gates - Comprehensive Testing: pytest with async support and high coverage
- Type Safety: Static type checking with mypy
- Coverage Visualization: Coverage Gutters integration
Architecture
Laborious uses a Temporal-based architecture with strong separation of concerns and defensive error handling for production ML.
Architecture Principles
1. Separation of Concerns
- Worker Layer: Temporal workers, task queues, lifecycle
- Workflow Layer: Business orchestration and coordination
- Activity Layer: External system interactions and isolated operations
- Data Layer: Persistence, caching, connectors
2. Fault Tolerance & Resilience
- Automatic Retry Policies for transient failures
- Graceful Degradation and circuit breaking for dependencies
- Detailed Error Handling with notifications
3. Scalability & Performance
- Horizontal Scaling: Multiple worker instances for load distribution
- Task Queue Isolation: Separate queues for different workflow types
- Connection Pooling: Optimized database and external service connections
- Asynchronous Processing: Non-blocking operations for improved throughput
4. Observability & Monitoring
- Prometheus Metrics: Comprehensive system and business metrics
- Structured Logging: Consistent log format with correlation IDs
- Health Checks: Endpoint health monitoring and alerting
- Performance Tracing: Request flow tracking and bottleneck identification
Key Components
Worker (laborious/worker/worker.py)
- Temporal client setup, worker lifecycle, task queues
- Metrics server initialization, notification handler setup
- Graceful shutdown and autoscaling-friendly behavior
Workflows (laborious/workflows/)
predictions_batch.py: Batch prediction entry pointsub_workflows/prediction_process.py: Core prediction pipelinesub_workflows/format_and_export_prediction.py: Formatting and exportminimal_retrain.py: Automated model retraining and production update
Activities (laborious/activities/)
gates.py: Data quality validation and filteringmlflow.py: Transform and predict operationsopc.py: OPC UA export to industrial systems (optional)activities.py: Aggregates activity interfaces
Data Services (laborious/utils/)
connectors_config.py: Env-driven configuration buildersrepository/model_repository.py: MLFlow operations and retrainingrepository/opc_repository.py: OPC communication and writesfilters/conditional_filters.pyandfilters/mlflow_filters.py
Data Flow Architecture
1. Batch Prediction Pipeline
Input Data (PostgreSQL) → Data Quality Gates → MLFlow Transform →
MLFlow Prediction → Response Validation → Export (PostgreSQL [+ OPC])
2. Model Retraining Pipeline
Training Data → Model Retraining → Quality Validation →
Production Update → Notification & Monitoring
Security Architecture
Authentication & Authorization
- MLFlow API Authentication: Username/password
- Database Security: Encrypted connections and credential management
- OPC Certificates (if enabled): Client/server certs
- Kubernetes Secrets: Secure secret storage
Network Security
- TLS/SSL, network policies, service mesh, firewalls, VPN
Data Security
- At-rest/in-transit encryption, RBAC, audit logging, lifecycle management
Workflows
1. Predictions Batch Workflow (predictions_batch.py)
The PredictionsBatch workflow is the main entry point for batch prediction pipelines. It orchestrates the complete prediction process and implements a robust data loading and processing pattern.
Purpose
- Batch Prediction Orchestration: Coordinates data loading and prediction processing
- Data Preparation: Loads data using custom SQL queries with configurable schemas
- Workflow Delegation: Delegates actual prediction processing to the PredictionProcess workflow
- Configuration Management: Handles model configuration, filters, and retention policies
Execution Flow
- Data Loading: Executes custom SQL query to load data from PostgreSQL
- Input Preparation: Prepares prediction input with metadata and configuration
- Workflow Delegation: Spawns PredictionProcess child workflow for actual processing
- Error Handling: Implements comprehensive error handling with retry policies
Key Features
- Custom Query Support: Flexible SQL-based data loading
- Schema Configuration: Configurable data schema definitions
- Automatic Retry: Implements Temporal retry policies for fault tolerance
- Timeout Management: 60-second timeout for all activities
- Comprehensive Error Handling: Detailed error reporting and notification integration
Input Parameters
{
"schedule_name": "hourly_predictions",
"model_name": "temperature_prediction_model",
"model_id": "temp_pred_001",
"query": "SELECT * FROM sensor_data WHERE timestamp > NOW() - INTERVAL '1 hour'",
"schema": {
"timestamp": "datetime",
"temperature": "float",
"humidity": "float"
},
"table_name": "predictions",
"input_filters": {
"EMPTY_DATA": {"POLICY": "STOP"}
},
"mlflow_transform_filters": {
"API_ERROR": {"POLICY": "STOP"}
},
"mlflow_predict_filters": {
"API_ERROR": {"POLICY": "STOP"}
},
"model_retention": 60,
"path_priority": ["STOP", "CONTINUE", "REPEAT"],
"opc_output_config": {
"server_id": "opc_server_1",
"tags": ["prediction_output"]
}
}
Architecture Diagram
flowchart LR
A[1. load_custom_query] --> B[2. prediction_process 🔃]
A -.-> Database[(Database)]
2. Prediction Process Workflow (prediction_process.py)
The PredictionProcess workflow implements the core prediction pipeline for ML model inference. It handles data quality validation, MLFlow model interactions, and prediction processing.
Purpose
- Data Quality Validation: Applies configurable filters for data integrity
- MLFlow Integration: Manages model transformation and prediction requests
- Response Validation: Filters MLFlow API responses for quality assurance
- Prediction Export: Delegates prediction formatting and export operations
Execution Flow
- Timestamp Retrieval: Gets the last processed timestamp for incremental processing
- Input Data Gate: Applies configured filters for data quality validation
- Path Decision: Determines processing path based on filter results
- MLFlow Transform: Requests data transformation using MLFlow models
- Response Validation: Filters transform responses for quality assurance
- MLFlow Prediction: Executes prediction using transformed data
- Content Validation: Filters prediction responses for final quality check
- Export Delegation: Delegates to FormatAndExportPrediction workflow
Key Features
- Configurable Quality Gates: Multiple filter types with policy-based configuration
- Flexible Path Handling: Configurable decision paths (STOP, CONTINUE, REPEAT)
- MLFlow Integration: Comprehensive model management and inference
- Incremental Processing: Timestamp-based data processing optimization
- Comprehensive Monitoring: Detailed metrics and error reporting
Input Parameters
{
"metadata": {
"schedule_name": "hourly_predictions",
"model_name": "temperature_prediction_model",
"model_id": "temp_pred_001",
"workflow_name": "predictions_batch"
},
"data": {...},
"schema": {...},
"table_name": "predictions",
"model_id": "temp_pred_001",
"model_name": "temperature_prediction_model",
"input_filters": {
"EMPTY_DATA": {"POLICY": "STOP"},
"SPECIFIC_VARIABLES_NULL_VALUES": {
"POLICY": "STOP",
"config": {"variables": ["temperature", "humidity"]}
}
},
"mlflow_transform_filters": {
"API_ERROR": {"POLICY": "STOP"}
},
"mlflow_predict_filters": {
"API_ERROR": {"POLICY": "STOP"},
"NAN_VALUES": {"POLICY": "STOP"}
},
"model_retention": 60,
"path_priority": ["STOP", "CONTINUE", "REPEAT"],
"opc_output_config": {...}
}
Architecture Diagram
flowchart LR
A[1. get_last_timestamp] --> B[2. input_gate] --> C[3. request_transform] --> D[4. mlflow_response_gate] --> E[5. mlflow_content_gate] --> F[6. request_predict] --> G[7. mlflow_response_gate] --> H[8. format_and_export_prediction🔃]
A -.-> Redis[(Redis)]
C -.-> MLFlow[MLFlow]
F -.-> MLFlow[MLFlow]
G -.-> Filters[MLFlow Filters]
3. Format and Export Prediction Workflow (format_and_export_prediction.py)
The FormatAndExportPrediction workflow handles prediction data formatting and export operations to multiple destinations.
Purpose
- Data Formatting: Formats prediction data for different output destinations
- PostgreSQL Export: Persists predictions to database with metrics
- OPC Integration: Writes predictions to OPC servers for real-time access
- Metrics Recording: Tracks export operations and performance metrics
Execution Flow
- Path Decision: Determines formatting path based on configuration
- Data Formatting: Formats prediction data for specific output requirements
- PostgreSQL Export: Writes formatted predictions to database
- OPC Export: Writes predictions to OPC servers
- Metrics Recording: Records export performance and success metrics
Key Features
- Flexible Formatting: Configurable output formats for different destinations
- Multi-Destination Export: PostgreSQL and OPC server integration
- Performance Monitoring: Comprehensive metrics for export operations
- Error Handling: Robust error handling with notification integration
Architecture Diagram
flowchart LR
A[1. format_prediction/format_default_prediction] --> B[2. write_opc_data] --> C[3. export_data_to_postgres] --> D[4. write_metrics]
A -.-> Format[Data Formatting]
B -.-> OPC[OPC Servers]
C -.-> PostgreSQL[(PostgreSQL)]
D -.-> Prometheus[Prometheus]
4. Minimal Retrain Workflow (minimal_retrain.py)
The MinimalRetrain workflow handles automated model retraining and production model updates.
Purpose
- Model Retraining: Automates ML model retraining processes
- Production Updates: Manages production model version updates
- Data Export: Exports training data for model development
- Quality Assurance: Ensures model quality before production deployment
Execution Flow
- Data Loading: Loads training data using custom queries
- Model Retraining: Executes model retraining process
- Quality Validation: Validates retrained model performance
- Production Update: Updates production model if quality criteria met
- Data Export: Exports training data for analysis
Architecture Diagram
flowchart LR
A[1. load_custom_query] --> B[2. retrain_model] --> C[3. update_production_model] --> D[4. export_data_to_postgres]
A -.-> Database[(Database)]
B -.-> MLFlow[MLFlow]
C -.-> MLFlow[MLFlow]
D -.-> PostgreSQL[(PostgreSQL)]
📋 Prerequisites
- Python 3.11+
- Temporal server/cluster
- PostgreSQL database
- MLFlow server
- MinIO object storage (for MLFlow artifacts)
- MongoDB server (for notifications)
- OPC server(s) if using OPC export
Note: External dependencies must be available either through:
- Kubernetes cluster deployment
- Docker Compose setup
- Cloud-managed services
- Local installations
🚀 Installation
Local Development Setup
-
Clone the repository
git clone <repository-url> cd sientia-dataops-laborious -
Create virtual environment
python3.11 -m venv venv source ./venv/bin/activate -
Install dependencies
-
Install github cli
bash sudo apt update sudo apt install gh -y -
Authenticate with github
bash gh auth login -
Run the install_dependencies.sh script
bash chmod +x install_dependencies.sh ./install_dependencies.sh -
Create environment configuration file
cp .env.example .env # Edit .env with your connection details -
Configure external dependencies
You'll need to set up port forwarding or connections to external services. For example:
# Port forwarding from Kubernetes cluster kubectl port-forward svc/postgresql 5432:5432 kubectl port-forward svc/mlflow 5000:5000 kubectl port-forward svc/mongodb 27017:27017 # Or connect to external services # Ensure services are accessible on localhost with appropriate ports
📦 How to Run
Running the Laborious Application
Use the provided script to run the application locally:
# Make script executable (first time only)
chmod +x run_local.sh
# Run the application
./run_local.sh
The script will:
- Activate the virtual environment
- Load environment variables from
.env - Start the laborious worker application
Running Tests and Coverage
Use the provided script to run tests with coverage:
# Make script executable (first time only)
chmod +x run_coverage.sh
# Run tests with coverage
./run_coverage.sh
The script will:
- Activate the virtual environment
- Run pytest with coverage reporting
- Generate HTML coverage report
- Open the coverage report in your browser
Manual Test Execution
You can also run tests manually:
# Activate virtual environment
source ./venv/bin/activate
# Run all tests
pytest
# Run with coverage
pytest --cov=laborious --cov-report=html
# Run specific test categories
pytest tests/activities/
pytest tests/workflow/
Manual Application Execution
For manual execution without scripts:
# Activate virtual environment
source ./venv/bin/activate
# Load environment variables (if using .env file)
if [ -f .env ]; then
export $(cat .env | grep -v '^#' | xargs)
fi
# Start the laborious worker
python -m laborious.worker.worker
Code Quality & Validation
Overview
Since Python is not compiled, we validate quality, security, and correctness before execution.
Validation Tools
- Ruff: Linting and formatting
- mypy: Static type checking
- Bandit: Security analysis
- pytest: Unit/integration testing with coverage
Tools Installation
pip install -r requirements-dev.txt
Complete Validation
Option 1 (recommended):
./validate.sh
The script runs, in order:
- Format check (Ruff)
- Linting (Ruff)
- Type checking (mypy)
- Security analysis (Bandit)
- Tests with coverage (pytest)
Option 2 (individual commands):
ruff format --check laborious/ tests/
ruff check laborious/ tests/
mypy laborious/
bandit -r laborious/ -ll
pytest tests/ --cov=laborious --cov-report=term-missing
Automatic Fixes
ruff format laborious/ tests/
ruff check --fix laborious/ tests/
Configuration
All settings reside in pyproject.toml (Ruff, mypy, pytest, Bandit).
CI/CD Integration
The workflow at .github/workflows/quality-gate.yml executes validations on each push/PR.
Best Practices
- Run
./validate.shbefore committing - Use
ruff check --watchfor continuous feedback - Add type hints and tests for new code
🧪 Testing
Test Structure
tests/
├── activities/ # Activity implementation tests
├── workflow/ # Workflow orchestration tests
├── utils/ # Utility function tests
└── integration/ # End-to-end workflow tests
Test Execution
# Install test dependencies
pip install pytest pytest-cov pytest-asyncio
# Run tests with coverage
pytest --cov=laborious --cov-report=html
# Run specific test modules
pytest tests/activities/test_gates.py
pytest tests/workflow/test_predictions_batch.py
📊 Monitoring and Metrics
The Laborious system exposes comprehensive Prometheus metrics for operational visibility and performance monitoring:
Application Health Metrics
app_up: Application health status (1=healthy, 0=unhealthy)- Labels:
pod_id
- Labels:
Prediction Operation Metrics
laborious_predictions_written_count: Counter for successful prediction exports- Labels:
pod_id,model_name,pipeline_name
- Labels:
laborious_prediction_confidence_monitor: Gauge for current prediction confidence levels- Labels:
pod_id,model_name,pipeline_name
- Labels:
laborious_prediction_response_time_monitor: Histogram for prediction response times- Labels:
pod_id,model_name,pipeline_name - Buckets: [0.01, 0.05, 0.1, 0.2, 0.5, 1.0, 2.0, 5.0, 10.0]
- Labels:
OPC Export Metrics
laborious_prediction_opc_writing_count: Counter for OPC server write operations- Labels:
pod_id,model_name,pipeline_name,opc_server_id
- Labels:
laborious_prediction_opc_writing_response_time_monitor: Histogram for OPC write response times- Labels:
pod_id,model_name,pipeline_name,opc_server_id - Buckets: [0.01, 0.05, 0.1, 0.2, 0.5, 1.0, 2.0, 5.0, 10.0]
- Labels:
Data Quality Metrics
- Filter pass/fail rates through notification system
- MLFlow API response validation metrics
- Data quality gate performance tracking
⚙️ Configuration
Environment Variables
| Variable | Description | Default | Required |
|---|---|---|---|
TEMPORAL_HOST |
Temporal server address | localhost:7233 |
Yes |
TEMPORAL_NAMESPACE |
Temporal namespace | laborious |
No |
POSTGRES_HOST |
PostgreSQL hostname | localhost |
Yes |
POSTGRES_PORT |
PostgreSQL port | 5432 |
Yes |
POSTGRES_USER |
PostgreSQL username | sientia |
Yes |
POSTGRES_PASSWORD |
PostgreSQL password | sientia |
Yes |
POSTGRES_DBNAME |
PostgreSQL database | sientia |
Yes |
POSTGRES_MIN_CONNECTIONS |
Minimum PostgreSQL connections | 5 |
No |
POSTGRES_MAX_CONNECTIONS |
Maximum PostgreSQL connections | 20 |
No |
MLFLOW_HOST |
MLFlow server hostname | http://localhost |
Yes |
MLFLOW_PORT |
MLFlow server port | 5080 |
Yes |
MLFLOW_USERNAME |
MLFlow username | aignosi |
Yes |
MLFLOW_PASSWORD |
MLFlow password | aignosi |
Yes |
OPC_CONFIG |
OPC server configuration (JSON) | {} |
No |
OPC_ID |
OPC server identifier | 1 |
No |
OPC_URL |
OPC server URL | opc.tcp://localhost:4840 |
No |
OPC_SERVER_URI |
OPC server URI | opc.tcp://localhost:4840 |
No |
OPC_CERT_PATH |
OPC client certificate path | None |
No |
OPC_PRIVATE_KEY_PATH |
OPC private key path | None |
No |
OPC_SERVER_CERT_PATH |
OPC server certificate path | None |
No |
OPC_RECONNECTION_INTERVAL |
OPC reconnection interval (ms) | 120 |
No |
MONGODB_URL |
MongoDB connection URI | localhost:27018 |
Yes |
MONGODB_USERNAME |
MongoDB username | root |
Yes |
MONGODB_PASSWORD |
MongoDB password | wKZDbMNU1c |
Yes |
MONGODB_DATABASE_NAME |
MongoDB database name | sientia |
Yes |
MONGODB_TTL_INDEX_HOURS |
MongoDB TTL index hours | 1 |
No |
LOG_LEVEL |
Application log level | INFO |
No |
PROJECT_NAME |
Project name for metrics | laborious |
No |
HTTP_METRICS_PORT |
Prometheus metrics port | 9090 |
No |
HTTP_SDK_METRICS_PORT |
Temporal SDK metrics port | 9091 |
No |
POD_ID |
Kubernetes pod identifier | None |
No |
OPC Configuration
For multiple OPC servers, use the OPC_CONFIG environment variable:
{
"opc_server_1": {
"url": "opc.tcp://server1:4840",
"name": "Server1",
"server_uri": "urn:server1:opcua",
"cert_path": "/path/to/cert.pem",
"private_key_path": "/path/to/key.pem",
"server_cert_path": "/path/to/server_cert.pem",
"reconnection_interval": 5000
},
"opc_server_2": {
"url": "opc.tcp://server2:4840",
"name": "Server2",
"server_uri": "urn:server2:opcua",
"cert_path": "/path/to/cert.pem",
"private_key_path": "/path/to/key.pem",
"server_cert_path": "/path/to/server_cert.pem",
"reconnection_interval": 5000
}
}
For single OPC server, use individual environment variables:
OPC_URLOPC_NAMEOPC_SERVER_URIOPC_CERT_PATHOPC_PRIVATE_KEY_PATHOPC_SERVER_CERT_PATHOPC_RECONNECTION_INTERVAL
Workflow Configuration
MongoDB pipeline configuration:
Predictions Batch Workflow configuration sample
This is the configuration for the Predictions Batch Workflow, to be inserted into the MongoDB pipeline collection.
{
"schedule_name": "laborious-orchestrated-pipeline",
"model_id": "1",
"workflow_type": "predictions_batch",
"frequency": "30s",
"max_retry_policy": 1,
"query": "select * from sientia_data.laborious_data where model_id = 1 and \"timestamp\" > NOW() - INTERVAL '5 minutes' order by \"timestamp\" desc limit 30;",
"write_tags": [
{
"server_id": "server1",
"type": "prediction",
"addr": "ns=2;i=5",
"data_type": "double"
},
{
"server_id": "server1",
"type": "confidence",
"addr": "ns=2;i=6",
"data_type": "double"
}
],
"input_filters": {
"EMPTY_DATA": {"POLICY": "STOP"},
"SPECIFIC_VARIABLES_NULL_VALUES": {
"POLICY": "CONTINUE",
"config": {"variables": ["Counter"]}
}
},
"mlflow_transform_filters": {
"API_ERROR": {"POLICY": "REPEAT"},
"NAN_VALUES": {"POLICY": "STOP"}
},
"mlflow_predict_filters": {
"API_ERROR": {"POLICY": "CONTINUE"}
},
"path_priority": ["STOP", "CONTINUE", "REPEAT"],
"active": true,
"updated_at": {
"$date": "2025-09-16T10:00:00.000Z"
},
"datetime_columns": ["timestamp", "created_at"],
"predictions_storage_policy": "lts:1"
}
This is the configuration created by the Orchestrator in Temporal.
{
"datetime_columns":["timestamp","created_at"],
"frequency":"15m",
"input_filters":{"EMPTY_DATA":{"config":{},"policy":"STOP"}},
"max_retry_policy":1,
"mlflow_predict_filters":{"API_ERROR":{"config":{},"policy":"CONTINUE"}},
"mlflow_transform_filters":{
"API_ERROR":{"config":{},"policy":"CONTINUE"},
"EMPTY_DATA":{"config":{},"policy":"STOP"}
},
"model_config":{
"is_compressed":true,
"predict_flavor":"pyfunc",
"retention_minutes":60,
"retention_target":"artifact",
"transform_function_keyword":"transform"
},
"model_id":"352",
"model_name":"courier",
"opc_output_config":{},
"path_priority":["STOP","CONTINUE","REPEAT"],
"predictions_storage_policy":"lts:1",
"query":"select * from sientia_data.laborious_data where model_id = 352 order by \"timestamp\" desc limit 300;",
"retention_time":3600,
"schedule_name":"laborious-courier",
"schema":"sientia_data",
"table_name":"predictions",
"updated_at":"2025-09-12 19:35:01.600000+0000",
"workflow_type":"predictions_batch"
}
🔧 Development
Project Structure
laborious/
├── activities/ # Temporal activity implementations
│ ├── activities.py # Main activities orchestrator
│ ├── gates.py # Data quality gates and filtering
│ ├── mlflow.py # MLFlow model operations
│ └── opc.py # OPC server operations
├── workflows/ # Temporal workflow definitions
│ ├── predictions_batch.py # Main batch prediction workflow
│ ├── minimal_retrain.py # Model retraining workflow
│ └── sub_workflows/ # Sub-workflow implementations
│ ├── prediction_process.py # Core prediction workflow
│ └── format_and_export_prediction.py # Export workflow
├── worker/ # Worker implementation
│ └── worker.py # Main worker orchestrator
├── utils/ # Utility functions
│ ├── connectors_config.py # Database configuration
│ ├── filters/ # Data quality filters
│ │ ├── conditional_filters.py # Conditional data filters
│ │ └── mlflow_filters.py # MLFlow response filters
│ └── repository/ # Data access layer
│ ├── model_repository.py # MLFlow model operations
│ └── opc_repository.py # OPC server operations
├── metrics.py # Prometheus metrics definitions
└── __init__.py
Adding New Features
- Follow Temporal patterns for new workflows and activities
- Add comprehensive docstrings for all public methods
- Include Prometheus metrics for monitoring
- Add unit tests for new functionality
- Update this README with new features and configuration
🐛 Troubleshooting
Common Issues
-
Temporal Connection Failures
- Verify Temporal server is running and accessible
- Check namespace configuration and permissions
- Review server logs for connection issues
-
MLFlow Connection Issues
- Verify MLFlow server is running and accessible
- Check authentication credentials and permissions
- Ensure model names and versions exist
-
Database Connection Issues
- Verify PostgreSQL service is running
- Check connection credentials and network access
- Ensure proper connection pool configuration
-
OPC Connection Failures
- Verify OPC server is accessible
- Check certificate and key file paths
- Review OPC server logs for connection issues
-
Workflow Execution Failures
- Review activity error logs and notifications
- Check data quality filter configurations
- Verify input data format and required fields
Debug Mode
Enable debug logging by setting the log level:
export LOG_LEVEL=DEBUG
⚡ Performance Tuning
Key Parameters
- Worker Concurrency: Adjust
max_concurrent_workflow_tasksandmax_concurrent_activities - Connection Pools: Optimize database connection pool sizes
- Model Retention: Configure MLFlow model retention based on requirements
- Batch Sizes: Adjust data processing batch sizes for optimal throughput
Scaling Considerations
- Horizontal Scaling: Deploy multiple worker instances
- Task Queue Distribution: Use multiple task queues for different workflow types
- Database Performance: Optimize indexes and connection pooling
- MLFlow Performance: Configure appropriate model serving resources
🤝 Contributing
- Fork the repository
- Create a feature branch
- Make your changes with comprehensive testing
- Update documentation and docstrings
- Submit a pull request
Code Quality Standards
- Follow PEP 8 style guidelines
- Include comprehensive docstrings for all public methods
- Maintain test coverage above 80%
- Use type hints where appropriate
- Follow Temporal.io best practices
📄 License
This project is licensed under the terms specified in the LICENSE file.
🆘 Support
For support and questions:
- Check the troubleshooting section above
- Review the metrics and logs for error patterns
- Open an issue in the project repository
- Contact the development team
Note: The Laborious system is designed for production use in industrial ML environments. Ensure proper security configuration and network isolation for production deployments.