SIENTIAPDE-1231
Enhance MLFlow and MLFlowRepository with improved data handling and logging - Refactored MLFlow class to sort data by 'created_at' and drop duplicates for better input preparation. - Updated MLFlowRepository methods to include detailed logging for artifact downloads and model predictions. - Introduced LzmaPayloadCodec for efficient payload compression in the worker, optimizing data handling for large payloads. - Enhanced timestamp handling in treated data to ensure compatibility with model expectations.
This commit is contained in:
@@ -232,14 +232,23 @@ class MLFlow(BaseActivity):
|
||||
timestamp = data['timestamp'].max()
|
||||
self.debug(f'Timestamp: {timestamp}', metadata)
|
||||
|
||||
# Sort by created_at in descending order and keep first occurrence of each variable/timestamp pair
|
||||
data = data.sort_values('created_at', ascending=False).drop_duplicates(
|
||||
subset=['variable', 'timestamp'], keep='first'
|
||||
)
|
||||
|
||||
data.drop(columns=['model_id'], inplace=True, errors='ignore')
|
||||
data.drop(columns=['created_at'], inplace=True, errors='ignore')
|
||||
|
||||
data = data.pivot(index='timestamp', columns='variable',
|
||||
values='value')
|
||||
data.sort_index(inplace=True)
|
||||
data.reset_index(inplace=True)
|
||||
# Pivot data for model input format
|
||||
data = data.pivot(
|
||||
index='timestamp', columns='variable',
|
||||
values='value')
|
||||
data.fillna(np.nan, inplace=True)
|
||||
# data.reset_index(inplace=True)
|
||||
data.columns.name = None
|
||||
|
||||
data['timestamp'] = data.index
|
||||
data['timestamp'] = to_datetime(
|
||||
data['timestamp'], format=DATETIME_FORMAT_WITH_TZ).dt.strftime(DATETIME_FORMAT)
|
||||
data['timestamp'] = to_datetime(
|
||||
|
||||
Reference in New Issue
Block a user