Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to Predictive AIOps
- Overview of predictive analytics within IT operations.
- Identifying data sources for prediction, including logs, metrics, and events.
- Key concepts in time-series forecasting and anomaly pattern recognition.
Designing Incident Prediction Models
- Labelling historical incidents and system behaviour for training.
- Selecting and training models, such as LSTM, Random Forest, or AutoML.
- Assessing model performance and managing false positives.
Data Collection and Feature Engineering
- Ingesting and aligning log and metric data for optimal model input.
- Extracting meaningful features from both structured and unstructured data.
- Managing noise and missing data within operational pipelines.
Automating Root Cause Analysis (RCA)
- Utilising graph-based correlation to map services and infrastructure.
- Applying ML to infer probable root causes from event chains.
- Visualising RCA outcomes through topology-aware dashboards.
Remediation and Workflow Automation
- Integrating with automation platforms, such as Ansible or Rundeck.
- Triggering automated actions like rollbacks, restarts, or traffic redirection.
- Auditing and documenting automated interventions for compliance and insight.
Scaling Intelligent AIOps Pipelines
- Applying MLOps to observability: focusing on retraining and model versioning.
- Executing real-time predictions across distributed nodes.
- Best practices for deploying AIOps in production environments.
Case Studies and Practical Applications
- Analysing real-world incident data using predictive AIOps models.
- Deploying RCA pipelines using both synthetic and production data.
- Reviewing industry use cases, including cloud outages, microservices instability, and network degradations.
Summary and Next Steps
Requirements
- Experience with monitoring systems such as Prometheus or ELK.
- Proficiency in Python and a foundational understanding of machine learning.
- Familiarity with incident management workflows.
Audience
- Senior site reliability engineers (SREs).
- IT automation architects.
- DevOps and observability platform leads.