Get in Touch
 Duration 14 hours

Course Outline

Introduction to Predictive AIOps

  • Overview of predictive analytics within IT operations.
  • Identifying data sources for prediction, including logs, metrics, and events.
  • Key concepts in time-series forecasting and anomaly pattern recognition.

Designing Incident Prediction Models

  • Labelling historical incidents and system behaviour for training.
  • Selecting and training models, such as LSTM, Random Forest, or AutoML.
  • Assessing model performance and managing false positives.

Data Collection and Feature Engineering

  • Ingesting and aligning log and metric data for optimal model input.
  • Extracting meaningful features from both structured and unstructured data.
  • Managing noise and missing data within operational pipelines.

Automating Root Cause Analysis (RCA)

  • Utilising graph-based correlation to map services and infrastructure.
  • Applying ML to infer probable root causes from event chains.
  • Visualising RCA outcomes through topology-aware dashboards.

Remediation and Workflow Automation

  • Integrating with automation platforms, such as Ansible or Rundeck.
  • Triggering automated actions like rollbacks, restarts, or traffic redirection.
  • Auditing and documenting automated interventions for compliance and insight.

Scaling Intelligent AIOps Pipelines

  • Applying MLOps to observability: focusing on retraining and model versioning.
  • Executing real-time predictions across distributed nodes.
  • Best practices for deploying AIOps in production environments.

Case Studies and Practical Applications

  • Analysing real-world incident data using predictive AIOps models.
  • Deploying RCA pipelines using both synthetic and production data.
  • Reviewing industry use cases, including cloud outages, microservices instability, and network degradations.

Summary and Next Steps

Requirements

  • Experience with monitoring systems such as Prometheus or ELK.
  • Proficiency in Python and a foundational understanding of machine learning.
  • Familiarity with incident management workflows.

Audience

  • Senior site reliability engineers (SREs).
  • IT automation architects.
  • DevOps and observability platform leads.

Number of participants


Price per participant

Provisional Upcoming Courses (Require 5+ participants)

Related Categories