Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps
- Defining AIOps and its importance
- Contrasting traditional monitoring with AIOps-driven observability
- Understanding AIOps architecture and key components
Collecting and Normalising Operational Data
- Types of observability data: metrics, logs, and traces
- Ingesting data from diverse sources (servers, containers, cloud)
- Utilising agents and exporters (Prometheus, Beats, Fluentd)
Data Correlation and Anomaly Detection
- Time series correlation and statistical methodologies
- Applying ML models for anomaly detection
- Identifying incidents within distributed systems
Alerting and Noise Reduction
- Designing intelligent alert rules and thresholds
- Implementing suppression, deduplication, and alert grouping
- Integrating with Alertmanager, Slack, PagerDuty, or Opsgenie
Root Cause Analysis and Visualisation
- Using dashboards to visualise metrics and detect trends
- Analysing events and timelines for RCA
- Tracing issues across layers with distributed tracing tools
Automation and Remediation
- Triggering automated scripts or workflows from incidents
- Integrating with ITSM systems (ServiceNow, Jira)
- Use cases: self-healing, scaling, traffic rerouting
Open Source and Commercial AIOps Platforms
- Overview of tools: Prometheus, Grafana, ELK, Moogsoft, Dynatrace
- Evaluation criteria for selecting an AIOps platform
- Demonstration and hands-on practice with a selected stack
Summary and Next Steps
Requirements
- A solid understanding of IT operations and system monitoring concepts
- Experience working with monitoring tools or dashboards
- Familiarity with standard log and metric formats
Audience
- Operations teams managing infrastructure and applications
- Site Reliability Engineers (SREs)
- Teams dedicated to IT monitoring and observability