Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps with Open Source Tools
- An overview of AIOps concepts and their organisational benefits
- The role of Prometheus and Grafana within the observability stack
- Positioning machine learning in AIOps: contrasting predictive and reactive analytics
Setting Up Prometheus and Grafana
- Installing and configuring Prometheus for effective time series data collection
- Building dashboards in Grafana using real-time metrics
- Exploring exporters, relabeling, and service discovery mechanisms
Data Preprocessing for Machine Learning
- Extracting and transforming metrics from Prometheus
- Preparing datasets suitable for anomaly detection and forecasting models
- Leveraging Grafana’s native transformations or Python-based pipelines
Applying Machine Learning for Anomaly Detection
- Implementing basic ML models for outlier detection (e.g., Isolation Forest, One-Class SVM)
- Training and evaluating models using time series data
- Visualising detected anomalies within Grafana dashboards
Forecasting Metrics with Machine Learning
- Developing simple forecasting models (introductory ARIMA, Prophet, and LSTM)
- Predicting system load and resource utilisation patterns
- Utilising predictions for early alerting and scaling decisions
Integrating Machine Learning with Alerting and Automation
- Defining alert rules based on machine learning outputs or defined thresholds
- Configuring Alertmanager and routing notifications effectively
- Triggering scripts or automation workflows upon anomaly detection
Scaling and Operationalising AIOps
- Integrating external observability tools (e.g., ELK stack, Moogsoft, Dynatrace)
- Operationalising ML models within observability pipelines
- Best practices for implementing AIOps at scale
Summary and Next Steps
Requirements
- A solid understanding of system monitoring and observability principles.
- Practical experience working with Grafana or Prometheus.
- Familiarity with Python and fundamental machine learning concepts.
Target Audience
- Observability Engineers.
- Infrastructure and DevOps Teams.
- Monitoring Platform Architects and Site Reliability Engineers (SREs).