Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps using Open Source Tools
- Overview of AIOps core concepts and associated benefits
- The role of Prometheus and Grafana in the observability stack
- The place of ML in AIOps: contrasting predictive and reactive analytics
Configuring Prometheus and Grafana
- Setting up and configuring Prometheus for time series data collection
- Building Grafana dashboards powered by real-time metrics
- Investigating exporters, relabeling mechanisms, and service discovery
Data Preprocessing for Machine Learning
- Extracting and transforming Prometheus metrics
- Structuring datasets for anomaly detection and forecasting tasks
- Leveraging Grafana transformations or Python-based pipelines
Leveraging Machine Learning for Anomaly Detection
- Fundamental ML models for outlier identification (e.g., Isolation Forest, One-Class SVM)
- Training and assessing models on time series datasets
- Visualizing detected anomalies within Grafana dashboards
Forecasting Metrics with ML
- Constructing basic forecasting models (introduction to ARIMA, Prophet, and LSTM)
- Anticipating system load and resource consumption
- Utilizing predictions to inform early alerting and scaling strategies
Integrating ML with Alerting and Automation
- Establishing alert rules based on ML outputs or defined thresholds
- Implementing Alertmanager and managing notification routing
- Initiating scripts or automation workflows upon anomaly detection
Scaling and Operationalizing AIOps
- Connecting external observability solutions (e.g., ELK stack, Moogsoft, Dynatrace)
- Integrating ML models into observability pipelines for production use
- Best practices for managing AIOps at scale
Recap and Future Directions
Requirements
- A solid grasp of system monitoring and observability principles
- Practical experience with Grafana or Prometheus
- Proficiency in Python and foundational knowledge of machine learning concepts
Target Audience
- Observability engineers
- Infrastructure and DevOps teams
- Monitoring platform architects and site reliability engineers (SREs)