Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps using Open Source Tools

  • Overview of AIOps core concepts and associated benefits
  • The role of Prometheus and Grafana in the observability stack
  • The place of ML in AIOps: contrasting predictive and reactive analytics

Configuring Prometheus and Grafana

  • Setting up and configuring Prometheus for time series data collection
  • Building Grafana dashboards powered by real-time metrics
  • Investigating exporters, relabeling mechanisms, and service discovery

Data Preprocessing for Machine Learning

  • Extracting and transforming Prometheus metrics
  • Structuring datasets for anomaly detection and forecasting tasks
  • Leveraging Grafana transformations or Python-based pipelines

Leveraging Machine Learning for Anomaly Detection

  • Fundamental ML models for outlier identification (e.g., Isolation Forest, One-Class SVM)
  • Training and assessing models on time series datasets
  • Visualizing detected anomalies within Grafana dashboards

Forecasting Metrics with ML

  • Constructing basic forecasting models (introduction to ARIMA, Prophet, and LSTM)
  • Anticipating system load and resource consumption
  • Utilizing predictions to inform early alerting and scaling strategies

Integrating ML with Alerting and Automation

  • Establishing alert rules based on ML outputs or defined thresholds
  • Implementing Alertmanager and managing notification routing
  • Initiating scripts or automation workflows upon anomaly detection

Scaling and Operationalizing AIOps

  • Connecting external observability solutions (e.g., ELK stack, Moogsoft, Dynatrace)
  • Integrating ML models into observability pipelines for production use
  • Best practices for managing AIOps at scale

Recap and Future Directions

Requirements

  • A solid grasp of system monitoring and observability principles
  • Practical experience with Grafana or Prometheus
  • Proficiency in Python and foundational knowledge of machine learning concepts

Target Audience

  • Observability engineers
  • Infrastructure and DevOps teams
  • Monitoring platform architects and site reliability engineers (SREs)

Number of participants


Price per participant

Upcoming Courses

Related Categories