Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI in Operations
- The evolution of IT automation: shifting from static runbooks to reasoning agents
- Understanding agent anatomy: the reasoning loop, tool utilization, memory mechanisms, and planning strategies
- Determining when to automate tasks and when to retain human oversight
Agent Frameworks and Architectural Design
- Single-agent patterns: exploring ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm-based patterns
- Comparative analysis of frameworks: LangGraph, CrewAI, AutoGen, and custom agent solutions
- Developing your first operational agent: querying monitoring systems, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to APIs from Prometheus, Grafana, Datadog, and PagerDuty
- Agent-driven log querying: integration with Elasticsearch, Loki, and Splunk
- Leveraging infrastructure tools: executing kubectl, Terraform, and Ansible via agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Automating Incident Response
- Automated incident triage: classifying severity and routing tickets
- Generating root cause hypotheses and gathering supporting evidence
- Executing automated remediations: restarting, scaling, rolling back, and failover actions
- Creating an incident runbook agent with progressive autonomy levels
Safety, Guardrails, and Human-in-the-Loop Mechanisms
- Classifying actions: read-only, low-risk, high-risk, and destructive operations
- Establishing approval gates and escalation policies for critical operations
- Implementing guardrail patterns: action allowlists, blast radius limitations, and rollback guarantees
- Maintaining audit trails and decision provenance for compliance purposes
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose contradictory actions
- Simulating end-to-end major incidents with multi-agent response strategies
Observability and Evaluation
- Tracing agent reasoning chains for debugging and audit purposes
- Evaluating agent decision quality: measuring precision, recall, and time-to-resolution
- Establishing feedback loops: learning from operator overrides and final outcomes
- Tracking costs and analyzing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: utilizing APIs, webhooks, and scheduled jobs
- Rolling out gradual autonomy: transitioning from shadow mode to full auto-remediation
- Creating runbooks for agent failures: procedures for when the agent itself breaks
- Building the business case and measuring ROI for autonomous operations
Requirements
- Professional experience with IT operations, DevOps, or SRE practices.
- Familiarity with Python scripting and REST APIs.
- Basic understanding of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation.
- Platform engineers focused on building self-healing infrastructure.
- IT operations leads evaluating agentic AI for incident management.
14 Hours