Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Cloud Operations on AWS
- Defining operational roles and responsibilities in a cloud context
- Structuring AWS accounts, organizations, and multi-account strategies
- Leveraging core operational services: CloudWatch, CloudTrail, and AWS Config
Infrastructure as Code and Provisioning
- Understanding IaC principles and the concept of immutable infrastructure
- Executing provisioning tasks using Terraform and AWS CloudFormation
- Managing state, modular components, and environment promotion workflows
CI/CD and Deployment Strategies
- Architecting CI/CD pipelines optimized for cloud-native applications
- Implementing blue/green, canary, and rolling deployment models
- Automating rollbacks, health checks, and release validation processes
Monitoring, Observability, and Alerting
- Handling metrics, logs, and traces: ingestion, storage, and analysis
- Utilizing CloudWatch, X-Ray, and third-party observability solutions
- Establishing SLOs/SLIs, alerting policies, and on-call protocols
Security Operations and Identity Management
- Applying IAM best practices, least privilege principles, and cross-account access controls
- Managing secrets, KMS, and secure parameter stores
- Maintaining operational security through patching strategies, vulnerability scanning, and audit trails
Resilience, Backup, and Disaster Recovery
- Designing systems for fault tolerance and high availability
- Developing backup strategies, automating snapshots, and executing restore procedures
- Planning disaster recovery and creating comprehensive runbooks
Cost Optimization and Governance
- Gaining cost visibility through billing, tagging, and cost allocation strategies
- Rightsizing resources, utilizing reserved instances/savings plans, and enforcing budget controls
- Implementing governance via policies, guardrails, and compliance automation
Containers, Serverless, and Runtime Operations
- Addressing operational considerations for ECS, EKS, and Lambda
- Managing service discovery, autoscaling, and resource limits
- Logging, tracing, and debugging containerized workloads
Incident Response, Playbooks, and Chaos Engineering
- Executing runbook-driven incident response and conducting postmortems
- Automating remediation and implementing self-healing patterns
- Introduction to chaos experiments for validating system resilience
Hands-on Workshop: Operating a Sample Workload
- Deploying a sample application utilizing IaC and a CI/CD pipeline
- Setting up monitoring, alerts, and automated remediation scripts
- Simulating incidents to practice runbook-based response mechanisms
Summary and Next Steps
Requirements
- A fundamental grasp of cloud concepts and networking principles
- Proficiency with the Linux command line and scripting capabilities
- Practical experience with source control (Git) and a basic understanding of CI/CD workflows
Target Audience
- Cloud operations engineers
- Site Reliability Engineers (SREs) and platform engineers
- DevOps engineers and technical team leads
21 Hours
Testimonials (1)
I've find out new interesting things about Lambda and Serverless