Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Cloud Operations on AWS
- Defining operational roles and responsibilities within the cloud
- AWS account structures, organisations, and multi-account strategies
- Essential operational services: CloudWatch, CloudTrail, and AWS Config
Infrastructure as Code and Provisioning
- Core principles of IaC and immutable infrastructure
- Provisioning workflows using Terraform and AWS CloudFormation
- Managing state, modules, and environment promotion processes
CI/CD and Deployment Strategies
- Architecting CI/CD pipelines for cloud-native applications
- Blue/green, canary, and rolling deployment methodologies
- Automating rollback mechanisms, health checks, and release validation
Monitoring, Observability, and Alerting
- Handling metrics, logs, and traces: shipping, storage, and analysis
- Utilising CloudWatch, X-Ray, and third-party observability tools
- Establishing SLOs/SLIs, alerting policies, and on-call procedures
Security Operations and Identity Management
- IAM best practices, least privilege models, and cross-account access
- Managing secrets, KMS, and secure parameter stores
- Operational security measures: patching strategies, vulnerability scanning, and audit trails
Resilience, Backup, and Disaster Recovery
- Designing for fault tolerance and high availability
- Backup strategies, snapshot automation, and restoration procedures
- Disaster recovery planning and the creation of operational runbooks
Cost Optimization and Governance
- Enhancing cost visibility through billing, tagging, and cost allocation strategies
- Rightsizing workloads, reserved instances/savings plans, and budgeting controls
- Governance frameworks: policies, guardrails, and compliance automation
Containers, Serverless, and Runtime Operations
- Operational considerations for ECS, EKS, and Lambda
- Service discovery, autoscaling, and setting resource limits
- Logging, tracing, and debugging containerised workloads
Incident Response, Playbooks, and Chaos Engineering
- Runbook-driven incident response and post-mortem practices
- Automating remediation and implementing self-healing patterns
- Introduction to chaos experiments for validating system resilience
Hands-on Workshop: Operating a Sample Workload
- Deploying a sample application using IaC and a CI/CD pipeline
- Implementing monitoring, alerts, and automated remediation scripts
- Simulating incidents and practising runbook-based response procedures
Summary and Next Steps
Requirements
- Foundational knowledge of cloud concepts and networking
- Proficiency with the Linux command line and scripting
- Experience with source control (Git) and fundamental CI/CD principles
Target Audience
- Cloud operations engineers
- SREs and platform engineers
- DevOps engineers and technical team leads
21 Hours
Testimonials (1)
I've find out new interesting things about Lambda and Serverless