Production & Scaling
High availability, disaster recovery, chaos engineering, and the ops side of running things.
“Duolingo for DevOps. One concept = one 2-4 minute video. Never introduce a concept before its prerequisites. Every lesson builds on the last.”
0 / 21 lessons readyaudience: sre21 lessons planned
1. Production, Scaling & DR
21 lessonsThe last mile from 'it works in staging' to 'it survives a bad Tuesday.'
High Availability Fundamentals
- What HA actually means3m 30sComing soon
- Stateless-first design4m 00sComing soon
- Multi-AZ deployment4m 00sComing soon
- Load balancer HA4m 00sComing soon
- Graceful shutdown and termination handlers4m 00sComing soon
- Zero-downtime schema migrations4m 00sComing soon
Disaster Recovery
- RPO and RTO — the DR vocabulary3m 30sComing soon
- 3-2-1 backup strategy3m 30sComing soon
- Pilot light, warm standby, multi-site4m 00sComing soon
- Running a DR drill4m 00sComing soon
- Velero — Kubernetes backup and restore4m 00sComing soon
- Database backup and restore drills4m 00sComing soon
Cost Optimization
- Where cloud money actually goes3m 30sComing soon
- Right-sizing nodes and pods4m 00sComing soon
- Spot instances at scale4m 00sComing soon
- Kubecost — attributing cost to workloads4m 00sComing soon
- Reserved vs Savings Plans4m 00sComing soon
Performance
- p50, p95, p99 — reading latency4m 00sComing soon
- Flame graphs — profiling intuitively4m 00sComing soon
- Capacity planning for scaling events4m 00sComing soon
- Load testing with k64m 00sComing soon