Production & Scaling

High availability, disaster recovery, chaos engineering, and the ops side of running things.

Duolingo for DevOps. One concept = one 2-4 minute video. Never introduce a concept before its prerequisites. Every lesson builds on the last.

0 / 21 lessons readyaudience: sre21 lessons planned
  1. 1. Production, Scaling & DR

    21 lessons

    The last mile from 'it works in staging' to 'it survives a bad Tuesday.'

    High Availability Fundamentals

    • What HA actually means3m 30sComing soon
    • Stateless-first design4m 00sComing soon
    • Multi-AZ deployment4m 00sComing soon
    • Load balancer HA4m 00sComing soon
    • Graceful shutdown and termination handlers4m 00sComing soon
    • Zero-downtime schema migrations4m 00sComing soon

    Disaster Recovery

    • RPO and RTO — the DR vocabulary3m 30sComing soon
    • 3-2-1 backup strategy3m 30sComing soon
    • Pilot light, warm standby, multi-site4m 00sComing soon
    • Running a DR drill4m 00sComing soon
    • Velero — Kubernetes backup and restore4m 00sComing soon
    • Database backup and restore drills4m 00sComing soon

    Cost Optimization

    • Where cloud money actually goes3m 30sComing soon
    • Right-sizing nodes and pods4m 00sComing soon
    • Spot instances at scale4m 00sComing soon
    • Kubecost — attributing cost to workloads4m 00sComing soon
    • Reserved vs Savings Plans4m 00sComing soon

    Performance

    • p50, p95, p99 — reading latency4m 00sComing soon
    • Flame graphs — profiling intuitively4m 00sComing soon
    • Capacity planning for scaling events4m 00sComing soon
    • Load testing with k64m 00sComing soon