Modern DevOps Engineering
From zero to production-ready DevOps Engineer, one concept at a time.
“Duolingo for DevOps. One concept = one 2-4 minute video. Never introduce a concept before its prerequisites. Every lesson builds on the last.”
1. Networking
40 lessonsEvery distributed system talks over the network. You cannot debug Kubernetes, an Ingress, or an outage without understanding what's actually happening between machines.
Networking Fundamentals
- What is the Internet, really?1m 49sWatch →
- IP addresses — every device's phone number2m 24sWatch →
- The client-server model1m 48sWatch →
- DNS — the internet's phonebook2m 27sWatch →
- Ports — apartment numbers on an IP2m 26sWatch →
- The OSI model — layers, intuitively3m 56sWatch →
- MAC vs IP addresses2m 41sWatch →
- Subnet masks and CIDR notation2m 40sWatch →
TCP/IP
DNS Deep Dive
HTTP
- What happens in that half-second URL load2m 10sWatch →
- GET, POST, PUT, DELETE — HTTP verbs3m 24sWatch →
- HTTP status codes decoded4m 15sWatch →
- HTTP headers — the metadata layer3m 26sWatch →
- HTTP vs HTTPS2m 10sWatch →
- HTTP/1.1 vs HTTP/2 vs HTTP/32m 59sWatch →
- CORS — same-origin, cross-origin, and preflight3m 28sWatch →
TLS/SSL
Firewalls, NAT, Routing
- What a firewall actually does3m 30sWatch →
- iptables basics4m 00sComing soon
- nftables — iptables' modern successor3m 30sComing soon
- NAT — how one public IP serves many devices3m 16sWatch →
- Load balancers explained3m 06sWatch →
- Reverse proxies vs load balancers3m 41sWatch →
- VPNs — a private tunnel over a public net3m 55sWatch →
2. Git
20 lessonsEvery DevOps workflow starts with git. Not a git internals course — just the mental models that let you use it confidently and unblock yourself when things go wrong.
The Git Model
- Why version control exists3m 08sWatch →
- Working dir → staging → commit1m 47sWatch →
- git init vs git clone2m 27sWatch →
- git log and git diff2m 42sWatch →
- .gitignore — what NOT to commit3m 29sWatch →
- Branches — parallel realities3m 12sWatch →
- git merge — combining branches3m 04sWatch →
- Merge vs rebase2m 43sWatch →
Working with Others
3. Docker
47 lessonsContainers are the fundamental packaging unit of modern deployment. Everything after (Kubernetes, CI/CD, Helm) assumes you understand what a container is and why it matters.
Why Docker Exists
The Dockerfile
- Your first Dockerfile2m 31sWatch →
- FROM, RUN, COPY — the essentials3m 04sWatch →
- WORKDIR and ENV2m 49sWatch →
- CMD vs ENTRYPOINT3m 05sWatch →
- EXPOSE and LABEL2m 31sWatch →
- Multi-stage builds — smaller images4m 00sComing soon
Running Containers
- docker run — the anatomy3m 09sWatch →
- docker ps, stop, rm2m 39sWatch →
- docker exec — getting a shell in a container3m 00sComing soon
- docker logs — reading container output2m 30sComing soon
- docker inspect — the full metadata dump3m 00sComing soon
Volumes and Data
- Volumes — data that outlives the container3m 00sComing soon
- Bind mounts vs named volumes3m 30sComing soon
- Volume drivers3m 30sComing soon
- Backing up volumes3m 30sComing soon
- tmpfs mounts3m 00sComing soon
Docker Networks
- Docker networks — the big picture3m 30sComing soon
- Bridge network — the default3m 30sComing soon
- Host network mode3m 00sComing soon
- Overlay networks4m 00sComing soon
- -p and port publishing3m 00sComing soon
Docker Compose
- What is docker-compose?3m 00sComing soon
- services — the compose primitive3m 00sComing soon
- How compose wires networks3m 30sComing soon
- Volumes in compose3m 00sComing soon
- Env files and .env3m 00sComing soon
- depends_on — startup order3m 00sComing soon
- Compose profiles3m 00sComing soon
Registries and Distribution
- Docker Hub — the default registry3m 00sComing soon
- docker login and docker push3m 00sComing soon
- Running a private registry3m 30sComing soon
- Image tags and semver discipline3m 30sComing soon
- Image signing — verifying what you pull4m 00sComing soon
Best Practices and Security
- Slimming an image — alpine, distroless3m 30sComing soon
- Non-root user in containers3m 00sComing soon
- Read-only filesystems3m 30sComing soon
- Image vulnerability scanning3m 30sComing soon
- Never bake secrets into images3m 30sComing soon
- CPU and memory limits at runtime3m 30sComing soon
- HEALTHCHECK — proving your container is alive3m 30sComing soon
- Mini project: containerize a Node app5m 00sComing soon
4. Kubernetes Core
65 lessonsKubernetes is where most beginners drown. This module carves it into tiny lessons — one primitive at a time, each with a concrete visual mental model.
Why Kubernetes Exists
- What problem does Kubernetes solve?2m 30sComing soon
- From Borg to Kubernetes — the lineage3m 00sComing soon
- Desired state vs actual state3m 30sComing soon
- Declarative vs imperative3m 00sComing soon
Cluster Architecture
- Cluster, nodes, control plane3m 30sComing soon
- The API server — Kubernetes' front door3m 30sComing soon
- etcd — the cluster's source of truth3m 30sComing soon
- The scheduler — placing pods on nodes3m 30sComing soon
- Controller manager and control loops3m 30sComing soon
- kubelet — the node agent3m 00sComing soon
- kube-proxy and iptables rules3m 30sComing soon
- CRI — how kubelet runs containers3m 30sComing soon
YAML and Manifests
- YAML basics3m 30sComing soon
- Manifest structure — apiVersion, kind, metadata, spec3m 30sComing soon
- kubectl apply — declarative in action3m 00sComing soon
- Labels and selectors3m 30sComing soon
- Annotations vs labels3m 00sComing soon
Pods
- What is a Pod?2m 30sComing soon
- Pod lifecycle — pending, running, succeeded, failed3m 30sComing soon
- Multi-container pods3m 30sComing soon
- Init containers3m 30sComing soon
- Sidecar containers3m 30sComing soon
- Liveness, readiness, startup probes4m 00sComing soon
- resources.requests vs resources.limits3m 30sComing soon
Deployments and Workloads
- ReplicaSets3m 00sComing soon
- Deployments — desired state3m 30sComing soon
- Rolling updates without downtime4m 00sComing soon
- Rollback in one command3m 00sComing soon
- DaemonSets — one pod per node3m 30sComing soon
- Jobs — run to completion3m 30sComing soon
- CronJobs — jobs on a schedule3m 30sComing soon
Kubernetes Networking
- Services — a stable address for pods3m 30sComing soon
- ClusterIP — the default3m 00sComing soon
- NodePort — exposing on every node3m 00sComing soon
- LoadBalancer service3m 30sComing soon
- Headless services3m 30sComing soon
- Endpoints and EndpointSlices3m 30sComing soon
- Ingress — routing traffic in3m 30sComing soon
- Ingress controllers — nginx, traefik3m 30sComing soon
- CoreDNS and service discovery3m 30sComing soon
- Network policies — firewalls for pods4m 00sComing soon
ConfigMaps and Secrets
- ConfigMaps — config as data3m 30sComing soon
- Secrets — sensitive config3m 30sComing soon
- envFrom and env — injecting into pods3m 30sComing soon
- Config as mounted files3m 30sComing soon
- Secret types — Opaque, TLS, dockerconfigjson3m 30sComing soon
- Rotating secrets safely4m 00sComing soon
Storage
- Volumes in Kubernetes3m 30sComing soon
- PersistentVolumes4m 00sComing soon
- PersistentVolumeClaims3m 30sComing soon
- StorageClasses and dynamic provisioning4m 00sComing soon
- CSI — the storage plug-in interface4m 00sComing soon
- Reclaim policies — delete vs retain3m 30sComing soon
Namespaces and RBAC
- Namespaces — logical isolation3m 00sComing soon
- ResourceQuotas per namespace3m 30sComing soon
- LimitRanges3m 30sComing soon
- ServiceAccounts3m 30sComing soon
- RBAC — Role and RoleBinding4m 00sComing soon
- ClusterRole vs Role3m 30sComing soon
Debugging Kubernetes
- kubectl get and kubectl describe3m 30sComing soon
- kubectl logs — and when they lie3m 30sComing soon
- kubectl exec — a shell in a pod3m 30sComing soon
- kubectl get events — the debug goldmine3m 30sComing soon
- kubectl port-forward for local debugging3m 30sComing soon
- CrashLoopBackOff — reading the signal4m 00sComing soon
5. Kubernetes Advanced
34 lessonsOnce you can deploy a stateless service, the next tier: scheduling, autoscaling, stateful workloads, and extending Kubernetes itself.
Advanced Scheduling
- Node affinity — where should this pod land?4m 00sComing soon
- Pod affinity and anti-affinity4m 00sComing soon
- Taints and tolerations4m 00sComing soon
- PriorityClasses and preemption4m 00sComing soon
- Topology spread constraints4m 00sComing soon
- PodDisruptionBudgets4m 00sComing soon
Autoscaling
- metrics-server — the source of truth3m 30sComing soon
- Horizontal Pod Autoscaler4m 00sComing soon
- Vertical Pod Autoscaler4m 00sComing soon
- Cluster Autoscaler4m 00sComing soon
- KEDA — event-driven autoscaling4m 00sComing soon
- QoS classes and eviction4m 00sComing soon
StatefulSets and Storage
- StatefulSets vs Deployments4m 00sComing soon
- Stable network identities for pods4m 00sComing soon
- PVC-to-PV binding rules4m 00sComing soon
- Ordered pod startup and shutdown4m 00sComing soon
- Backup and restore patterns4m 00sComing soon
- Should you run a database in Kubernetes?4m 00sComing soon
Operators and CRDs
- Custom Resource Definitions4m 00sComing soon
- The Operator pattern4m 00sComing soon
- operator-sdk and kubebuilder4m 00sComing soon
- Finalizers — cleanup on delete4m 00sComing soon
- Admission controllers4m 00sComing soon
- Mutating vs validating webhooks4m 00sComing soon
- Gatekeeper and OPA policies4m 00sComing soon
Service Mesh
- What is a service mesh?4m 00sComing soon
- Istio — a quick tour4m 00sComing soon
- Linkerd — the lightweight alternative4m 00sComing soon
- Envoy sidecars4m 00sComing soon
- mTLS between services4m 00sComing soon
- Canary and traffic splitting4m 00sComing soon
Gateway API and Modern Routing
- Gateway API vs Ingress4m 00sComing soon
- HTTPRoute and TCPRoute4m 00sComing soon
- Multi-cluster Kubernetes — the shape4m 00sComing soon
6. Helm
15 lessonsManaging raw YAML at scale is unmaintainable. Helm packages Kubernetes apps into reusable, versioned charts.
Charts and Templates
- What is Helm and why?3m 00sComing soon
- Chart anatomy — Chart.yaml, values.yaml, templates/4m 00sComing soon
- Go templates in Helm4m 00sComing soon
- Pipelines, functions, and filters4m 00sComing soon
- _helpers.tpl — DRY templates4m 00sComing soon
Values, Environments, Lifecycle
- Layering values across environments4m 00sComing soon
- helm install and helm upgrade3m 30sComing soon
- helm rollback3m 30sComing soon
- Helm hooks for pre-install, post-install4m 00sComing soon
- Chart dependencies and subcharts4m 00sComing soon
Publishing Charts
- Helm repositories3m 30sComing soon
- OCI registries for Helm4m 00sComing soon
- ChartMuseum vs OCI3m 30sComing soon
- Chart testing pre-release4m 00sComing soon
- Mini project: chart your own app5m 00sComing soon
7. Terraform
26 lessonsClick-ops doesn't scale. Terraform is the industry standard for describing infrastructure declaratively.
Infrastructure as Code
- What is Infrastructure as Code?3m 00sComing soon
- Declarative vs imperative IaC3m 00sComing soon
- Terraform vs Pulumi vs CloudFormation3m 30sComing soon
HCL Basics
- HCL syntax quick tour3m 30sComing soon
- Providers — how Terraform talks to APIs3m 30sComing soon
- resource blocks3m 30sComing soon
- Input variables3m 30sComing soon
- Outputs — passing values out3m 00sComing soon
- Data sources — reading existing infra3m 30sComing soon
- locals — internal DRY3m 00sComing soon
State
- What terraform.tfstate actually is4m 00sComing soon
- State backends — local, S3, GCS4m 00sComing soon
- State locking with DynamoDB4m 00sComing soon
- terraform import — adopting existing resources4m 00sComing soon
- terraform state mv, rm — surgery on state4m 00sComing soon
- Workspaces — dev, staging, prod4m 00sComing soon
Modules
- Modules — reusable building blocks4m 00sComing soon
- Terraform Registry3m 30sComing soon
- Versioning your modules4m 00sComing soon
- Composition patterns — thin root modules4m 00sComing soon
- Testing modules with terratest4m 00sComing soon
Production Patterns
- plan → apply — the safe workflow4m 00sComing soon
- Terraform in CI/CD4m 00sComing soon
- Terragrunt — Terraform at scale4m 00sComing soon
- Drift detection4m 00sComing soon
- Mini project: three-tier AWS stack5m 00sComing soon
8. AWS
49 lessonsAWS is the default cloud in most DevOps interviews. Not a certification prep course — the concepts you need to design and operate production infra.
AWS Fundamentals
- Regions, Availability Zones, Edge Locations3m 30sComing soon
- Accounts and Organizations3m 30sComing soon
- The shared responsibility model3m 00sComing soon
- AWS billing basics3m 30sComing soon
EC2 and Compute
- EC2 — a VM in the cloud3m 30sComing soon
- Instance types decoded4m 00sComing soon
- AMIs — VM templates3m 30sComing soon
- Security groups — stateful firewalls4m 00sComing soon
- On-demand vs Spot vs Reserved4m 00sComing soon
- Auto Scaling Groups4m 00sComing soon
VPC and Networking
- What is a VPC?3m 30sComing soon
- Public vs private subnets4m 00sComing soon
- Internet Gateway and NAT Gateway4m 00sComing soon
- Route tables3m 30sComing soon
- VPC peering vs Transit Gateway4m 00sComing soon
- VPC endpoints — private AWS API access4m 00sComing soon
- NACLs vs Security Groups4m 00sComing soon
S3
- S3 — object storage3m 30sComing soon
- Buckets, keys, prefixes3m 30sComing soon
- Storage classes — Standard, IA, Glacier4m 00sComing soon
- Versioning and lifecycle rules3m 30sComing soon
- Pre-signed URLs3m 30sComing soon
IAM
- IAM — the who, what, and where of AWS3m 30sComing soon
- Users vs Groups vs Roles4m 00sComing soon
- Policies — JSON, actions, resources4m 00sComing soon
- Managed vs inline policies3m 30sComing soon
- AssumeRole and cross-account access4m 00sComing soon
- Least-privilege in practice4m 00sComing soon
- Permission boundaries4m 00sComing soon
RDS and Databases
- RDS — managed databases3m 30sComing soon
- Multi-AZ vs Read Replicas4m 00sComing soon
- DynamoDB — key-value at scale4m 00sComing soon
- RDS backups and snapshots3m 30sComing soon
- Parameter groups and tuning4m 00sComing soon
EKS
- EKS — managed Kubernetes3m 30sComing soon
- EKS vs self-managed k8s3m 30sComing soon
- Node groups — managed and self-managed4m 00sComing soon
- EKS on Fargate4m 00sComing soon
- IRSA — IAM roles for service accounts4m 00sComing soon
- AWS Load Balancer Controller4m 00sComing soon
- EKS add-ons3m 30sComing soon
CloudWatch
- CloudWatch — metrics, logs, alarms3m 30sComing soon
- Log groups and log streams3m 30sComing soon
- CloudWatch alarms and actions4m 00sComing soon
- CloudWatch Logs Insights4m 00sComing soon
Route 53 and DNS
- Route 53 — AWS DNS3m 30sComing soon
- Hosted zones3m 30sComing soon
- Routing policies — simple, weighted, latency4m 00sComing soon
- Health checks and failover4m 00sComing soon
9. CI/CD
37 lessonsContinuous integration and delivery is where DevOps meets developer experience. Every one of these tools has the same shape — trigger, pipeline, jobs — but different quirks.
CI/CD Concepts
- Continuous integration vs delivery vs deployment3m 30sComing soon
- Pipeline anatomy — triggers, stages, jobs3m 30sComing soon
- Pipeline-as-code — why YAML wins3m 00sComing soon
- Artifacts and registries3m 30sComing soon
- Semver and release notes3m 30sComing soon
GitHub Actions
- GitHub Actions — the mental model3m 30sComing soon
- Workflow file anatomy3m 30sComing soon
- Events and triggers4m 00sComing soon
- Runners — hosted vs self-hosted4m 00sComing soon
- Using actions from the marketplace3m 30sComing soon
- Secrets and environments4m 00sComing soon
- Matrix builds4m 00sComing soon
- Reusable workflows4m 00sComing soon
- Composite actions4m 00sComing soon
- OIDC to AWS — no long-lived keys4m 00sComing soon
GitLab CI
- GitLab CI — the shape3m 30sComing soon
- Stages and jobs3m 30sComing soon
- Runners in GitLab4m 00sComing soon
- DAG pipelines with needs:4m 00sComing soon
- Environments and deployments4m 00sComing soon
ArgoCD
- ArgoCD — pull-based deployment3m 30sComing soon
- Applications and app-of-apps4m 00sComing soon
- Sync policies and auto-sync4m 00sComing soon
- Custom health checks4m 00sComing soon
- ArgoCD Projects and RBAC4m 00sComing soon
- Argo Notifications3m 30sComing soon
- Argo Image Updater4m 00sComing soon
FluxCD
- Flux — the other GitOps engine3m 30sComing soon
- Sources and Kustomizations4m 00sComing soon
- Flux vs ArgoCD — pick one4m 00sComing soon
- Flux image automation4m 00sComing soon
- Flux notifications3m 30sComing soon
Deployment Strategies
- Rolling deployment3m 30sComing soon
- Blue-green deployment4m 00sComing soon
- Canary deployment4m 00sComing soon
- Shadow deployment4m 00sComing soon
- Feature flags — decoupling deploy from release4m 00sComing soon
10. Monitoring
30 lessonsYou can't fix what you can't see. Metrics, alerting, and dashboards — the reactive half of observability.
Observability Concepts
- Metrics, logs, traces — the three pillars3m 30sComing soon
- What is a metric?3m 00sComing soon
- Counters, gauges, histograms4m 00sComing soon
- Pull-based vs push-based metrics3m 30sComing soon
- Cardinality — the metric budget4m 00sComing soon
Prometheus
- Prometheus — the reference model3m 30sComing soon
- scrape_configs — telling Prometheus what to poll4m 00sComing soon
- Exporters — node_exporter, cAdvisor, kube-state4m 00sComing soon
- Service discovery in Kubernetes4m 00sComing soon
- PromQL basics4m 00sComing soon
- rate, irate, and the common gotchas4m 00sComing soon
- Alertmanager4m 00sComing soon
- Federation and Thanos4m 00sComing soon
- Prometheus HA patterns4m 00sComing soon
- Recording rules — pre-computing4m 00sComing soon
Grafana
- Grafana — the eyes on the data3m 30sComing soon
- Data sources3m 30sComing soon
- Panels and visualizations4m 00sComing soon
- Dashboard variables4m 00sComing soon
- Dashboard provisioning as code4m 00sComing soon
- Grafana alerts vs Alertmanager4m 00sComing soon
- Annotations — marking deploys4m 00sComing soon
Alerting Discipline
- Alert fatigue and how to avoid it3m 30sComing soon
- Symptom-based vs cause-based alerts4m 00sComing soon
- SLOs and error budgets4m 00sComing soon
- On-call rotation basics3m 30sComing soon
- Incident playbooks4m 00sComing soon
- Routing and inhibition rules4m 00sComing soon
- Wiring alerts to Slack + PagerDuty3m 30sComing soon
- Postmortem culture4m 00sComing soon
11. Logging
15 lessonsStructured logging is a superpower. Loki is Prometheus-for-logs. The DevOps job is often reading logs from many services at once.
Log Fundamentals
- What makes a good log line?3m 30sComing soon
- Structured logs — JSON everywhere3m 30sComing soon
- Log levels used correctly3m 30sComing soon
- Correlation IDs across services4m 00sComing soon
- Where logs go — stdout, file, sink3m 30sComing soon
Loki
- Loki — the mental model3m 30sComing soon
- Promtail — the log shipper4m 00sComing soon
- LogQL basics4m 00sComing soon
- Labels done right in Loki4m 00sComing soon
- Retention and storage classes4m 00sComing soon
- Loki HA architecture4m 00sComing soon
Aggregation Patterns
- Fluent Bit — the swiss army log agent4m 00sComing soon
- The ELK stack briefly4m 00sComing soon
- Multi-tenant logging pitfalls4m 00sComing soon
- PII in logs — the compliance minefield4m 00sComing soon
12. Tracing & OpenTelemetry
16 lessonsDistributed tracing turns 'the request was slow' into 'the DB call in service B took 3 seconds.' OpenTelemetry is the vendor-neutral standard.
Tracing Concepts
- Traces, spans, and context propagation4m 00sComing soon
- Sampling — head vs tail4m 00sComing soon
- Baggage and cross-cutting attributes4m 00sComing soon
Tempo
- Tempo — traces at scale4m 00sComing soon
- TraceQL basics4m 00sComing soon
- Tempo + Grafana integration4m 00sComing soon
OpenTelemetry
- OpenTelemetry — one API to rule them all4m 00sComing soon
- SDK vs API vs collector4m 00sComing soon
- Auto-instrumentation for common runtimes4m 00sComing soon
- The Collector — receivers, processors, exporters4m 00sComing soon
- Context propagation across languages4m 00sComing soon
- Semantic conventions — why they matter4m 00sComing soon
SLIs, SLOs, SLAs
- What makes a good SLI?3m 30sComing soon
- SLOs — the reliability target4m 00sComing soon
- SLA vs SLO — legal vs engineering3m 30sComing soon
- Error budget policies4m 00sComing soon
13. Security
29 lessonsSecurity is a shared responsibility, but DevOps often owns the infra security posture — RBAC, secrets, network policies, image supply chain.
Container Security
- The image supply chain4m 00sComing soon
- Trivy — scanning images and IaC4m 00sComing soon
- Grype + Syft for SBOMs4m 00sComing soon
- Signing images with cosign4m 00sComing soon
- Distroless base images4m 00sComing soon
- Falco — runtime security4m 00sComing soon
- seccomp and AppArmor profiles4m 00sComing soon
RBAC Deep Dive
- RBAC verbs and resources4m 00sComing soon
- RoleBindings vs ClusterRoleBindings4m 00sComing soon
- kubectl auth can-i3m 30sComing soon
- User impersonation for testing4m 00sComing soon
- Auditing RBAC in production4m 00sComing soon
Secrets Management
- Secrets vs env vars vs config3m 30sComing soon
- Sealed Secrets — commit-safe secrets4m 00sComing soon
- SOPS — encrypted YAML4m 00sComing soon
- External Secrets Operator4m 00sComing soon
- Secret rotation patterns4m 00sComing soon
- SSH key management at scale4m 00sComing soon
HashiCorp Vault
- Vault — the mental model4m 00sComing soon
- Auth methods — how Vault knows who you are4m 00sComing soon
- KV engine — static secrets4m 00sComing soon
- Dynamic secrets — the killer feature4m 00sComing soon
- Transit engine — encryption as a service4m 00sComing soon
- Vault policies4m 00sComing soon
- Vault HA and Raft storage4m 00sComing soon
Network Policies and Zero Trust
- Default-deny network policies4m 00sComing soon
- Ingress vs egress rules4m 00sComing soon
- Cilium — eBPF-powered network policies4m 00sComing soon
- Zero-trust in Kubernetes4m 00sComing soon
14. GitOps
15 lessonsGitOps unifies the philosophy: git is the source of truth, an operator continuously reconciles cluster state.
GitOps Philosophy
- What GitOps actually means4m 00sComing soon
- Pull-based vs push-based deployment4m 00sComing soon
- Declarative state as source of truth4m 00sComing soon
- The reconciliation loop4m 00sComing soon
- The four GitOps principles4m 00sComing soon
GitOps in Practice
- Mono-repo vs multi-repo for GitOps4m 00sComing soon
- Environment promotion — dev → staging → prod4m 00sComing soon
- The config-repo pattern4m 00sComing soon
- Secrets in a GitOps world4m 00sComing soon
- PR-driven deploys4m 00sComing soon
- Rollback via git revert4m 00sComing soon
- Progressive delivery with Flagger4m 00sComing soon
Multi-Environment
- Kustomize overlays for envs4m 00sComing soon
- Multi-cluster GitOps patterns4m 00sComing soon
- Bootstrap pattern for new clusters4m 00sComing soon
15. Production, Scaling & DR
21 lessonsThe last mile from 'it works in staging' to 'it survives a bad Tuesday.'
High Availability Fundamentals
- What HA actually means3m 30sComing soon
- Stateless-first design4m 00sComing soon
- Multi-AZ deployment4m 00sComing soon
- Load balancer HA4m 00sComing soon
- Graceful shutdown and termination handlers4m 00sComing soon
- Zero-downtime schema migrations4m 00sComing soon
Disaster Recovery
- RPO and RTO — the DR vocabulary3m 30sComing soon
- 3-2-1 backup strategy3m 30sComing soon
- Pilot light, warm standby, multi-site4m 00sComing soon
- Running a DR drill4m 00sComing soon
- Velero — Kubernetes backup and restore4m 00sComing soon
- Database backup and restore drills4m 00sComing soon
Cost Optimization
- Where cloud money actually goes3m 30sComing soon
- Right-sizing nodes and pods4m 00sComing soon
- Spot instances at scale4m 00sComing soon
- Kubecost — attributing cost to workloads4m 00sComing soon
- Reserved vs Savings Plans4m 00sComing soon
Performance
- p50, p95, p99 — reading latency4m 00sComing soon
- Flame graphs — profiling intuitively4m 00sComing soon
- Capacity planning for scaling events4m 00sComing soon
- Load testing with k64m 00sComing soon
16. Real Projects
39 lessonsNothing lands the concepts like building the actual stack end-to-end. Five projects, each broken into lessons that reinforce prior modules.
Project 1: Deploy a Static Website
- Project scope — what we're building3m 00sComing soon
- Containerize a static site4m 00sComing soon
- Push image to a registry3m 30sComing soon
- Write the Deployment YAML4m 00sComing soon
- Expose it with a Service4m 00sComing soon
- Add an Ingress and TLS4m 00sComing soon
- CI pipeline to auto-deploy on push4m 00sComing soon
- Add basic monitoring4m 00sComing soon
Project 2: Backend API + PostgreSQL
- Project scope3m 00sComing soon
- Dockerize the API4m 00sComing soon
- Local dev with docker-compose4m 00sComing soon
- Managed Postgres in the cloud4m 00sComing soon
- Wire secrets via External Secrets4m 00sComing soon
- Zero-downtime migrations4m 00sComing soon
- HPA on request rate4m 00sComing soon
- Real health probes for the API4m 00sComing soon
- Expose Prometheus metrics4m 00sComing soon
- Write the first alert + playbook4m 00sComing soon
Project 3: Microservices Demo
- Project scope — 3-service split3m 30sComing soon
- Deciding service boundaries4m 00sComing soon
- gRPC vs HTTP between services4m 00sComing soon
- Service discovery in Kubernetes4m 00sComing soon
- Wire OpenTelemetry across services4m 00sComing soon
- Circuit breakers and timeouts4m 00sComing soon
- Adding a message queue4m 00sComing soon
- Canary rollout with Argo Rollouts4m 00sComing soon
- Incident walkthrough — a real bug5m 00sComing soon
Project 4: Production SaaS
- Project scope — multi-tenant SaaS4m 00sComing soon
- Multi-env infrastructure with Terraform5m 00sComing soon
- Bootstrap ArgoCD across envs4m 00sComing soon
- Ship the app as a Helm chart4m 00sComing soon
- Vault-backed prod secrets4m 00sComing soon
- Domains, DNS, TLS in prod4m 00sComing soon
- Full observability stack — Prom, Loki, Tempo5m 00sComing soon
- Define real SLOs and burn alerts4m 00sComing soon
- Backups and DR drill4m 00sComing soon
- Cost review + tuning4m 00sComing soon
- Scale test — find the ceiling5m 00sComing soon
- The production launch checklist5m 00sComing soon
17. Interview Preparation
19 lessonsThe interview is a separate skill. Behavioral answers, system design, and hands-on debugging drills.
Behavioral
- STAR method for behavioral answers3m 30sComing soon
- Telling a good production-incident story4m 00sComing soon
- Handling disagreement questions4m 00sComing soon
- Questions to ask the interviewer3m 30sComing soon
System Design
- A framework for DevOps system design4m 00sComing soon
- Design: deploy a SaaS to prod5m 00sComing soon
- Design: migrate a monolith to Kubernetes5m 00sComing soon
- Design: multi-region HA5m 00sComing soon
- Design: an observability stack5m 00sComing soon
- Design: CI/CD for 500 engineers5m 00sComing soon
- Design: a secret management platform5m 00sComing soon
Debugging Drills
- Drill: pod is in CrashLoopBackOff5m 00sComing soon
- Drill: cluster DNS is broken5m 00sComing soon
- Drill: pod is Pending forever5m 00sComing soon
- Drill: Ingress returns 5035m 00sComing soon
- Drill: node disk is full5m 00sComing soon
- Drill: p99 latency spiked5m 00sComing soon
- Drill: TLS cert just expired5m 00sComing soon
- Drill: DB connection pool exhausted5m 00sComing soon
18. Supporting Concepts
2 lessonsCross-cutting mental models that appear across modules — put here so lessons in other modules can cite them.
Data Structures Aside
Future modules
Azure (overview) · GCP (overview) · Chaos Engineering · Service Level Objective platform (Nobl9-style) · Serverless (Lambda, Cloud Functions) · Edge (Cloudflare Workers) · Data platform DevOps (Kafka, Spark, warehouse) · ML platform DevOps (Kubeflow, MLflow) · Compliance (SOC2, HIPAA, PCI) · Advanced Networking (BGP, WireGuard) · Advanced storage (Ceph, MinIO) · Backup Systems (Restic, Barman)