Observability
Metrics, logs, traces — the three pillars that let you debug production.
“Duolingo for DevOps. One concept = one 2-4 minute video. Never introduce a concept before its prerequisites. Every lesson builds on the last.”
0 / 61 lessons readyaudience: sre & devops61 lessons planned
1. Monitoring
30 lessonsYou can't fix what you can't see. Metrics, alerting, and dashboards — the reactive half of observability.
Observability Concepts
- Metrics, logs, traces — the three pillars3m 30sComing soon
- What is a metric?3m 00sComing soon
- Counters, gauges, histograms4m 00sComing soon
- Pull-based vs push-based metrics3m 30sComing soon
- Cardinality — the metric budget4m 00sComing soon
Prometheus
- Prometheus — the reference model3m 30sComing soon
- scrape_configs — telling Prometheus what to poll4m 00sComing soon
- Exporters — node_exporter, cAdvisor, kube-state4m 00sComing soon
- Service discovery in Kubernetes4m 00sComing soon
- PromQL basics4m 00sComing soon
- rate, irate, and the common gotchas4m 00sComing soon
- Alertmanager4m 00sComing soon
- Federation and Thanos4m 00sComing soon
- Prometheus HA patterns4m 00sComing soon
- Recording rules — pre-computing4m 00sComing soon
Grafana
- Grafana — the eyes on the data3m 30sComing soon
- Data sources3m 30sComing soon
- Panels and visualizations4m 00sComing soon
- Dashboard variables4m 00sComing soon
- Dashboard provisioning as code4m 00sComing soon
- Grafana alerts vs Alertmanager4m 00sComing soon
- Annotations — marking deploys4m 00sComing soon
Alerting Discipline
- Alert fatigue and how to avoid it3m 30sComing soon
- Symptom-based vs cause-based alerts4m 00sComing soon
- SLOs and error budgets4m 00sComing soon
- On-call rotation basics3m 30sComing soon
- Incident playbooks4m 00sComing soon
- Routing and inhibition rules4m 00sComing soon
- Wiring alerts to Slack + PagerDuty3m 30sComing soon
- Postmortem culture4m 00sComing soon
2. Logging
15 lessonsStructured logging is a superpower. Loki is Prometheus-for-logs. The DevOps job is often reading logs from many services at once.
Log Fundamentals
- What makes a good log line?3m 30sComing soon
- Structured logs — JSON everywhere3m 30sComing soon
- Log levels used correctly3m 30sComing soon
- Correlation IDs across services4m 00sComing soon
- Where logs go — stdout, file, sink3m 30sComing soon
Loki
- Loki — the mental model3m 30sComing soon
- Promtail — the log shipper4m 00sComing soon
- LogQL basics4m 00sComing soon
- Labels done right in Loki4m 00sComing soon
- Retention and storage classes4m 00sComing soon
- Loki HA architecture4m 00sComing soon
Aggregation Patterns
- Fluent Bit — the swiss army log agent4m 00sComing soon
- The ELK stack briefly4m 00sComing soon
- Multi-tenant logging pitfalls4m 00sComing soon
- PII in logs — the compliance minefield4m 00sComing soon
3. Tracing & OpenTelemetry
16 lessonsDistributed tracing turns 'the request was slow' into 'the DB call in service B took 3 seconds.' OpenTelemetry is the vendor-neutral standard.
Tracing Concepts
- Traces, spans, and context propagation4m 00sComing soon
- Sampling — head vs tail4m 00sComing soon
- Baggage and cross-cutting attributes4m 00sComing soon
Tempo
- Tempo — traces at scale4m 00sComing soon
- TraceQL basics4m 00sComing soon
- Tempo + Grafana integration4m 00sComing soon
OpenTelemetry
- OpenTelemetry — one API to rule them all4m 00sComing soon
- SDK vs API vs collector4m 00sComing soon
- Auto-instrumentation for common runtimes4m 00sComing soon
- The Collector — receivers, processors, exporters4m 00sComing soon
- Context propagation across languages4m 00sComing soon
- Semantic conventions — why they matter4m 00sComing soon
SLIs, SLOs, SLAs
- What makes a good SLI?3m 30sComing soon
- SLOs — the reliability target4m 00sComing soon
- SLA vs SLO — legal vs engineering3m 30sComing soon
- Error budget policies4m 00sComing soon