Sally Roth

SRE Lab

Small, working tools and explainers for reliability engineering: setting reliability targets, knowing when something is really broken, learning from the failures that keep coming back, and how Kubernetes and Linux actually work, including what breaks at scale. More on the way: incident management, cardinality, delivery, observability, security, Terraform, and the same patterns in Python, Go and Ruby.

SLO Burn SimulatorSee how fast an error budget drains, which burn-rate alerts fire and when, and get matching Prometheus rules. Down or MeUptime checks with an AI classifier that decides whether a failure is real, and only alerts when it's sure. The Incident Files310 real postmortems, searchable by failure mode, plus the nine failure patterns that keep coming back and how companies warded them off. Kubernetes, DrawnHow Kubernetes works in 18 step-through diagrams, from kubectl apply to DNS, storage and upgrades, each with a 30-second answer, gotchas, a quiz and what breaks at scale. Linux, DrawnWhat a Linux node really does with memory and packets: page cache, the OOM killer, cgroups, a packet's path, conntrack, TCP queues and the first 60 seconds of performance analysis.

Still growing