Small, working tools and explainers for reliability engineering: setting reliability targets, knowing when something is really broken, learning from the failures that keep coming back, and how Kubernetes and Linux actually work, including what breaks at scale. More on the way: incident management, cardinality, delivery, observability, security, Terraform, and the same patterns in Python, Go and Ruby.
Sprouting nowSane Incident ManagementSeverity, roles, the incident lifecycle, communication, on-call, blameless reviews, action items and readiness.
Sprouting nowTaming CardinalityHow one innocent label multiplies time series and cost, how Prometheus stores them, and how to keep metrics useful without the bill.
Sprouting nowDelivery, DrawnCI pipelines, artifact promotion, GitOps with Argo CD, canaries, feature flags, safe migrations and supply-chain security.
Sprouting nowObservability, DrawnOpenTelemetry and the Collector, Prometheus and PromQL, tracing and sampling, logs, and alerting that wakes the right person.
Sprouting nowSecurity for SREs, DrawnHow AWS decides allow or deny, TLS and mTLS, secrets, network segmentation, least privilege and audit trails.
Sprouting nowTerraform, Five Years OnMoved, import and removed blocks, checks and tests, the license change and OpenTofu, ephemeral values, and Stacks.
Sprouting nowOne Pattern, Three LanguagesRetries, timeouts, bounded concurrency, error handling, graceful shutdown and circuit breakers, side by side in Python, Go and Ruby.