Real outages. Real breaches. Told by the companies that lived through them.
Common failures repeat: expired certificates, config pushes that go everywhere at once, retry storms, databases running out of IDs, keys left where someone could find them. Describe a failure mode and this page pulls the matching case files, leading with what caused each one and what fixed it. It's meant for the time around incidents, not the middle of one: readiness and design reviews, on-call training, and postmortems that point to known patterns and preventions. Mid-incident, the job is capturing data from your own system, because every system is its own thing.
Recurring patterns across these files, and how the companies involved warded them off afterwards.
Public postmortems lean toward availability incidents a company is comfortable telling: a clever root cause, a fast fix, lessons learned. Security incidents, often the costliest, are published less often, later, and with less detail, because lawyers, regulators and disclosure rules shape what gets said. This collection includes a set of credential leaks and supply-chain compromises, but like anything built from public writing, it underrepresents the failures companies would rather not discuss.