Observability & SRE
-
Root Cause Analysis Is the Wrong Question
Complex systems fail through combinations of conditions, not single causes. Searching for the root cause produces a satisfying story and…
Read More » -
Five Nines Is a Vanity Metric: Setting SLOs That Mean Something
Availability targets set without measuring what users experience produce dashboards nobody trusts. Error budgets turn reliability from an argument into…
Read More » -
Trace Sampling Decides Which Incidents You Can Investigate
Keeping one percent of traces at random means the trace you need during an incident was almost certainly discarded.
Read More » -
Grepping Logs Stops Working at About Three Services
Free-text log lines are readable by a human on one machine and unqueryable across a distributed system.
Read More » -
Alerting on Causes Produces Noise; Alerting on Symptoms Produces Pages
High CPU is not a problem. Users unable to complete a purchase is a problem. Most alert fatigue comes from…
Read More » -
One Alert Threshold Cannot Catch Both Fast and Slow Failures
A threshold sensitive enough to catch a severe five-minute outage pages constantly on noise. One tuned to avoid noise misses…
Read More » -
Most Postmortem Action Items Never Get Done
A review that produces fourteen items and completes three has not improved reliability. It has produced a document.
Read More » -
Orderlek Unveils Multi-Storefront Architecture for Global B2B Enterprises
Discover how Orderlek’s Multi-Storefront Architecture empowers global B2B brands to launch hundreds of localized, unique storefronts from a single backend…
Read More »