Posts tagged operations.
This is one of my saltier posts. You've been warned.
Having too many alerts that drive everyone insane is still better than having no alerts at all. I've complained about alert fatigue plenty of times before, but here's the uncomfortable truth: that statement is completely backwards.
Engineering leadership's most expensive monitoring decision isn't choosing the wrong tool. It's falling into the monitoring trap that costs organizations in wasted engineering time and preventable downtime.
Most teams focus on faster incident response. The real solution is preventing incidents from happening in the first place.
Traditional war games work great for single teams. But what happens when you have three subsystem teams, a dedicated SRE group, and multiple stakeholders who all need to respond to incidents together?
Here’s a bit of a paradox: the better you are at solving SaaS production incidents, the harder each incident is to solve.
A little excitement in your job is usually a good thing. It could be learning a new development language, preparing to release a new feature, or taking on new responsibilities as part of a promotion. That’s great for most jobs, but not operations.
Over the last two articles (SaaS War Games - Part 1 and SaaS War Games - Part 2), we dove into the value of War Games. In this article, the rubber meets the road on how to run one.
This article builds on SaaS War Games - Part 1, so I recommend reading that article before diving into this.
It’s 3 AM and your phone is ringing. There’s only one number you let ring through your Do Not Disturb settings. You open one eye and look at the first of 12 on-call notifications.