Posted by mooreds 10 hours ago
It is understandably a completely different task compared to what those skilled people are specialised to do. You probably need a dedicated role to do it.
Perhaps you could call them a manager. Their job is to see the multiple parts of the system. They should ask for the Details of what happened so they can determine why the problem occurred.
Consider one of the problems listed in the article
>"The alert fired, but the on-call engineer had already dealt with twenty low-value alerts that evening".
The engineer can say they were busy, they didn't see the alert, that they are swamped with things they think are low-value. Someone else can say the alert fired. Each person involved may have their own perspective, with different ideas as to what the problem actually is.
It's easy when you see problem described in terms of what the solution is. Someone needs to figure that out, to do that they need the details.
And maybe this is just a part that the article dodges, but for something described as non-catastrophic, I would suggest that there shouldn't be a presumption of changes. In my view there should be an assessment of the risk of re-occurrence, and if it was a near miss, what is the risk had it not been a near miss. And then even should changes be part of the outcomes, how expensive are those to implement compared to whatever the issue was, and if they're too expensive compared to the risk, don't put resources into it.
I usually do want the details, but I also want to communicate my expectations up top and unambiguously.
When "Something had gone wrong that shouldn't have" bubbles up that far, you have an organizational issue, not a technical issue.
Everything else is a conversation that takes up space on a post-mortem or runbook.
I think i’m going to call our ticket support system “folklore” from now on. It sure is used that way.
That line is absolute gold