Top
Best
New

Posted by mooreds 10 hours ago

I don't want the details(michaelheap.com)
323 points | 187 commentspage 3
baud9600 9 hours ago|
I think the problem is when the execs aren’t looking at change-for-the-better. Instead they’re looking to appoint blame. Moving to a position of assumed competency in your staff and taking the perspective of how-to-improve-next-time is a real step forward in maturity. My guess is it’s rare and blame is easier: most management doesn’t reach this level, instead it’s immature.
Lerc 9 hours ago||
I think that's a problem domain in itself. Identify when failures that occur even when skilled people doing their jobs properly, adjust the environment so that they no longer happen.

It is understandably a completely different task compared to what those skilled people are specialised to do. You probably need a dedicated role to do it.

Perhaps you could call them a manager. Their job is to see the multiple parts of the system. They should ask for the Details of what happened so they can determine why the problem occurred.

Consider one of the problems listed in the article

>"The alert fired, but the on-call engineer had already dealt with twenty low-value alerts that evening".

The engineer can say they were busy, they didn't see the alert, that they are swamped with things they think are low-value. Someone else can say the alert fired. Each person involved may have their own perspective, with different ideas as to what the problem actually is.

It's easy when you see problem described in terms of what the solution is. Someone needs to figure that out, to do that they need the details.

monideas 9 hours ago||
An SVP of engineering should care about the details, this is the essence of leadership. When you don’t care about the details you are just a manager and worthless i.e. you should be replaced with a leader.
kevin_nisbet 6 hours ago||
Lots of folks are commenting on the internal dynamics, but the other part that didn't resonate with me was actually the what happens next part. This may be an overreaction on my part, stemming from watching lots of engineers suggest expensive solutions to relatively minor issues. And it's a mistake I've made myself before as well, I started in telecom with very strict standards for availability, so it's hard to beat out of me the desire to look through the most minor alarm for a potential problem brewing or get to a root cause of the most minor incidents.

And maybe this is just a part that the article dodges, but for something described as non-catastrophic, I would suggest that there shouldn't be a presumption of changes. In my view there should be an assessment of the risk of re-occurrence, and if it was a near miss, what is the risk had it not been a near miss. And then even should changes be part of the outcomes, how expensive are those to implement compared to whatever the issue was, and if they're too expensive compared to the risk, don't put resources into it.

miiiiiike 8 hours ago||
My version of this is: "This is the kind of mistake that we get to make once. How do we prevent it from happening again."

I usually do want the details, but I also want to communicate my expectations up top and unambiguously.

ape4 10 hours ago||
A Senior Vice President of ENGINEERING doesn't want the technical details?
kube-system 8 hours ago||
Yes, if a director and SVP are discussing the technical details of an incident, I would be concerned -- they should be concerned with putting the right people and processes in place to make those decisions.

When "Something had gone wrong that shouldn't have" bubbles up that far, you have an organizational issue, not a technical issue.

heisenbit 6 hours ago|||
Of course, details might point upwards - plausible deniability could get damaged.
cbg0 9 hours ago||
If the person who is charged with resolving the issue is capable, there's no point in a senior manager knowing the details outside of professional curiosity.
Shacharp 9 hours ago||
This is a really really smart SVP. The answer to “how do we prevent this in the future” is to identify decisions that reduce that problem from happening again. This doesn’t have to be perfect. You can say “we are going to do X next time” and also say “but we don’t know how far X will work”. When the failure happens again, you can retire process X and move on to Y. The aim is to make a series of decisions that eventually get to the heart of the issue.

Everything else is a conversation that takes up space on a post-mortem or runbook.

b3lvedere 6 hours ago||
“If your corrective action depends on people remembering a conversation from six months ago, you don't have a corrective action. You have organizational folklore.”

I think i’m going to call our ticket support system “folklore” from now on. It sure is used that way.

airstrike 8 hours ago||
> What they didn't want was for empathy to become the mechanism by which the organisation absolved itself of having to change.

That line is absolute gold

mfbx9da4 5 hours ago|
The assumption that something always has to change is the culture that leads startups to knee jerk their way into miles of red tape and performative bureaucracy
More comments...