Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

  > At the moment the team is focusing on fixing the issue
  > and I don't want to distract them from that.
That might be ok for feature teams, but for infrastructure tools/services, it's very frustrating for users (devs) to be kept in the dark on the progress of the fix.

At work, the incident response starts with identifying Investigators (to find and fix the problem) and a Communicator (to update channel topics, send the outage email & periodic updates, field first-line questions about the incident, and to contact those most affected by the incident so they don't get surprised/try to fix it themselves). The person who starts the incident is the Coordinator, who assigns the roles, escalates if more help is needed, tries to unblock investigations, and turns facts from the investigators into status updates for the communicator.



i will provide an opposing viewpoint which i'm sure many people do not agree with.

if a service i use is down, all i want is an acknowledgement and that "we are working on it right now with high priority". i want all available resources to be fixing the problem.

my anxiety over powerlessness in relying on others during a crisis manifests in other ways, like figuring out why i'm at the mercy of this thing in the first place, and putting alternatives in place.

but during the crisis i'll just go do something else for an hour and then read the post mortem when it comes out.


Updates that are communicated are not the details of what's happening in the investigation, but things that are expected to be useful to users/clients. They are things such as the estimated time for problem resolution, updates on the scope of the problem (e.g. "this is an instrumentation problem" vs "this is an actual outage") or mitigation steps that could be applied by the clients. It is often very useful to know such things during an outage.


i don't think anyone, anywhere would have given you an accurate estimated figure of nearly 5 hours to fix this problem.

furthermore, even if you somehow could divine the future, telling the customer that you think an outage will last over half a working day is going to turn an extremely shitty situation into something even worse.


sometimes the resources available are suitable for different problems.

I consider high quality communicators to be a significant resource, and the problem to solve is to make sure everyone affected has the right information to make their next decision to mitigate the issue.

These communicators might not know how to solve the actual technical problem, and it'd be like putting monkeys on a typewriter to tell them to "all hands in fixing a configuration issue with the web server" for example.

I'd say apply the right people to the right problems.


Keeping the affected users informed is part of 'fixing the problem'.

Sometimes the problem is even something that some users will be able to work around if they know what it is.


I agree with you and we have a similar process at Docker. Part of what went wrong in this particular case is precisely that our infrastructure incident response process did not kick in. I am guessing that one conclusion of the post-mortem will be "make sure to handle open-source package distribution infrastructure like the rest of our infrastructure, including incident response checklist".




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: