Incident Management: the Complete Guide

Technology is everywhere, and we depend on it. Whether it be home, work, school, health, or civic needs, technology is involved.

That’s why we get so frustrated with any disruption from normal operation. Social media is awash with stories of systems not performing as they should: banking systems, healthcare portals, airline booking, online shopping, even the social media platforms themselves.

For this reason, most relationships between any service provider—you—and your customers depend heavily on whether you can ensure minimal disruption. When the inevitable disruption does occur, you must manage the incident in a way that the consumer has agreed to tolerate.

Incident management is the formal name of this necessary business practice, and it’s not one for companies to take lightly, no matter your industry. This article will look at many parts of the incident management practice, including:

Let’s get started.

What is an incident?

When it comes to ensuring that operational services provide value to customers, incident management is among the most important disciplines. ITIL® 4 defines an incident as:

An unplanned interruption to a service or reduction in the quality of a service.

Here are some other definitions:

  • ISO 20000 defines an incident as an unplanned interruption to a service, a reduction in the quality of a service, or an event that has not yet impacted the service to the customer or user.
  • VeriSM broadens the term “issue” to covering situations when a customer perceives a service interruption, as well as actual interruptions.

How a service provider handles incidents plays a very significant role in determining customer satisfaction. Here are some examples of an incident in an online system:

  • Users not being able to log in
  • The system’s lack of responsiveness to commands
  • Perceived slowness compared to normal
  • Corrupted or hacked data

Of course, not all incidents are visible to the end user. But they still require your attention.

What is incident management?

The purpose of the incident management practice is to minimize the negative impact of incidents by restoring normal service operation as quickly as possible.

Whether it’s a crashed laptop, corrupted data or a painfully slow application, how we respond and deal with the interruption to service indicates whether we have an optimal incident management process.

This practice can be handled by an individual, teams or multiple organizations depending on the scale. Successful organizations designate a specific Incident Commander (IC)—one person who is responsible for leading a temporary cross-functional team to focus all energies and attention towards a swift resolution.

Your company’s ability to quickly address incidents is a key factor in:

  • User and customer satisfaction
  • Your credibility and reputation
  • The value you create in your relationships

Clearly, these are areas you must excel at—making incident management a critical activity. At its most essential, incident management involves two main activities:

  1. Record
  2. Manage

Once you identify, or get notified of, the incident, you would capture just enough information about it, including description, time, and source. This record then becomes the basis for analysis and decisions on managing the incident, including:

  • Communication
  • Resolution
  • Escalation
  • Handoff to other processes

The classic incident management activities as outlined in ITIL 4 are seen below:

Successful incident management relies on having a clear understanding of what the customer agreed to or is willing to tolerate regarding the duration and handling of any particular incident. This is usually defined in service level agreements (SLAs) or contracts, which include timelines for responding and resolving incidents based on some criteria, usually priority, as a function of impact and urgency.

As the service provider, how you structure your organization to handle different types of incidents is a major driver in your incident management execution:

  • Some incidents may be repeatable, with their causes known. In these cases, you can define and use incident models for handling and resolution. An incident model is a repeatable approach to managing a particular type of incident. Models help reduce both resolution time and the learning curve for new employees.
  • Where a solution to an incident is not easily found, a workaround may be applied to try and lessen the impact and/or probability of recurrence. Workarounds, such as restarts or reconfigurations, can quickly restore services back to an acceptable level of quality.

(Importantly, incident management is different from problem management, which focuses on how you handle the problem in the future.)

Elena Rostova

Elena Rostova

Lead Health, Wellness & Medical Journalist

Elena Rostova holds a Master's degree in Public Health Journalism. She covers groundbreaking medical research, holistic wellness trends, mental health awareness, and nutritional science.

Share this article