Introduction to Mttd: Mean Time to Detect
Devops Teams Rely on Data-Driven Insights to Establish and Maintain an Efficient Sdlc Pipeline. a Variety of It Incident Metrics Support Evaluating the...
DevOps teams rely on data-driven insights to establish and maintain an efficient SDLC pipeline. A variety of IT incident metrics support evaluating the performance evaluation of your IT infrastructure overall as well as individual IT incidents. The challenge for ITOps and service desks centers around identifying the most meaningful metrics, especially those that contribute to minimizing the impact of IT incidents on end users.
In this article, we will discuss mean time to detect, a metric focused on finding IT incidents faster, regardless of the issue resolution capacity and performance. Both DevOps and ITOps teams can utilize MTTD.
What is mean time to detect?
A key performance indicator (KPI) within IT incident management, mean time to detect (MTTD) refers to the average time passed between the onset of an IT incident and its discovery.
MTTD can be calculated mathematically with the following formula:
Though calculating the MTTD mathematically is simple, it does require an accurate knowledge of the start time of IT incidents. This may require an evaluation of historical infrastructure KPI data. Take the average of time passed between the start and actual discovery of multiple IT incidents. These calculations can be performed across different periods (e.g., daily, weekly, or quarterly) to evaluate changes in MTTD performance over time.
Mean time to detect is one of several metrics that support system reliability and availability
MTTD for effective incident management
Incident management comes down not merely to detecting incidents, but to determining what incidents may impact end-user performance or your revenue stream. Incidents that do not affect performance and revenue earning are likely prioritized. MTTD can help you track both:
- Numerous infrastructure metrics, not a single one, should report revenue-impacting incidents. That’s a good thing, but it has some failings: the vast volume of log metrics generated at every IT means there’s a lot of noise. IT teams using traditional ITSM tooling may struggle to identify IT issues proactively.
- End users also help identify IT incidents when they report performance-impacting incidents such as service outages, disruptions, or other performance issues. If the underlying incidents are not already known or visible to your team, the types and scale of end-user reports might indicate a discrepancy around your incident management methods and/or an inadequacy of your monitoring solutions to recognize the issues proactively.
Addressing these two MTTD-related issues can be both technical and strategic in nature:
- DevOps may need reevaluate how they define the impact tolerance of an incident.
- IT may need to invest in technology solutions capable of producing granular insights using the log metrics big data. Some ITSM event tools are capable of correlating information from multiple sources of IT incident information in order to identify hidden patterns of insights.
Uncovering these can help you identify a significant IT incident before it makes its impact on performance and revenue. Once an IT incident is recognized as higher priority, DevOps teams can follow pre-designed issue resolution protocols to address issues based on priority and impact on SDLC performance.