Measure What Matters

Mark Overholser (Corelight TME + Threat Hunter)

SAINTCON 2025 · Day 2 · Main Track 1

Overview

In his SAINTCON talk, "Measure What Matters," Mark Overholser, a Technical Marketing Engineer and Threat Hunter at Corelight, delivers a foundational yet critical message for blue teams and Security Operations Centers (SOCs): the imperative of data-driven decision-making. Overholser argues that without robust and relevant metrics, security efforts are reduced to mere guesswork, making it impossible to accurately assess performance, justify investments, or communicate value to the broader business. The presentation serves as a guide for security professionals, particularly those early in their careers or building out security teams, on how to establish effective measurement practices.

Watch on YouTube

Visual summary for Measure What Matters by Mark Overholser
Visual summary for Measure What Matters by Mark Overholser

Key moments

  1. 1:19 Why measure? Importance of data and stakeholder alignment.
  2. 2:07 How to measure: Importance of clear definitions.
  3. 2:59 NIST definition of a cybersecurity incident example.
  4. 4:00 Applying SMART goals to security measurement.
  5. 5:08 Common mistake: Changing definitions mid-stream.
  6. 6:00 Common mistake: Measuring what's merely convenient.
  7. 6:59 Common mistake: Not realizing a metric is a proxy.

Measure What Matters

Speakers: Mark Overholser, Corelight TME + Threat Hunter

Conference: SAINTCON

YouTube: https://www.youtube.com/watch?v=ImywQ0mzmCk

Overview

In his SAINTCON talk, "Measure What Matters," Mark Overholser, a Technical Marketing Engineer and Threat Hunter at Corelight, delivers a foundational yet critical message for blue teams and Security Operations Centers (SOCs): the imperative of data-driven decision-making. Overholser argues that without robust and relevant metrics, security efforts are reduced to mere guesswork, making it impossible to accurately assess performance, justify investments, or communicate value to the broader business. The presentation serves as a guide for security professionals, particularly those early in their careers or building out security teams, on how to establish effective measurement practices.

Overholser emphasizes that proper measurement extends beyond simply collecting data; it necessitates clear definitions, a focus on impactful metrics over convenient ones, and a strategic understanding of how these metrics influence team behavior and business outcomes. He draws parallels from project management principles and basic statistics to construct a framework for quantifying risk and evaluating operational efficiency. The talk highlights common pitfalls in security measurement and provides actionable strategies to overcome them, ultimately empowering security teams to articulate their value in terms the business understands.

This discussion is crucial because security teams are often perceived as cost centers. By systematically measuring "what matters," organizations can transform their security function from a reactive expense into a proactive, value-generating entity. Overholser's insights enable defenders to not only improve their operational effectiveness but also to confidently advocate for the resources and support necessary to bolster their organization's overall security posture against an ever-evolving threat landscape.

Background

▶ Watch: Why measure? Importance of data and stakeholder alignment. (1:19)

The fundamental premise of Overholser’s talk is that "without data, you're just guessing." This statement underscores a common challenge in cybersecurity: the difficulty in objectively assessing the effectiveness of security controls, incident response processes, and overall risk posture. Many organizations operate on instinct, anecdotal evidence, or metrics that are easy to collect rather than those that truly reflect security outcomes. This lack of empirical grounding can lead to misallocated resources, an inability to demonstrate return on investment (ROI) for security spending, and a perpetual struggle to gain executive buy-in.

A significant part of the background context involves the pitfalls of poor measurement. Overholser identifies several common mistakes. First, a lack of clear, written definitions for key terms like "incident" or "severity" leads to ambiguity and disagreement, hindering consistent reporting and decision-making. He advocates for adopting established frameworks, such as the NIST framework for improving critical infrastructure cybersecurity version 1.1, which provides a starting point for defining an incident as "a cybersecurity event that has been determined to have an impact on the organization prompting the need for response and recovery." This clarity ensures that everyone, from analysts to executives, speaks the same language.

Second, he cautions against changing definitions mid-stream. If metrics are to be compared over time (e.g., week-over-week incident counts), the underlying definitions must remain constant to ensure an "apples to apples" comparison. Changing a definition requires resetting the measurement baseline. Third, and perhaps most prevalent, is the error of measuring what’s convenient rather than what’s important. Security teams often report statistics readily available from appliances (e.g., "7 million blocked connections on the firewall"), without critically evaluating what question these numbers answer or their significance to the business's agreed-upon values.

Another critical concept introduced is the idea of a proxy metric. Overholser explains that sometimes, a direct measure of what's truly desired (e.g., emotional maturity) is impractical, so a substitute (e.g., physical age) is used. While proxies can be useful, security teams must be aware when they are relying on them and understand their limitations and potential for misrepresentation. Finally, a profound warning comes in the form of Goodhart's Law: "When a measure becomes a target, it ceases to be a good measure." This phenomenon occurs when an incentive is tied to a metric, causing individuals or teams to alter their behavior to hit the target, often at the expense of the true objective. For instance, if analysts are rewarded for quick ticket acknowledgment, they might click "acknowledge" without fully reading the ticket, skewing the metric and undermining the actual goal of prompt incident review.

To counter these issues, Overholser suggests applying principles from project management, specifically the SMART criteria for goal setting: Specific, Measurable, Relevant, and Time-bound. He intentionally crosses out "Achievable" for measurement purposes, arguing that the goal is to observe and understand, not necessarily to set achievable targets that might inadvertently trigger Goodhart's Law. By adhering to these principles, security teams can develop a more robust and meaningful measurement strategy.

Key Findings

▶ Watch: NIST definition of a cybersecurity incident example. (2:59)

Overholser's talk distills several key findings that are fundamental to establishing effective security measurement practices for blue teams:

  1. The Indispensability of Data: The overarching finding is that data is non-negotiable for effective security operations. Without it, security decisions, resource allocation, and performance evaluations are based on conjecture. Data provides the objective foundation for understanding current posture, identifying areas for improvement, and demonstrating value.
  2. Definitions are Paramount: The talk strongly emphasizes that clear, unambiguous, and agreed-upon definitions are the bedrock of any meaningful measurement system. Disagreements about what constitutes an "incident" or how "severity" is assessed undermine the consistency and reliability of all metrics. Adopting external standards like the NIST framework is a recommended starting point.
  3. Focus on Value, Not Convenience: A critical insight is the distinction between measuring what is convenient (e.g., appliance statistics) and measuring what is truly valuable to the organization and its stakeholders. Security metrics must directly answer questions about the business's risk profile and the security team's contribution to mitigating that risk.
  4. Goodhart's Law is a Potent Threat: Overholser highlights the significant danger of Goodhart's Law, where turning a measure into a target inadvertently distorts behavior and renders the metric useless. This finding serves as a warning against overly incentivizing specific metrics without considering the broader implications for honest reporting and genuine improvement.
  5. Risk Quantification in Monetary Terms is Transformative: A pivotal finding is the proposed method for quantifying risk using expected value (probability x impact), specifically by translating impact into dollars and probability into a percentage over a defined period (e.g., 90 days). This approach allows security teams to speak the language of business, enabling direct comparison of control costs against potential financial losses and facilitating informed investment decisions.
  6. Embrace Imperfection and Iteration: Finally, Overholser's advice to "take your best guess" and "give yourself permission to be wrong" is a powerful finding. It acknowledges that perfect data is often unattainable, especially initially. The key is to start measuring, learn from the data, re-evaluate, and iteratively refine the metrics and underlying assumptions. This iterative approach fosters continuous improvement in measurement accuracy and security posture.

Technical Deep Dive

▶ Watch: Applying SMART goals to security measurement. (4:00)

The technical deep dive into Overholser's methodology for "Measuring What Matters" revolves around establishing robust KPIs and a pragmatic approach to risk quantification. The core framework begins with precise definitions. He champions the NIST definition of a cybersecurity incident as "a cybersecurity event that has been determined to have an impact on the organization prompting the need for response and recovery." This provides a foundational, universally understood starting point for incident tracking. Furthermore, all metrics must adhere to SMART principles (Specific, Measurable, Relevant, Time-bound) to ensure clarity and utility.

Overholser then outlines a suite of Key Performance Indicators (KPIs) specifically tailored for blue teams and SOCs:

  1. Mean Time To Detection (MTTD): This metric measures the average time from the estimated start of an incident to when the organization becomes aware of it. It directly answers the question: "Do we have good visibility? How quickly are we detecting things?" A decreasing MTTD indicates improved monitoring and threat intelligence capabilities.
  2. Mean Time To Acknowledgement (MTTA): This measures the average time from a ticket's creation or assignment to its acknowledgment by an analyst. It addresses: "Do we have enough analysts?" A rising MTTA could signal an increasing workload or insufficient staffing, providing data to justify team expansion or process optimization.
  3. Mean Time To Response/Recovery/Resolution/Remediation (MTRs): Overholser stresses the ambiguity of a generic "MTR." Instead, he advocates for distinct, clearly defined MTRs:
  • Mean Time To Response: Time from detection to when incident response activities begin.
  • Mean Time To Remediation: Time from the start of incident response to when incident response ends or closes.
  • Mean Time To Resolution: Time from the start of incident response to when affected systems return to a defined operational state (e.g., "90% of pre-incident capability"). Each MTR answers the question: "Are we dealing with incidents quickly and effectively?"
  1. Percentage of Closure & Types of Closure: This metric tracks the proportion of incidents closed and categorizes them by reason. Overholser suggests categories like "Expired" (e.g., not resolved within 7 days), "Solved," and "Not Solvable." The "Not Solvable" category is particularly powerful. It acknowledges a lack of sufficient information or tooling, providing concrete data to justify investments in new security tools, telemetry, or training. For instance, if 15% of incidents are closed as "Not Solvable" due to insufficient endpoint logging, this becomes a clear business case for an EDR solution.

The most technically profound aspect of the talk is Overholser's method for quantifying risk. Initially, he discusses a basic qualitative risk register using a 1-10 scale for both probability and impact, yielding a risk score from 1-100. While useful for internal ranking, this lacks business context. He then introduces a more sophisticated, financially-driven model:

Risk = Probability (%) × Impact ($)

Here, Probability is defined as the percentage likelihood of an event occurring within a specific time-bound period (e.g., 90 days). Impact is quantified directly in dollars, representing the estimated financial loss if the event occurs (e.g., data breach costs, downtime, regulatory fines). The resulting "Risk in Dollars" provides a tangible, financial value for each identified vulnerability or exposure. This allows for:

  • Stack Ranking: Prioritizing risks by their potential financial exposure.
  • Cost-Benefit Analysis: Directly comparing the cost of implementing a control (also in dollars, including person-hours translated to monetary value) against the dollar value of the risk it mitigates. This empowers security leaders to say, "Investing $50,000 in this control will reduce a $250,000 risk by 80%," making a compelling business case.

Overholser emphasizes that this can be started simply in a spreadsheet (like Excel or Google Sheets), acknowledging that larger organizations might eventually outgrow it for dedicated databases or GRC platforms. The key is the shift in perspective from abstract security scores to concrete financial implications, enabling security to communicate its needs and successes in the universal language of business.

Demo / Proof of Concept

▶ Watch: Common mistake: Measuring what's merely convenient. (6:00)

While the talk does not feature a live software demonstration or a complex technical exploit, Mark Overholser effectively uses incident timelines and a risk register spreadsheet as practical, illustrative "proofs of concept" for his measurement principles. These examples serve as a clear demonstration of how the discussed metrics are applied in real-world scenarios.

Incident Timeline Example 1: Ransomware Attack

Overholser presents a simplified incident timeline to illustrate Mean Time To Detection (MTTD):

  • User clicked on phishing link.
  • Credentials entered into phishing site.
  • Thread actor logs into virtual desktop using fished credentials (Incident Start: January 15th). This is defined as the incident start because it's the first point of actual impact on the organization.
  • Failed RDP sessions to domain controller.
  • Data exfiltration from ERP system.
  • Successful RDP to domain controller.
  • Ransomware deployed (Detection: January 19th). This is when the organization finally detected the incident.

Through this example, Overholser visually highlights the time to detect as the interval between January 15th and January 19th. This practical illustration underscores the importance of defining the "start" of an incident based on organizational impact, not just the initial malicious action, and how a significant gap can exist between the actual compromise and its detection.

Incident Timeline Example 2: Compromised User Device

A more detailed and realistic incident timeline is used to demonstrate multiple MTR metrics:

  • Pre-detection (unknown start): User might have clicked a link or downloaded a file, but no telemetry exists.
  • Detection: An alert fires for a user's computer scanning the network on several ports, alongside alerts for multiple failed SSH authentication attempts to a core switch.
  • Acknowledgement: A security analyst reviews the alert and promotes it to an incident. The time from detection to this point is the Mean Time To Acknowledgement (MTTA).
  • Response: The security analyst initiates quarantine of the end-user's device. The time from acknowledgment to this action is the Mean Time To Response (MTR).
  • Resolution: A hardware technician replaces the device, and the user's credentials are reset. The time from the start of response to this point is the Mean Time To Resolution (MTR).

This second example vividly demonstrates the sequential nature of incident response and how different time-bound metrics capture distinct phases of the process. It also realistically acknowledges "question marks" (unknown initial incident start) when telemetry is lacking, reinforcing the idea of making educated guesses and improving over time.

Risk Register Spreadsheet Example

Overholser presents a basic spreadsheet, callable in Excel or Google Sheets, as a "risk register." This is a practical, low-cost "proof of concept" for his risk quantification method:

  • Columns: Vulnerability/Exposure, Probability (initially 1-10, then a percentage over 90 days), Impact (initially 1-10, then in dollars), Calculated Risk (Probability x Impact), Cost to Address, Control Description.
  • Functionality: The spreadsheet allows users to list identified risks, assign subjective or estimated probabilities and impacts, and then calculate a raw risk score. Critically, he modifies this to show how converting Impact to dollars and Probability to a percentage yields a Risk in dollars.

This simple spreadsheet demonstrates how any organization can immediately begin quantifying risk in financial terms. It provides a tangible tool for stack-ranking risks, comparing them against the cost of mitigation, and making data-backed arguments for security investments. He emphasizes that while simple, this spreadsheet approach is powerful for getting started and can be implemented by anyone in about "15 minutes."

Defensive Implications

▶ Watch: Common mistake: Not realizing a metric is a proxy. (6:59)

Mark Overholser's talk provides critical insights that directly translate into actionable defensive strategies for security teams:

  1. Standardize Definitions Across the Organization: Defenders must proactively establish and disseminate clear, unambiguous definitions for key security terms, especially "incident," "severity," and phases of incident response (detection, acknowledgment, response, resolution, remediation). Adopting frameworks like the NIST Cybersecurity Framework or ISO 27000 series for definitions provides a common language that prevents miscommunication and ensures consistent measurement. This standardization is foundational for all subsequent defensive efforts.
  1. Prioritize and Implement Meaningful Metrics: Instead of relying on convenient but irrelevant metrics, security teams should focus on measuring what truly matters to their operational effectiveness and the business. Key metrics like Mean Time To Detection (MTTD), Mean Time To Acknowledgement (MTTA), and the various Mean Times To Response/Resolution/Remediation (MTRs) are crucial. By tracking these, defenders can:
  • Assess Visibility Gaps: A high MTTD indicates a need for better logging, monitoring tools, or threat intelligence.
  • Evaluate Staffing & Workload: A rising MTTA suggests a need for more analysts, automation, or process improvements.
  • Measure Incident Handling Efficiency: MTRs provide objective data on how quickly and effectively incidents are contained and resolved.
  • Leverage "Not Solvable" Closures: This specific closure reason is a powerful tool for justifying budget requests for new tooling (e.g., EDR, network telemetry, SIEM correlation rules) or training, as it quantifies the impact of current visibility gaps.
  1. Quantify Risk in Financial Terms: The most impactful defensive implication is the adoption of Overholser's financial risk quantification model: Risk = Probability (%) × Impact ($). Defenders should:
  • Estimate Probabilities: Assess the likelihood of specific threats materializing within a defined period (e.g., 90 days), using threat intelligence, vulnerability data, and internal incident history.
  • Monetize Impact: Work with business stakeholders (finance, legal, operations) to assign a credible dollar value to the impact of various security incidents (e.g., data breach costs, regulatory fines, reputational damage, downtime losses).
  • Justify Investments: Use the "Risk in Dollars" to perform cost-benefit analyses for proposed security controls. Presenting that a control costing $X will reduce a $Y risk by Z% makes a much stronger case for budget approval than abstract security scores. This aligns security spending with business objectives.
  1. Beware of Metric Manipulation (Goodhart's Law): Defenders must design metrics carefully to avoid unintentionally incentivizing undesirable behavior. When setting performance targets, consider the potential for gaming the system. Focus on outcomes and quality of work rather than just speed or quantity. For example, instead of just MTTA, also track the quality of initial triage or false positive rates.
  1. Embrace Iterative Improvement and Educated Guesswork: When data is incomplete, defenders should not be paralyzed. Overholser encourages making the "best guess" for unknown variables (e.g., incident start time, likelihood, impact). The key is to:
  • Start Measuring: Begin collecting and analyzing data, even if imperfect.
  • Re-evaluate Regularly: Periodically review assumptions and refine estimates based on new information and experience. This continuous feedback loop leads to increasingly accurate measurements and a more mature security program.

By implementing these defensive implications, security teams can move beyond reactive firefighting to a proactive, data-driven security posture that is not only more effective but also better understood and supported by the broader organization.

Key Takeaways

  • Data is essential for effective security operations: Without objective measurement, security decisions, resource allocation, and performance evaluations are based on guesswork.
  • Clear, consistent definitions are foundational: Establish and adhere to precise, agreed-upon definitions for terms like "incident," "severity," and incident response phases (detection, acknowledgment, response, resolution) to ensure consistent measurement and communication.
  • Measure what matters, not just what's convenient: Focus on metrics that directly answer questions about business value, risk reduction, and operational efficiency (e.g., MTTD, MTTA, MTRs), rather than easily available but irrelevant appliance statistics.
  • Quantify risk in monetary terms: Translate risk into financial impact (Probability % x Impact $) to speak the language of business, justify security investments, and perform effective cost-benefit analyses for controls.
  • Beware of Goodhart's Law: Do not turn measures into targets, as this can inadvertently distort behavior and undermine the integrity and utility of the metric itself.
  • Embrace iterative improvement and educated guesses: When data is imperfect, make your best guess, start measuring, and continuously refine your metrics and assumptions over time through experience and re-evaluation.

About the Speaker(s)

Mark Overholser is a seasoned cybersecurity professional with over a decade of experience in the field. He currently serves as a Technical Marketing Engineer (TME) and Threat Hunter at Corelight, a role he has held for six years. Prior to his tenure at Corelight, Mark began his career in IT at the age of 15, demonstrating a long-standing passion for technology and security. Beyond his corporate role, he has also had the distinct honor of serving as a threat hunter in the demanding environment of the Black Hat conference network operations center for several years, working alongside other prominent security professionals. His extensive experience spans both practical threat hunting and the strategic development of security measurement frameworks, making him a knowledgeable and practical voice in the cybersecurity community.

All talks from SAINTCON 2025