Managed SOC Services
Aug 13, 2026
Karan Patel

The Moneyball Approach in Cybersecurity: Why It's Time to Think Differently

Learn how the Moneyball approach in cybersecurity replaces gut-feel spending with data-driven risk metrics, helping lean security teams outperform bigger budgets.

details hero

Baseball scouts spent a century evaluating players on things that felt meaningful: how a swing looked, how a player carried himself, how hard the ball sounded off the bat. The evaluations were made by experienced people with genuine expertise, and they were frequently wrong, because the traits being measured had a weaker relationship to winning games than everyone assumed.

The correction was not smarter scouts. It was a shift in what got measured. Teams that identified which statistics actually predicted run production, and which were expensive noise, could assemble competitive rosters at a fraction of the cost of teams still paying premiums for traits that looked impressive and did not correlate with outcomes.

Security operations has an uncomfortable amount in common with pre-Moneyball baseball. Considerable expertise, considerable spending, and a set of widely tracked metrics that have surprisingly little relationship to whether an organization is actually harder to compromise. This post is about which numbers are the security equivalent of batting average, which are on-base percentage, and how to shift from one to the other.

The Scouting Problem in Security Operations

Most security programs are evaluated on inputs and activity rather than on outcomes, for understandable reasons.

We Measure What Is Easy to Count

Alerts handled, events ingested, threats blocked, patches applied, training completions, tools deployed. Every one of these is easy to instrument and produces a satisfying upward trend. None of them tells you whether an attacker would succeed.

The parallel to batting average is close. It is a real statistic measuring a real thing, and it was used for decades as the primary evaluation metric despite correlating with run scoring less strongly than alternatives that were available and ignored.

Expertise Substitutes for Evidence

Security decisions are frequently made on the judgment of experienced practitioners, which is valuable and also fallible in specific, predictable ways. Recent incidents get overweighted. Vendor messaging shapes threat perception. Techniques that are prominent in the news feel more likely than techniques that are quietly more common in the environment.

None of this reflects poorly on the people involved. It reflects the well-documented behavior of expert judgment operating without a feedback loop.

The Feedback Loop Is Broken

Baseball has an unusually clean feedback mechanism: every decision produces an observable result within months, over thousands of repetitions. Security has the opposite. The most important outcome, the breach that did not happen, is unobservable. Feedback arrives rarely, late, and confounded by dozens of variables.

This is why security is unusually susceptible to spending on things that feel effective. Nothing contradicts the feeling until an incident does, and by then the attribution is muddled.

Organizations wanting to establish a measurable baseline before deciding what to change can start that work with FoxRadar360, where operational telemetry provides the observed data that gut assessment cannot.

Vanity Metrics: The Batting Averages of Security

Certain numbers appear in nearly every security report and carry almost no predictive value.

Total Events Ingested

A measure of your logging configuration and license tier. Larger is not better; it is simply larger. It appears in reporting because it is a big impressive number, not because anyone makes a decision from it.

Threats Blocked

Usually an aggregate of spam filtering, firewall denies, and automated endpoint quarantines. It measures commodity noise absorbed by controls that would have absorbed it regardless. Legitimate as an operational statistic, misleading as evidence of program effectiveness.

Alerts Handled

Reports on tuning quality more than security contribution. A program handling fifty thousand alerts monthly may simply have a noise problem. The relevant figures are how many required human judgment and how many were real.

Percentage of Alerts Closed

Approaches one hundred percent in any functioning operation, including one closing everything incorrectly.

Patch Compliance Percentage

Genuinely useful, frequently misused. Ninety-six percent patched sounds strong until you learn the remaining four percent includes the internet-facing systems with the longest exposure windows. Aggregate compliance obscures exactly where risk concentrates.

Training Completion Rates

Measures that people clicked through modules. Correlation with reduced susceptibility to actual phishing is weak enough that completion rate should never be presented as a risk indicator.

Tool Count and Coverage Claims

Owning more security products does not correlate with better outcomes, and licensed seats are not the same as deployed, tuned, and monitored assets.

The Metrics That Actually Predict Defensive Outcomes

The security equivalents of on-base percentage share a common trait: they measure things attackers actually have to defeat.

Detection Coverage by Adversary Technique

The single most predictive metric available. For each technique relevant to your environment, do you have detection logic, has it been validated to fire, and does it cover all the assets where the technique could be used?

This is predictive because it maps directly to what an attacker must do. Coverage gaps are the paths through your environment that generate no alert. Tracking coverage growth over time measures genuine program improvement in a way that alert counts never will.

Time to Detect, Measured From First Evidence

Not from when the alert fired, but from the earliest evidence of the activity that exists in your telemetry. The gap between those two timestamps is itself a measure of detection quality, and it is the number that determines how much an attacker accomplishes before you engage.

Time to Containment, Reported as a Distribution

Averages hide the incidents that matter. A mean of twenty minutes with a ninety-fifth percentile of nine hours describes a program that handles routine cases well and complex ones poorly, which is exactly the failure mode that produces breaches. Report percentiles and explain the outliers.

False Negative Rate From Sampling

The most important metric in security and the least frequently measured, because it requires deliberately looking for your own mistakes. Periodically sample closed alerts and re-examine them independently. What was dismissed that should not have been?

This is uncomfortable and irreplaceable. Everything else measures what you caught; this measures what you missed, which is the only thing an attacker cares about.

Validated Versus Configured Coverage

The percentage of your enabled detections that actually fire when the corresponding behavior is executed in a controlled test. The gap is consistently larger than teams expect, caused by field mapping changes, silent log source failures, and rules disabled during past noise reduction that nobody re-enabled.

Configured coverage is a claim. Validated coverage is a fact. Building validation testing into a regular cadence, rather than treating it as an occasional exercise, is a core part of how FoxRadar360 structures ongoing improvement in monitored environments.

Telemetry Completeness by Plane

For identity, endpoint, network, and cloud, what percentage of assets are sending the specific event types your detections depend on? Not whether a source is connected, but whether it sends the fields the logic requires. A partially reporting source produces confident silence.

Privilege Exposure

How many accounts hold standing administrative access, how long elevated permissions persist, and how much accumulated access exists relative to current role. This predicts blast radius, which determines whether a compromised account becomes an incident or a breach.

Attacker Interruption Rate in Testing

During purple team or adversary emulation exercises, at how many points in the attack chain did a control detect or interrupt the operator? A chain interrupted at three points is far more resilient than one interrupted once at the end, even if both technically caught the activity.

Finding Undervalued Assets in Your Own Program

The other half of the Moneyball insight was not just better measurement. It was recognizing that the market systematically overpaid for some traits and undervalued others, creating arbitrage for anyone willing to act on evidence.

Security has the same inefficiency.

Overvalued: New Tooling

The market prices novelty highly. A new platform in an emerging category commands premium spend while delivering marginal improvement over a properly tuned incumbent. Purchasing is fast and visible, which makes it the default response to any identified gap.

Undervalued: Tuning and Detection Engineering

Nobody announces that they spent a quarter improving detection quality on existing platforms. It produces no procurement event and no vendor relationship. It also delivers, reliably, more risk reduction per dollar than most acquisitions, because partially deployed and untuned tooling is the actual constraint in most programs.

Overvalued: Threat Intelligence Volume

Feed counts and indicator volumes are marketed as capability. Most indicators are stale, irrelevant to your environment, or duplicative. Intelligence that changes a detection or a decision has value; intelligence that fills a database does not.

Undervalued: Asset Inventory Accuracy

Deeply unglamorous, and foundational to everything. You cannot detect on assets you do not know exist, cannot patch them, and cannot assess them. Time spent reconciling inventory sources improves the effectiveness of every other control simultaneously.

Overvalued: Prevention Investment at the Margin

Prevention matters enormously up to a point, after which incremental spend delivers diminishing returns while detection and response capability remains thin. Many programs are heavily weighted toward preventing intrusion and lightly weighted toward what happens when prevention fails, which it will.

Undervalued: Containment Authority and Runbooks

Pre-agreed authority to isolate hosts and revoke sessions at any hour costs essentially nothing and directly determines how much damage a detected intrusion causes. It is a governance decision, not a purchase, which is precisely why it gets deferred.

Undervalued: Log Retention

Discovering an intrusion that began ninety days ago and finding thirty days of logs means scoping becomes guesswork. Retention is cheap relative to the cost of being unable to state what was accessed.

Building an Evidence-Driven Security Program

Shifting from instinct to evidence requires structure, not just intent.

Establish Baselines Before Changing Anything

Measure current detection coverage, response times by hour of day, telemetry completeness, and validated versus configured coverage. Without a baseline, no subsequent claim of improvement is verifiable, and the program has no way to distinguish real progress from activity.

Form Hypotheses and Test Them

Treat security investments as testable propositions. If a control is expected to reduce a specific exposure, state the expected effect and the metric that would demonstrate it, then check afterward. This is unusual in security and routine in disciplines that improve reliably.

Run Controlled Validation Regularly

Purple team exercises and automated adversary emulation are the closest thing security has to a clean feedback loop. They generate observable results for defensive decisions on a timescale short enough to learn from. Any program serious about evidence needs this running continuously rather than annually.

Report What You Missed

Include false negative sampling results in regular reporting. Programs that only report successes are optimizing for reassurance, and reassurance is not a security outcome. Leadership that sees honest miss data makes better resourcing decisions than leadership that sees only handled alert counts.

Resist Metrics That Cannot Change a Decision

Before adding a number to a report, ask what decision it would inform and what value would prompt a different choice. Metrics failing that test are consuming attention that better metrics need.

What This Looks Like in Practice

A program that made the shift reports differently. Instead of alerts handled and threats blocked, it reports detection coverage by technique with the specific gaps named, response time distributions segmented by hour of day, validation test results showing which detections fired and which did not, false negative sampling findings, and telemetry completeness by plane with known blind spots stated plainly.

It spends differently as well. Less on additional platforms, more on completing and tuning existing deployments. Less on intelligence volume, more on inventory accuracy and log retention. Less on marginal prevention, more on containment authority and validated detection.

The uncomfortable part is that this reporting looks worse before it looks better. Honest coverage measurement reveals gaps. False negative sampling finds misses. Validation testing shows that detections you believed in do not fire. That is the point. Those findings existed before you measured them; measurement simply moved them from unknown to addressable.

Teams that want an external read on where their measurable gaps actually sit, rather than relying on internal assessment alone, can work through that review with FoxRadar360 and use the results to redirect effort toward what demonstrably moves outcomes.

Final Thoughts

The Moneyball lesson was never that expertise does not matter. Experienced scouts remained valuable; they simply needed to be pointed at questions where human judgment outperforms statistics, while statistics handled the questions where they outperform intuition.

Security needs the same division. Experienced practitioners are irreplaceable for investigating ambiguous activity, understanding business context, and making judgment calls during incidents. They are less reliable, as everyone is, at estimating probability without data, at knowing which of their controls actually work, and at assessing what they missed.

Fill those gaps with measurement. Track detection coverage against techniques rather than alert volume. Report response times as distributions rather than averages. Sample your closed alerts to find what you got wrong. Validate that detections fire rather than trusting that they are enabled. Then direct spending toward the undervalued fundamentals: inventory accuracy, tuning, retention, and the authority to contain a threat at three in the morning without waiting for an approval meeting.

None of that produces an impressive slide. It produces a program that improves measurably each year, which is the only version of winning available in this game. To establish that measurement baseline and see where your current coverage genuinely stands, start a conversation with the team at FoxRadar360.

Your Threat-Free Future Is One Click Away

Let FoxRadar360 transform your business into a secure, monitored, and threat-resilient operation. Schedule your SOC demo in seconds, simple and stress-free.  

title-icon
Cloud Monitoring
title-icon
Incident Response
title-icon
Compliance Support
title-icon
Threat Intelligence
title-icon
Intelligent TDIR + CTEM
title-icon
SIEM Integration
title-icon
Endpoint Detection and Response
title-icon
Proactive Cyber Risk Management