← CYBERNETICSINTERN.COMCYBERNETIC INTERN
DetectionDefensive guide

The detection audit you can run in a week

Five checks that tell you whether your detections work, each producing evidence rather than a coverage number. No vendor needed, and none takes a quarter.

Researched and drafted with AI assistance, then reviewed and edited by Shreyas Lipare before publication. Every source below was checked against the original.

Somebody asks how good your detection is. You have a number ready, and if you are honest, you know the number is softer than it sounds when you say it out loud.

The number counts rules that exist and carry a label. That is easy to measure, which is why it gets measured. Whether any of those rules would catch anything is hard to measure, and it is the only part the person asking cares about.

Here are five checks that produce evidence instead. None needs a vendor, a budget line or a quarter.

The core issue

Every detection is a chain of assumptions, and a coverage number tests none of them.

The event has to be generated. It has to reach your platform. The rule has to still match the shape of the data. It has to run often enough to matter. And when it fires, somebody has to see it and know what to do.

A rule can exist, be well written, carry a technique label, and fail at any one of those points without producing a single symptom. The console stays green either way, because nothing in the console is watching the chain. It is watching the rule.

That is why detection programmes tend to degrade quietly rather than break loudly, and why the degradation is usually discovered during an incident.

Why this matters

The failure mode is silence, and silence reads as success. A broken detection and a working detection with nothing to catch produce identical output: nothing. Every other part of a security programme fails noisily. This one does not, so it needs deliberate checking in a way that a firewall or a backup does not.

Coverage numbers create a false sense of position. They are not dishonest, they are answering a question about configuration and being read as an answer about capability. The gap between those two things is where the incident happens.

The frameworks moved and most programmes have not noticed. If your coverage model is built on ATT&CK data sources, that model was deprecated. In the October 2025 release MITRE states that detections in techniques have been replaced with Detection Strategies, and data sources were deprecated across the Enterprise, Mobile and ICS domains [2]. The data sources page carries the deprecation notice [3]. Anything built on that structure is now built on something MITRE has moved on from.

I went to cite the old model and found it retired. That is worth checking against your own tooling before you next present a coverage slide.

Technical breakdown: the five checks

Run them in order. Each one is cheap, each produces a written answer, and each is useless as a feeling.

1. Does the data still arrive?

The check. For every log source feeding a detection, find the most recent event. Not the source’s health status, the actual newest record. Sort ascending and look at the bottom of the list.

What bad looks like. Any source whose newest event is older than its normal interval. A source that reports healthy while sending nothing is common, because health usually means the connection is up rather than that data is flowing.

Most detections fire when something bad arrives, and almost none fire when nothing arrives at all, which is the state an attacker is working towards. That asymmetry is covered in alerting on log source silence.

The joint guidance from ASD’s ACSC, CISA, FBI, NSA and international partners makes centralised logging a baseline rather than an aspiration, and frames it around living off the land techniques, where the absence of malware means the log is the only evidence there is [1].

2. Does the rule still match the data?

The check. Take a sample of rules and, for each condition, confirm the field it references still exists and still carries the values the rule expects. Query the field directly rather than reading the rule.

What bad looks like. A field that returns nothing. A rule matching a field that no longer exists does not error, it matches nothing, and the rule stays enabled and green throughout. This is the quietest failure in the whole chain and it usually arrives with a vendor upgrade. See why the same field has three names.

While you are there, confirm the data exists at all. A rule that reads process command lines is inert if command line logging was never enabled, and that is a setting rather than a default on most estates. More in you cannot detect what you never collected.

3. Has it ever fired, and did anyone check on purpose?

The check. Sort every rule by the date it last fired. Then pick your ten highest-value rules and either search historical data for the pattern they describe, or generate the benign version of the behaviour in a controlled place and watch what happens.

What bad looks like. Rules with no last-fired date at all. Those are not quiet successes, they are unanswered questions, because a working rule and a broken one look the same from the console. The reasoning is in an untested detection is a hypothesis.

Test the whole path, not the match. A rule matching is not the same as somebody seeing an alert, and the pipeline, the routing and whatever suppression exists all sit in between.

4. How long does it actually take?

The check. For each of your top detections, find the ingest delay on its source and add the schedule interval. Write the total next to the rule.

Ingest delay is rarely documented anywhere, so measure it rather than asking. Most platforms record both the time an event happened and the time it was indexed. Subtract one from the other across a day of data and take the worst case, not the average, because the worst case is what an attacker gets.

What bad looks like. Any total you have never calculated, which is most of them. A rule scheduled hourly against a source running fifteen minutes behind can be seventy five minutes behind reality, and neither number describes that on its own. Whether that is acceptable depends entirely on what it is chasing, which is the point: a detection has two clocks.

This is the check most teams have never run, and it is the one that most often changes a roadmap. A detection slower than the attack it describes is documentation.

5. When it fires, does the right thing happen?

The check. Take twenty recent alerts. For each, count how many separate consoles an analyst had to open to close it, and confirm the severity matched what should have happened rather than how alarming the technique sounds.

What bad looks like. Alerts that begin with a pivot, because every field the alert omits gets rebuilt by hand once per firing. A rule firing twenty times a day with sparse context is a two hundred minute problem rather than a twenty alert problem. See the cost of an alert is the pivot.

Also bad: an estate where almost everything is medium. That is not a severity scheme, it is a field. Severity should answer what happens when this fires, which is why the same match can be an emergency on a domain controller and routine on a build agent. More in severity is a routing decision.

The sixth question, which is really about the other five

Can you prove what any of this looked like six months ago?

If detections live only in a console, the answer is no. There is no history, so a broken rule cannot be traced to the edit that broke it, and an incident review asking whether a rule was live in March gets a guess. Detections belong in version control, and the migration is unpleasant if you wait and nearly free if you start with the next rule you write.

There is no attack-chain diagram in this post because there is no adversary sequence in it. This is the plumbing underneath the diagrams.

What defenders should do Monday morning

Immediate, within 24 hours. Run check one against every source feeding a detection you would name in a board report. It is a single query per source and it is the only check that can find a problem that is currently active. If you find a silent source, that is an incident until proven otherwise.

Near term, this week. Run checks two through four and write the answers down in one place. Resist fixing anything until all three are done, because the value is in seeing the pattern rather than in the individual findings. Teams that fix as they go usually stop after the first interesting thing.

Structural, this quarter. Make checks one and three continuous rather than annual: alert on source silence, and record the last-fired date for every rule where somebody will see it. Then decide what your coverage number is actually allowed to claim, and change how it is reported if the honest version is smaller. If you are a small team with no central logging at all, CISA publishes Logging Made Easy as a free log management option aimed at exactly that gap [4].

Checklist

  • Newest actual event confirmed for every source feeding a named detection
  • Alerting in place for source silence, not only for source health
  • Fields referenced by sampled rules confirmed to exist and carry expected values
  • Command line and other optional logging confirmed enabled where rules depend on it
  • Every rule has a last-fired date, and rules with none are listed
  • Top ten detections tested end to end, through routing, not only matched
  • Ingest delay plus schedule interval calculated and recorded per key detection
  • Twenty recent alerts measured for how many consoles a close required
  • Severity distribution checked, and a mostly-medium estate treated as a finding
  • Detections in version control, or a dated plan to get them there
  • Coverage reporting reworded to claim only what the evidence supports

Why nobody runs these

None of this is difficult. That is the uncomfortable part.

Every check above is a query, a sort, or a count, and none needs a product nobody owns. They go unrun because they produce numbers that are worse than the number currently on the slide, and because there is no external forcing function. Nobody audits whether your detections work. Regulators ask whether you have a SIEM. Frameworks ask whether you map to techniques. Both questions are about existence, and existence is exactly what these five checks refuse to accept as an answer.

So the incentive runs the wrong way, and the profession has quietly agreed to measure the thing that is easy instead of the thing that matters. A coverage percentage is a statement about your configuration. It has never been a statement about what would happen.

The five answers are worse. They are also true, and they are the only version of this you can defend when somebody eventually asks whether the detection worked.

FREQUENTLY ASKED

How long does this actually take?
The first four checks are a day of work each at most, and less if your data is in one place. The fifth is longer because it usually surfaces process problems rather than technical ones. The point of framing it as a week is that none of this needs a project, a budget line, or a vendor engagement. If a check cannot be run in a day, the answer to that check is almost certainly no.
We have a coverage percentage from our platform. Is that not the same thing?
It answers a different question. A coverage number counts rules that exist and carry a label. None of these five checks can be passed by a rule existing. They ask whether data arrives, whether the rule still matches, whether it has ever fired, how long it takes, and whether anybody acts. A programme can score highly on coverage and fail all five.
What if the answers are bad?
Then you have found the work, which is the purpose. Bad answers here are more useful than good ones because they are specific: this log source has been silent for eleven days, these forty rules have never fired, this alert takes ninety minutes to reach a human. Those are fixable statements. A vague sense that detection could be better is not.
Does this replace purple teaming or breach and attack simulation?
No, and it is not competing with them. Those exercise detections against adversary behaviour, which is a stronger test than anything here. This is the layer underneath: whether the plumbing works at all. Running an attack simulation against a pipeline that has been dropping a log source for a fortnight tells you something you could have learned for free.
Our detections are in a vendor platform we cannot export. Does any of this apply?
The first four do, because they are about behaviour rather than authorship. The fifth is harder, and that difficulty is itself the finding. If you cannot answer what a rule looked like six months ago, you have a dependency worth naming in a risk register rather than a detail worth shrugging at.

REFERENCES

  1. [1]ASD's ACSC, CISA, FBI, and NSA, with the support of International Partners Release Best Practices for Event Logging and Threat DetectionCybersecurity and Infrastructure Security Agency (CISA) · Published August 21, 2024 · Accessed August 22, 2026PRIMARY
  2. [2]Updates: October 2025 (ATT&CK v18)MITRE ATT&CK · Published October 1, 2025 · Accessed August 22, 2026PRIMARY
  3. [3]Data SourcesMITRE ATT&CK · Accessed August 22, 2026PRIMARY
  4. [4]Logging Made EasyCybersecurity and Infrastructure Security Agency (CISA) · Accessed August 22, 2026PRIMARY

BEFORE YOU GO

Was this useful?

Tell me what you'd change, what was unclear, or what you'd want covered next. Replies shape what gets written.

SEND FEEDBACK ↗

The newsletter

Follow along

Notes between posts, and whatever I'm breaking in the lab.

LINKEDIN ↗