Overview
Curated: · Written: · Reviewed:
Security monitoring is an evidence and response system
Security monitoring turns trustworthy telemetry into decisions that reduce risk. Logging is the lifecycle of generating, transmitting, normalizing, storing, accessing, analyzing, retaining and disposing of event records. Monitoring combines those records with asset, identity, vulnerability and threat context to identify conditions worth investigation. A SIEM can centralize search and correlation, but buying one does not create detection capability: owners, use cases, data contracts, reliable pipelines, tested analytics and response paths do.
Design from decisions and threats
Start with critical business services, data, identities, trust boundaries and credible abuse cases. For each detection or investigation question, identify the events and fields required, which system can authoritatively produce them, expected volume and latency, retention need, owner and response action. Useful records normally include event and observation time, source, actor and authentication context, action, target, outcome, reason, tenant or boundary, request/session/trace identifiers and schema version. Application audit events often provide business and authorization context unavailable from infrastructure logs. Never log passwords, session tokens, private keys or unnecessary personal and payload data; mask, tokenize or omit sensitive fields by design.
Coverage spans identity and privilege changes, cloud and SaaS control planes, applications and data stores, endpoints, networks, containers, orchestration, CI/CD, secrets and security controls. Centralization enables cross-source correlation, but raw evidence and provenance must remain recoverable. Normalize semantics without erasing source-specific meaning. Synchronize clocks where possible and retain original timestamps, timezone, receipt time and known uncertainty. Correlation by stable identifiers is stronger than timestamp proximity alone.
Make telemetry trustworthy and survivable
Treat the telemetry plane as a security-critical production system. Authenticate and encrypt transport, separate administrative duties, restrict and monitor read access, make destructive changes difficult, preserve immutable or independently controlled copies where risk requires, and define legal/privacy-aware retention and deletion. Protect parsers and viewers from log injection and untrusted content. Backpressure, throttling, sampling and cost controls must fail visibly and must not silently discard the exact high-risk events a detection needs.
Measure source coverage, freshness, completeness, parse failures, duplicates, clock skew, queue lag, storage pressure and end-to-end canary arrival. Alert on logging being disabled, retention reduced, collectors reconfigured or high-value sources going silent. A dashboard that is green while a tenant, region or schema is missing is false confidence. Document gaps explicitly and provide alternate evidence paths.
Engineer and operate detections
A detection has a threat hypothesis, required telemetry, query or model, window, thresholds, exclusions, severity, owner, runbook, expected evidence and tests. Version and review it like code. Validate with historical replay, synthetic events, controlled simulations and production-safe exercises; test positive, negative, boundary, delayed, duplicate and missing-data cases. ATT&CK can organize behaviors and detection strategies, but mapped technique labels do not prove effective coverage. Prevention and detection complement each other: an alert neither blocks an action nor proves malicious intent.
Triage combines signal fidelity, affected asset/data criticality, identity privilege, exposure, scope, confidence and current impact. Correlate independent evidence, preserve raw events and distinguish fact from inference. Tune noisy rules by finding root causes, adding reliable context, fixing data and splitting use cases—not by permanently suppressing unexplained alerts. Threat hunting is hypothesis-driven analysis that can discover gaps and produce detections; it is not random searching or proof that an environment is clean.
Measure outcomes rather than alert volume alone: detection coverage with tested evidence, data quality and freshness, time to acknowledge and reach a defensible disposition, escalation precision, missed or reopened incidents, recurrence, runbook effectiveness and corrective-action completion. Segment metrics by scenario and severity and resist incentives that encourage premature closure. When monitoring fails, record the exact sources, fields, tenants, regions and times affected, lower confidence, use alternate evidence, restore visibility without destroying artifacts and treat intentional tampering as a potential incident.
Detection coverage is a claim that has to be tested rather than asserted. A pipeline that ingests every log and has thirty rules may still miss the technique an attacker would actually use, and nobody discovers that during the incident. Map rules to the techniques you care about, then generate each one deliberately in a controlled way — an atomic test, a purple-team exercise — and confirm the alert fires with enough context to act on. A rule that fires with a hostname and no user, process, or parent process shifts the investigation back to the analyst and lengthens every response that depends on it.
Log integrity matters as much as log content when the logs are evidence. An attacker with access to a host can usually edit or delete its local logs, so the record that matters is the one shipped off the host promptly to storage the host cannot write to. Ship rather than pull where possible, alert on a source that stops reporting since silence is indistinguishable from a quiet system, and retain long enough for the investigation window that real intrusions require — dwell times measured in months mean thirty days of retention regularly leaves an investigation with no data from the period that matters.
Time and identity are what make correlation possible, and both are commonly broken. Hosts with unsynchronised clocks produce a timeline that cannot be assembled, so discipline every source to a common time source and record the offset. Identity is harder: an event attributed to a service account, a shared role, or a NAT address stops the investigation at the point where it needs to name a person, so carry the originating identity through automation and network translation where the architecture allows, and know in advance which of your sources cannot answer that question.
Worked example: coverage is a test, not a rule count
A SIEM with thirty rules still misses credential dumping if nobody ever fired that technique on purpose. Run the atomic test; record what the alert actually contains.
| technique | rule exists | atomic test fired | alert fields | usable |
|---|---|---|---|---|
| brute force (T1110) | yes | yes | user, source IP, count | yes |
| OS credential dump (T1003) | yes | yes | hostname only | no |
| inbound C2 (T1071) | no | n/a | — | gap |
Thirty rules, two tested, one usable. Silence on T1003 is a broken detection, not a quiet estate. That table is the interview.
