platformengineering security sitereliabilityengineering

Source: PlatformCon 2026 Talk by Diptamay Sanyal (Principal Engineer, CrowdStrike)

Core Idea

Vendor severity is static and blind to your environment, so SOCs drown in alerts; Impact Lens is a streaming enrichment and scoring pipeline that turns the flood into a trusted, ranked queue.

  • The core problem: exploitation velocity is rising (exploited criticals up 105% year over year while EPSS scores fell, 42% of zero-days exploited before disclosure).
  • The Impact Lens framework and its architecture and data pipeline.
  • Key engineering trade-offs behind the design.

Most Security Operation Centers (SOCs) face a deluge of alerts where severity is static, assigned by vendors, and entirely disconnected from environmental reality. This talk presents Impact Lens, a streaming enrichment and scoring pipeline that sits on top of existing detections to turn a flood of alerts into a highly trusted, ranked queue.

1. The Core Problem

  • Exploitation Velocity: Industry data shows exploited critical CVSS parameters increased 105% year-over-year, while the average EPSS scores for those same parameters declined. Attackers are exploiting zero-days before public disclosure (up 42%).
  • Domain Silos: Attackers are leveraging valid account abuse (35% of incidents), making lateral cross-domain inclusion chains look completely benign when viewed within a single silo (identity, endpoint, or cloud).
  • CVSS Limitation: CVSS does not capture environmental context—such as whether a critical asset is internet-reachable, or if the associated identity is privileged.

2. The Impact Lens Framework

The platform pre-computes and bakes four core signals into each telemetry event before any detection rules fire:

  1. Asset Criticality: Identifies whether the target is a domain controller, CI/CD pipeline, production gateway, or merely a test box.
  2. Identity Exposure: Determines if the affected identity is a highly privileged service account or an unprivileged user.
  3. Blast Radius: A pre-computed graph property traversing end-dependency maps to determine how many downstream services or data stores are reachable if compromised.
  4. Reachability: Asks if the asset is internet-facing or isolated behind network segmentation. This field contributes the most to noise reduction. (e.g., A CVSS 9.8 on an completely unreachable asset is a standard maintenance task, not a live incident).

3. Architecture & Data Pipeline

The framework processes millions of events per second across three sequential stages:

  • Stage 1: Stateful Enrichment (Apache Flink & RocksDB)

    • Telemetry is keyed by Asset ID and enriched locally using a RocksDB state backend with microsecond latency (avoiding network hops).
    • State sizes hover around 1–2 TB per partition, necessitating isolated thread pools for memory flushing so background compactions never block the write path.
    • Uses a 10-minute TTL cache strategy. Expired keys pass to a side output with a rate-limited consumer to update the local RocksDB asynchronously, preventing thundering herd problems on the master KV store.
  • Stage 2: Context-Aware Detection & SIEM Fan-out

    • Enriched data streams directly to a unified Kafka topic where independent consumer groups handle SIEM indexing and the Detection Engine concurrently.
    • Detection rules become context-aware out of the box (e.g., only fire on CVSS > 9 if reachability > 0.7 and identity is privileged), killing noise at the root.
  • Stage 3: Alert Deduplication & LLM Summarization
     - Pass 1: Lightweight deterministic deduplication collapses identical CVSS/Asset IDs within a tumbling window.

    • Pass 2: Uses an embedding model (BGE-small-v1.5, 384-dimensions for optimal compute footprint) and OpenSearch K-NN (IVF index) to group semantically similar alerts stemming from the same attack chain.
       - Asynchronous LLM: High-ranking grouped incidents are routed to a 7B LLM running behind a proxy to generate human-readable summaries and action recommendations without blocking the main alert pipeline.

4. Key Engineering Trade-offs

  1. Enrich Before Detection: Contextualizing raw telemetry produces vastly quieter alert feeds than triaging post-detection output.
  2. Localize Pipeline State: Millions of events per second cannot tolerate network network requests; state must live adjacent to processing threads.
  3. Deterministic Fast-Paths First: Use math and graph algorithms for grouping/scoring; reserve expensive LLM inference for high-priority summaries.
  4. Design for Graceful Degradation: If the enrichment layer fails, the pipeline automatically falls back to raw CVSS scores. A noisier queue for a few minutes is always better than dropping data entirely.