devops/observability/prometheus sitereliabilityengineering telemetry

Core Idea

Prometheus is a pull-based metrics platform: it scrapes targets’ HTTP endpoints into its own time-series database, answers PromQL queries, and alerts when a metric crosses a threshold.

  • It watches numbers: CPU and memory, disk space, uptime, latency, error counts. It does not monitor events, traces, or logs.
  • Components: Prometheus server (scrape and store), exporters, pushgateway for short-lived jobs, Alertmanager for grouping/routing/deduping notifications, service discovery for targets, Grafana for dashboards.
  • See Pull Based and Push Based for the scrape model, Time Series Data & TSDB for storage, and Prometheus Metrics for the metric types.

Prometheus

Learn Prometheus Architecture: A Complete Guide

Prometheus is a Open-source monitoring solution that is responsible for collecting and aggregating metrics

  • Prometheus allows you to generate alerts when metrics reach a user specified threshold.
  • Prometheus collects metrics by scraping targets who expose metrics through an HTTP endpoint.
  • Scraped metrics are then stored in a time series database which can be queried using Prometheus’ built-in query language PromQl.
  • Prometheus uses its own specialized Time Series Database (TSDB) for storing and managing time-stamped data.
  • This TSDB is optimized for high-performance data ingestion, efficient storage, and fast querying of time series data, making it ideal for monitoring and alerting systems.

👉 Time Series Data & TSDB

Prometheus Moniters:

  • CPU/Memory Utilization
  • Disk space
  • Service Uptime
  • Application specific data
    • Number of exceptions
    • Latency
    • Pending Requests

Prometheus not monitor:

  • Events
  • SystemTraces
  • logs

Prometheus Architecture

Prometheus is a monitoring and alerting toolkit with a pull-based architecture designed for efficiency and scalability. Its main components include:
👉Pull Based and Push Based

  1. Prometheus Server: Scrapes metrics from targets and stores them in a time-series database for querying with PromQL.
  2. Exporters: Expose metrics from applications or systems in a Prometheus-compatible format.
  3. Push Gateway: Temporarily stores metrics from short-lived jobs, allowing Prometheus to scrape them.
  4. Alertmanager: Manages alerts from Prometheus, handling grouping, routing, deduplication, and notifications.
  5. Service Discovery: Dynamically identifies scrape targets from systems like Kubernetes, AWS, or Consul.
  6. Visualization: Uses PromQL for querying metrics and integrates with tools like Grafana for advanced dashboards.

Prometheus Metrics