devops/observability/prometheus sitereliabilityengineering telemetry
Core Idea
Prometheus is a pull-based metrics platform: it scrapes targets’ HTTP endpoints into its own time-series database, answers PromQL queries, and alerts when a metric crosses a threshold.
- It watches numbers: CPU and memory, disk space, uptime, latency, error counts. It does not monitor events, traces, or logs.
- Components: Prometheus server (scrape and store), exporters, pushgateway for short-lived jobs, Alertmanager for grouping/routing/deduping notifications, service discovery for targets, Grafana for dashboards.
- See Pull Based and Push Based for the scrape model, Time Series Data & TSDB for storage, and Prometheus Metrics for the metric types.
Prometheus
Prometheus is a Open-source monitoring solution that is responsible for collecting and aggregating metrics
- Prometheus allows you to generate alerts when metrics reach a user specified threshold.
- Prometheus collects metrics by scraping targets who expose metrics through an HTTP endpoint.
- Scraped metrics are then stored in a time series database which can be queried using Prometheus’ built-in query language PromQl.
- PromQL Basics : Querying basics | Prometheus
- PromQL CeatSheet : PromLabs | PromQL Cheat Sheet
- Prometheus uses its own specialized Time Series Database (TSDB) for storing and managing time-stamped data.
- This TSDB is optimized for high-performance data ingestion, efficient storage, and fast querying of time series data, making it ideal for monitoring and alerting systems.
Prometheus Moniters:
- CPU/Memory Utilization
- Disk space
- Service Uptime
- Application specific data
- Number of exceptions
- Latency
- Pending Requests
Prometheus not monitor:
- Events
- SystemTraces
- logs
Prometheus Architecture
Prometheus is a monitoring and alerting toolkit with a pull-based architecture designed for efficiency and scalability. Its main components include:
👉Pull Based and Push Based
- Prometheus Server: Scrapes metrics from targets and stores them in a time-series database for querying with PromQL.
- Exporters: Expose metrics from applications or systems in a Prometheus-compatible format.
- Push Gateway: Temporarily stores metrics from short-lived jobs, allowing Prometheus to scrape them.
- Alertmanager: Manages alerts from Prometheus, handling grouping, routing, deduplication, and notifications.
- Service Discovery: Dynamically identifies scrape targets from systems like Kubernetes, AWS, or Consul.
- Visualization: Uses PromQL for querying metrics and integrates with tools like Grafana for advanced dashboards.