Prometheus & Grafana
The standard open-source monitoring and observability stack — Prometheus collects metrics from infrastructure and applications; Grafana visualizes them in dashboards.
Prometheus and Grafana form the dominant open-source observability stack for cloud-native applications. Prometheus is a time-series metrics database that scrapes metrics from instrumented services and infrastructure, evaluates alerting rules, and stores data for querying with its PromQL language. Grafana is the visualization layer — connecting to Prometheus (and dozens of other data sources) to build real-time dashboards for system health, application performance, and business metrics. Together they are deployed at the majority of companies running Kubernetes workloads and appear consistently in DevOps, SRE, and platform engineering job descriptions.
Typical time to job-readiness: ~3 weeks.
Learning Prometheus & Grafana
Beginner
Run Prometheus and Grafana locally with Docker Compose and scrape metrics from a simple service. Learn the four core Prometheus metric types: Counter, Gauge, Histogram, and Summary. Build a basic Grafana dashboard with panels showing request rate, error rate, and latency.
Intermediate
PromQL for querying metrics (rate(), increase(), histogram_quantile()), configuring alerting rules and routing to Alertmanager (PagerDuty, Slack), and instrumenting your own application with the Prometheus client library for your language. Learn the USE method (Utilization, Saturation, Errors) for systematic infrastructure monitoring.
Advanced
Prometheus federation and remote write for large-scale deployments, recording rules for expensive queries, Grafana Loki for log aggregation alongside metrics, and Tempo for distributed tracing. SRE interviews involve designing an observability strategy — be ready to discuss SLIs, SLOs, and error budgets alongside the tooling.
Key concepts
- Prometheus scrapes metrics from HTTP endpoints exposed by services (/metrics path) at a configured interval
- Metric types: Counter (only increases), Gauge (up and down), Histogram (distribution), Summary (quantiles)
- PromQL: Prometheus query language — rate() for counters, avg_over_time() for gauges, histogram_quantile() for latency
- Alertmanager: routes firing alerts to email, Slack, PagerDuty, and handles deduplication and silencing
- Grafana connects to Prometheus (and other sources) to build dashboards from PromQL queries
- SLIs and SLOs: define what 'good' looks like before setting alerts — metrics without context are noise
Common interview topics
- What is the difference between a Counter and a Gauge in Prometheus
- How would you write a PromQL query to get the request error rate over the last 5 minutes
- How do you set up alerting in Prometheus and route alerts to the right team
- What is an SLO and how do you build dashboards around it
- Walk me through how you'd investigate a latency spike using Prometheus and Grafana