Metrics, logs and events, traces and spans, push vs pull, service discovery, SLOs
Metrics — Numeric measurements over time — cheap to store, easy to aggregate and alert on. Great for "how much / how fast / how many errors"; less suited to per-request detail, which is what logs and traces are for. Docs
Traces & spans — A trace follows one request across services; each step is a span. Prometheus does not store traces (Jaeger/Tempo do), but exemplars can link a metric sample to a trace ID. Docs
Pull model — Prometheus scrapes targets; targets do not push to it. Pull makes "is the target up?" free (the `up` metric) and keeps targets simple. The Pushgateway exists for short-lived batch jobs only. Docs
Service discovery — Finds scrape targets automatically: kubernetes_sd, consul_sd, file_sd…. Discovery produces __meta_* labels; relabel_configs decide which targets to keep and what labels they get. Docs
SLI / SLO / SLA — Indicator (what you measure) → Objective (your target) → Agreement (the contract). e.g. SLI = share of requests under 300 ms, SLO = 99.9% over 30 days, SLA = refund if below 99.5%. The gap is your error budget. Docs
Prometheus Fundamentals (20%)
Architecture, configuration and scraping, limitations, data model, exposition format
Prometheus server — Scrapes, stores in a local TSDB, evaluates rules and serves PromQL. One binary: retrieval, TSDB, rule evaluation and HTTP API. Alertmanager, exporters and the Pushgateway are separate components. Docs
scrape_configs — prometheus.yml jobs: targets, scrape_interval, metrics_path, relabeling. Each job gets a `job` label and each target an `instance` label. Reload config with SIGHUP or POST /-/reload (with --web.enable-lifecycle). Docs
Labels — A time series = metric name + a unique set of key/value labels. http_requests_total{method="GET",status="200"}. Every new label value is a new series, so avoid unbounded values like user IDs (cardinality). Docs
Exposition format — The plain-text /metrics format with # HELP and # TYPE lines. One sample per line: name{labels} value [timestamp]. OpenMetrics is its standardized successor. Docs
Local storage limits — One node’s TSDB: not clustered, not for long-term retention or 100% accuracy. Use remote_write to Thanos, Cortex/Mimir and the like for durability and global view. Prometheus is for metrics — not billing-grade counts or logs. Docs
PromQL (28%)
Selecting data, rates, aggregation over time and dimensions, binary operators, histograms
Instant & range vectors — up{job="api"} is an instant vector; up{job="api"}[5m] is a range vector. Matchers: =, !=, =~ (regex), !~. Range vectors feed functions like rate(); you cannot graph them directly. Docs
rate() — Per-second average increase of a counter over a range, reset-aware. rate(http_requests_total[5m]). irate uses only the last two samples (spiky); increase = rate × range. Use rate on counters, never on gauges (deriv() is for gauges). Docs
sum by — Aggregate across series, keeping only the listed labels. sum by (job) (rate(http_requests_total[5m])). Also avg, min, max, count, topk, quantile; `without` drops labels instead. Docs
avg_over_time() — Aggregate each series across time: avg/max/min/sum/count_over_time. max_over_time(node_load1[1h]) — the highest load per series in the last hour. Docs
Vector matching — Binary operators match series by labels: on(), ignoring(), group_left. errors / requests divides series with identical labels; use `on(job)` or `ignoring(code)` to control matching, group_left for many-to-one. Docs
histogram_quantile() — Estimates a percentile from histogram buckets. histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket[5m]))) — keep the `le` label when aggregating. Docs
Counter — Only goes up (or resets to zero on restart) — requests, errors, bytes. Name it with a _total suffix and query it with rate(). Never use a counter for a value that can go down. Docs
Gauge — Goes up and down — temperature, queue length, memory in use. Query it directly or with *_over_time functions and deriv(). Docs
Histogram — Counts observations into buckets: _bucket{le=…}, _sum and _count. Aggregatable across instances, unlike a summary’s client-side quantiles. Choose bucket boundaries around your SLO. Docs
Client libraries — Official Go, Java, Python, Ruby and Rust libraries to instrument your own code. Define metrics in code, increment them where things happen, and expose /metrics over HTTP. Docs
Exporters — Translate third-party systems into Prometheus metrics: node_exporter, blackbox…. For things you cannot instrument directly. node_exporter covers host metrics, blackbox_exporter probes endpoints from outside. Docs
Metric naming — snake_case, a single base unit as suffix: http_request_duration_seconds. Use base units (seconds, bytes), _total for counters, and an application prefix. Don’t put label values in the name. Docs
Alerting & Dashboarding (18%)
Dashboards, alerting rules, Alertmanager, when and what to alert on
Alerting rules — A PromQL expr plus `for:` — pending, then firing. expr: job:errors:ratio5m > 0.05, for: 10m, labels: {severity: page}, annotations: {summary: …}. Loaded via rule_files. Docs
Recording rules — Precompute expensive queries into new series, named level:metric:operations. Speeds up dashboards and alerts that reuse the same aggregation. Docs
Alertmanager — Groups, deduplicates, routes, silences and inhibits alerts. Prometheus only fires alerts; Alertmanager’s routing tree sends them to receivers (email, Slack, PagerDuty) with group_by, group_wait and repeat_interval. Docs
Alert on symptoms — Page on what users feel (errors, latency), not on every cause. Every page should be urgent and actionable; causes (high CPU) belong on dashboards. Docs
Grafana — The usual dashboarding layer on top of Prometheus. Panels run PromQL against a Prometheus data source; template variables (e.g. $job) make dashboards reusable. Docs
Practice questions
Which signal is best for "how many requests per second are failing right now"?
Answer: Metrics. Aggregated numeric time series answer rate questions cheaply.
Which signal is best for following one slow request across several services?
Answer: Traces. A trace links spans across services for one request.
What is an event, compared to a metric?
Answer: A discrete record of something that happened, with its own timestamp and details. Metrics summarize many events into numbers over time.
Why is high cardinality a problem in Prometheus?
Answer: Each unique label set is a separate series, consuming memory and storage. Avoid unbounded label values like user IDs or request IDs.
Which is an advantage of pull-based collection?
Answer: Prometheus knows a target is down when a scrape fails (the up metric). Pull also lets you scrape a target manually to debug it.
When is pushing metrics appropriate in the Prometheus world?
Answer: For short-lived batch jobs, via the Pushgateway. The Pushgateway is not a general-purpose push receiver.
What is an SLI?
Answer: A measured indicator of service level, e.g. the ratio of successful requests. SLI = what you measure; SLO = the target; SLA = the agreement.
What is an error budget?
Answer: The amount of unreliability allowed by the SLO over its window. With a 99.9% SLO, 0.1% of requests (or time) may fail.
Which service discovery mechanism lets Prometheus read targets from JSON/YAML files?
Answer: file_sd_configs. Another system can write the files; Prometheus reloads them automatically.
In kubernetes_sd_configs, which role discovers every Pod container port?
Answer: pod. Roles include node, service, pod, endpoints, endpointslice and ingress.
Which relabel action drops targets that do not match a regex?
Answer: keep. keep keeps only matching targets; drop removes matching ones.
What are labels starting with __meta_ used for?
Answer: Service discovery metadata available during relabeling, dropped afterwards. Use relabel_configs to copy useful __meta_ labels into real labels.
What is the difference between relabel_configs and metric_relabel_configs?
Answer: relabel_configs act on targets before scraping; metric_relabel_configs on scraped samples. Use metric_relabel_configs to drop expensive series at ingestion.
Which pillar explains why something is slow inside one process (CPU, memory hot paths)?
Answer: Profiles. Continuous profiling complements metrics, logs and traces.
Which component of the Prometheus ecosystem sends notifications?
Answer: Alertmanager. Prometheus evaluates alerting rules and forwards firing alerts to Alertmanager.
Where does the Prometheus server store data by default?
Answer: A local time series database (TSDB) on disk. Blocks are written to the --storage.tsdb.path directory.
Which flag controls how long Prometheus keeps local data?
Answer: --storage.tsdb.retention.time. There is also --storage.tsdb.retention.size.
Which setting sends samples to long-term storage such as Thanos Receive or Mimir?
What is the default scrape_interval if none is configured?
Answer: 1m. Many example configs set 15s explicitly, but the built-in default is one minute.
Which label does Prometheus attach to a target’s series from its address?
Answer: instance. job comes from the scrape job name.
What does the up metric equal when a scrape fails?
Answer: 0. up is 1 for a healthy scrape and 0 for a failed one.
Which metric types does the Prometheus data model define?
Answer: Counter; Gauge; Histogram; Summary. Meter and Timer are terms from other libraries, not Prometheus types.
In the text exposition format, what does a # TYPE line declare?
Answer: The metric type (counter, gauge, histogram, summary, untyped). # HELP adds the description.
What is federation in Prometheus?
Answer: One Prometheus scraping selected series from another’s /federate endpoint. Often used to pull aggregated series into a global view.
What is a known limitation of a single Prometheus server?
Answer: It is not clustered: storage is local to one node. HA is usually two identical servers; long-term storage goes to remote systems.
Why is Prometheus a poor fit for per-request billing data?
Answer: Scrapes sample counters periodically, so exactness is not guaranteed. Use an event-based system for data that must be 100% accurate.
How do you reload the configuration without restarting Prometheus?
Answer: Send SIGHUP or POST /-/reload (with --web.enable-lifecycle). The config is validated first; invalid config keeps the old one.
Which tool validates a Prometheus configuration or rule file?
Answer: promtool check config / promtool check rules. promtool also runs rule unit tests with promtool test rules.
What is the HTTP API path for instant queries?
Answer: /api/v1/query. /api/v1/query_range runs range queries.
What is the default global scrape_timeout?
Answer: 10s. A scrape is abandoned after scrape_timeout, which cannot be greater than scrape_interval.
What is OpenMetrics?
Answer: A standardized evolution of the Prometheus exposition format. It is supported by Prometheus scrapes and client libraries.
Which selector matches series where method is GET or POST?
Answer: http_requests_total{method=~"GET|POST"}. =~ is a regex match, anchored to the whole value.
Which selector excludes 2xx status codes?
Answer: http_requests_total{code!~"2.."}. !~ is a negative regex match.
What does the offset modifier do in `rate(http_requests_total[5m] offset 1h)`?
Answer: Evaluates the rate as it was one hour ago. Useful for comparing with an earlier period.
What does increase(http_requests_total[1h]) return?
Answer: The approximate number of requests in the last hour. increase is rate multiplied by the range in seconds (with extrapolation).
Why use irate instead of rate?
Answer: To see fast-moving spikes from the last two samples. rate is smoother and better for alerts and dashboards over time.
Which function should you use for the per-second change of a gauge?
Answer: deriv(). rate/increase assume counters and their resets.
Which query returns the 5 series with the highest request rate?
Answer: topk(5, rate(http_requests_total[5m])). topk keeps the original labels of the selected series.
What does `sum without (instance) (rate(x[5m]))` do?
Answer: Sums over instances, keeping every other label. without is the inverse of by.
What does count by (job) (up == 1) return?
Answer: The number of healthy targets per job. The comparison filters series; count counts them per job.
How do you return 1/0 instead of filtering when comparing, e.g. up > 0?
Answer: Use the bool modifier: up > bool 0. Without bool, comparison operators drop non-matching series.
Which query alerts when no series for job="api" exists at all?
Answer: absent(up{job="api"}). An empty result can’t trigger a comparison; absent returns 1 when nothing matches.
What is a subquery, e.g. max_over_time(rate(x[5m])[1h:1m])?
Answer: Evaluates an inner instant query over a range at a given resolution. Here: the maximum 5-minute rate seen over the last hour, sampled each minute.
Error ratio per job: which query divides correctly?
Answer: sum by (job) (rate(errors_total[5m])) / sum by (job) (rate(requests_total[5m])). Rate first, then aggregate both sides to the same labels.
When is group_left needed in a binary operation?
Answer: Many-to-one matching, where the left side has more series per match group. It also copies labels from the "one" side with group_left(label).
What does `on(instance)` do in a binary operation?
Answer: Matches series using only the instance label. ignoring(…) matches on all labels except the listed ones.
What does histogram_quantile(0.99, …) return when most requests fall in the highest finite bucket?
Answer: An estimate capped near that bucket’s upper bound — buckets limit accuracy. Quantile estimates are only as good as the bucket layout.
How do you compute the average request duration from a histogram?
Answer: rate(x_sum[5m]) / rate(x_count[5m]). Rate both, so the average reflects the recent window.
Which function returns the value of each series as of the last sample before now?
Answer: last_over_time(). Other *_over_time functions aggregate across the range instead.
What does timestamp(up) return?
Answer: The timestamp of each sample, in seconds. Useful, e.g., time() - timestamp(x) to see how stale a series is.
Which expression gives the number of seconds since a counter last reset?
Answer: time() - process_start_time_seconds. process_start_time_seconds is a standard client library metric.
What does resets(counter[1h]) count?
Answer: How many times the counter dropped (e.g. restarts) in the last hour. changes() counts value changes for gauges.
Why can’t you graph `http_requests_total[5m]` directly?
Answer: It is a range vector; graphs need instant vectors. Wrap it in a function such as rate().
What should a counter tracking requests be named?
Answer: http_requests_total. snake_case, a namespace prefix, and the _total suffix that marks a counter.
Which unit should a duration metric use?
Answer: Seconds. Prometheus conventions use base units: seconds and bytes.
A summary vs a histogram: which can be aggregated across instances for quantiles?
Answer: Histogram. Summary quantiles are computed in each client and cannot be combined.
Which labels are a bad idea on a request counter?
Answer: user_id; full request URL with IDs. Unbounded values explode cardinality; method and code are bounded.
Which metric type should track "items currently in a queue"?
Answer: Gauge. The value goes up and down.
What is the blackbox_exporter used for?
Answer: Probing endpoints from outside over HTTP, TCP, ICMP or DNS. It tests what users experience.
Which exporter exposes Linux host metrics such as CPU, memory and filesystems?
Answer: node_exporter. On Windows, windows_exporter fills that role.
What does kube-state-metrics expose?
Answer: The state of Kubernetes objects, e.g. Deployment replicas and Pod phases. cAdvisor (via the kubelet) provides container resource usage.
What is the recommended way to expose a constant like build version?
Answer: An info-style gauge set to 1 with the values as labels, e.g. app_build_info{version="1.2"}. Join it to other series when needed.
When instrumenting a library, which metrics are most useful to start with?
Answer: Requests, errors and duration (RED) for each operation. The RED method for services; USE (utilization, saturation, errors) for resources.
What does the _bucket{le="0.5"} series of a histogram count?
Answer: Observations less than or equal to 0.5. Buckets are cumulative; le="+Inf" equals _count.
Which alerting rule fields are required?
Answer: alert (the name); expr. for, labels and annotations are optional.
An alert with for: 5m has a true expression for 2 minutes. What is its state?
Answer: pending. It fires once the expression has stayed true for 5 minutes.
Which Alertmanager setting batches alerts that share labels into one notification?
Answer: group_by. group_wait, group_interval and repeat_interval tune the timing.
How do you mute alerts during planned maintenance?
Answer: Create a silence in Alertmanager (UI or amtool). Silences match alerts by labels for a time window.
What does an inhibition rule do?
Answer: Mutes some alerts while another, more important alert is firing. E.g. suppress all warnings for a cluster when "ClusterDown" fires.
Which Alertmanager setting resends a still-firing alert notification?
Answer: repeat_interval. Typically hours, to avoid notification fatigue.
Why does Alertmanager deduplicate alerts?
Answer: Highly available Prometheus pairs send the same alerts twice. Running Alertmanager in a cluster also deduplicates across its replicas.
Which tool lets you query and manage Alertmanager from the command line?
Where do you define which alerts go to which team?
Answer: The Alertmanager routing tree (route with matchers and receivers). Routes match on alert labels such as team or severity.
In Grafana, how do you make one dashboard work for every job?
Answer: A template variable (e.g. $job) used in the panel queries. Variables can be populated with label_values(up, job).
Which Grafana panel type suits a single current value like "requests per second now"?
Answer: Stat. Heatmaps suit histogram buckets over time.
A batch job runs for 20 seconds every night. How should it get its metrics into Prometheus?
Answer: Push them to a Pushgateway, which Prometheus scrapes. Short-lived jobs may finish before a scrape. The Pushgateway holds their last values for scraping.
Your SLO is 99.9% availability per 30 days. What is the error budget?
Answer: About 43 minutes of unavailability per 30 days. 0.1% of 30 days (43,200 min) ≈ 43.2 minutes.
Which labels does Prometheus add to every scraped series by default?
Answer: job; instance. job comes from the scrape config, instance from the target address; Kubernetes labels need service discovery + relabeling.
Why should you avoid a user_id label on a request counter?
Answer: Each value creates a new time series — unbounded cardinality blows up memory and storage. Series count = product of label values. Unbounded values belong in logs or traces.
Which query gives the per-second request rate over the last 5 minutes?
Answer: rate(http_requests_total[5m]). rate() handles counter resets and returns per-second values. delta() is meant for gauges.
Which query returns the 95th percentile latency across all instances?
Answer: histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket[5m]))). Rate the buckets, aggregate while keeping `le`, then estimate the quantile (0–1, not 0–100).
Total request rate per job, dropping every other label?
Answer: sum by (job) (rate(http_requests_total[5m])). Always rate first, then aggregate — summing raw counters breaks reset handling.
Which metric type fits "number of jobs currently waiting in the queue"?
Answer: Gauge. A queue length goes up and down, so it is a gauge. Counters only increase.
Which metric name follows Prometheus naming conventions?
Answer: http_request_duration_seconds. snake_case, base unit (seconds) as suffix, and _total reserved for counters.
You need CPU, memory and filesystem metrics from Linux hosts. What do you deploy?
An alert must fire only after the condition has held for 10 minutes. Which field?
Answer: for: 10m. `for` keeps the alert pending until the expression has been true that long. The others are Alertmanager timing settings.
Which jobs belong to Alertmanager rather than Prometheus?
Answer: Grouping and deduplicating alerts; Silencing and inhibiting alerts. Prometheus evaluates rules and sends firing alerts; Alertmanager decides who gets notified and how.