KubeRusted
Spinning up your cluster…

PCA exam prep — Prometheus Certified Associate

Free PCA (Prometheus Certified Associate) practice: the 5 official domains, exam-style questions, a timed practice exam and more.

Play the PCA map → · Official PCA exam page

Observability Concepts (18%)

Metrics, logs and events, traces and spans, push vs pull, service discovery, SLOs

Prometheus Fundamentals (20%)

Architecture, configuration and scraping, limitations, data model, exposition format

PromQL (28%)

Selecting data, rates, aggregation over time and dimensions, binary operators, histograms

Instrumentation and Exporters (16%)

Client libraries, instrumentation, exporters, naming metrics

Alerting & Dashboarding (18%)

Dashboards, alerting rules, Alertmanager, when and what to alert on

Practice questions

Which signal is best for "how many requests per second are failing right now"?

Answer: Metrics. Aggregated numeric time series answer rate questions cheaply.

Which signal is best for following one slow request across several services?

Answer: Traces. A trace links spans across services for one request.

What is an event, compared to a metric?

Answer: A discrete record of something that happened, with its own timestamp and details. Metrics summarize many events into numbers over time.

Why is high cardinality a problem in Prometheus?

Answer: Each unique label set is a separate series, consuming memory and storage. Avoid unbounded label values like user IDs or request IDs.

Which is an advantage of pull-based collection?

Answer: Prometheus knows a target is down when a scrape fails (the up metric). Pull also lets you scrape a target manually to debug it.

When is pushing metrics appropriate in the Prometheus world?

Answer: For short-lived batch jobs, via the Pushgateway. The Pushgateway is not a general-purpose push receiver.

What is an SLI?

Answer: A measured indicator of service level, e.g. the ratio of successful requests. SLI = what you measure; SLO = the target; SLA = the agreement.

What is an error budget?

Answer: The amount of unreliability allowed by the SLO over its window. With a 99.9% SLO, 0.1% of requests (or time) may fail.

Which service discovery mechanism lets Prometheus read targets from JSON/YAML files?

Answer: file_sd_configs. Another system can write the files; Prometheus reloads them automatically.

In kubernetes_sd_configs, which role discovers every Pod container port?

Answer: pod. Roles include node, service, pod, endpoints, endpointslice and ingress.

Which relabel action drops targets that do not match a regex?

Answer: keep. keep keeps only matching targets; drop removes matching ones.

What are labels starting with __meta_ used for?

Answer: Service discovery metadata available during relabeling, dropped afterwards. Use relabel_configs to copy useful __meta_ labels into real labels.

What is the difference between relabel_configs and metric_relabel_configs?

Answer: relabel_configs act on targets before scraping; metric_relabel_configs on scraped samples. Use metric_relabel_configs to drop expensive series at ingestion.

Which pillar explains why something is slow inside one process (CPU, memory hot paths)?

Answer: Profiles. Continuous profiling complements metrics, logs and traces.

Which component of the Prometheus ecosystem sends notifications?

Answer: Alertmanager. Prometheus evaluates alerting rules and forwards firing alerts to Alertmanager.

Where does the Prometheus server store data by default?

Answer: A local time series database (TSDB) on disk. Blocks are written to the --storage.tsdb.path directory.

Which flag controls how long Prometheus keeps local data?

Answer: --storage.tsdb.retention.time. There is also --storage.tsdb.retention.size.

Which setting sends samples to long-term storage such as Thanos Receive or Mimir?

Answer: remote_write. remote_read lets Prometheus query remote storage back.

What is the default scrape_interval if none is configured?

Answer: 1m. Many example configs set 15s explicitly, but the built-in default is one minute.

Which label does Prometheus attach to a target’s series from its address?

Answer: instance. job comes from the scrape job name.

What does the up metric equal when a scrape fails?

Answer: 0. up is 1 for a healthy scrape and 0 for a failed one.

Which metric types does the Prometheus data model define?

Answer: Counter; Gauge; Histogram; Summary. Meter and Timer are terms from other libraries, not Prometheus types.

In the text exposition format, what does a # TYPE line declare?

Answer: The metric type (counter, gauge, histogram, summary, untyped). # HELP adds the description.

What is federation in Prometheus?

Answer: One Prometheus scraping selected series from another’s /federate endpoint. Often used to pull aggregated series into a global view.

What is a known limitation of a single Prometheus server?

Answer: It is not clustered: storage is local to one node. HA is usually two identical servers; long-term storage goes to remote systems.

Why is Prometheus a poor fit for per-request billing data?

Answer: Scrapes sample counters periodically, so exactness is not guaranteed. Use an event-based system for data that must be 100% accurate.

How do you reload the configuration without restarting Prometheus?

Answer: Send SIGHUP or POST /-/reload (with --web.enable-lifecycle). The config is validated first; invalid config keeps the old one.

Which tool validates a Prometheus configuration or rule file?

Answer: promtool check config / promtool check rules. promtool also runs rule unit tests with promtool test rules.

What is the HTTP API path for instant queries?

Answer: /api/v1/query. /api/v1/query_range runs range queries.

What is the default global scrape_timeout?

Answer: 10s. A scrape is abandoned after scrape_timeout, which cannot be greater than scrape_interval.

What is OpenMetrics?

Answer: A standardized evolution of the Prometheus exposition format. It is supported by Prometheus scrapes and client libraries.

Which selector matches series where method is GET or POST?

Answer: http_requests_total{method=~"GET|POST"}. =~ is a regex match, anchored to the whole value.

Which selector excludes 2xx status codes?

Answer: http_requests_total{code!~"2.."}. !~ is a negative regex match.

What does the offset modifier do in `rate(http_requests_total[5m] offset 1h)`?

Answer: Evaluates the rate as it was one hour ago. Useful for comparing with an earlier period.

What does increase(http_requests_total[1h]) return?

Answer: The approximate number of requests in the last hour. increase is rate multiplied by the range in seconds (with extrapolation).

Why use irate instead of rate?

Answer: To see fast-moving spikes from the last two samples. rate is smoother and better for alerts and dashboards over time.

Which function should you use for the per-second change of a gauge?

Answer: deriv(). rate/increase assume counters and their resets.

Which query returns the 5 series with the highest request rate?

Answer: topk(5, rate(http_requests_total[5m])). topk keeps the original labels of the selected series.

What does `sum without (instance) (rate(x[5m]))` do?

Answer: Sums over instances, keeping every other label. without is the inverse of by.

What does count by (job) (up == 1) return?

Answer: The number of healthy targets per job. The comparison filters series; count counts them per job.

How do you return 1/0 instead of filtering when comparing, e.g. up > 0?

Answer: Use the bool modifier: up > bool 0. Without bool, comparison operators drop non-matching series.

Which query alerts when no series for job="api" exists at all?

Answer: absent(up{job="api"}). An empty result can’t trigger a comparison; absent returns 1 when nothing matches.

What is a subquery, e.g. max_over_time(rate(x[5m])[1h:1m])?

Answer: Evaluates an inner instant query over a range at a given resolution. Here: the maximum 5-minute rate seen over the last hour, sampled each minute.

Error ratio per job: which query divides correctly?

Answer: sum by (job) (rate(errors_total[5m])) / sum by (job) (rate(requests_total[5m])). Rate first, then aggregate both sides to the same labels.

When is group_left needed in a binary operation?

Answer: Many-to-one matching, where the left side has more series per match group. It also copies labels from the "one" side with group_left(label).

What does `on(instance)` do in a binary operation?

Answer: Matches series using only the instance label. ignoring(…) matches on all labels except the listed ones.

What does histogram_quantile(0.99, …) return when most requests fall in the highest finite bucket?

Answer: An estimate capped near that bucket’s upper bound — buckets limit accuracy. Quantile estimates are only as good as the bucket layout.

How do you compute the average request duration from a histogram?

Answer: rate(x_sum[5m]) / rate(x_count[5m]). Rate both, so the average reflects the recent window.

Which function returns the value of each series as of the last sample before now?

Answer: last_over_time(). Other *_over_time functions aggregate across the range instead.

What does timestamp(up) return?

Answer: The timestamp of each sample, in seconds. Useful, e.g., time() - timestamp(x) to see how stale a series is.

Which expression gives the number of seconds since a counter last reset?

Answer: time() - process_start_time_seconds. process_start_time_seconds is a standard client library metric.

What does resets(counter[1h]) count?

Answer: How many times the counter dropped (e.g. restarts) in the last hour. changes() counts value changes for gauges.

Why can’t you graph `http_requests_total[5m]` directly?

Answer: It is a range vector; graphs need instant vectors. Wrap it in a function such as rate().

What should a counter tracking requests be named?

Answer: http_requests_total. snake_case, a namespace prefix, and the _total suffix that marks a counter.

Which unit should a duration metric use?

Answer: Seconds. Prometheus conventions use base units: seconds and bytes.

A summary vs a histogram: which can be aggregated across instances for quantiles?

Answer: Histogram. Summary quantiles are computed in each client and cannot be combined.

Which labels are a bad idea on a request counter?

Answer: user_id; full request URL with IDs. Unbounded values explode cardinality; method and code are bounded.

Which metric type should track "items currently in a queue"?

Answer: Gauge. The value goes up and down.

What is the blackbox_exporter used for?

Answer: Probing endpoints from outside over HTTP, TCP, ICMP or DNS. It tests what users experience.

Which exporter exposes Linux host metrics such as CPU, memory and filesystems?

Answer: node_exporter. On Windows, windows_exporter fills that role.

What does kube-state-metrics expose?

Answer: The state of Kubernetes objects, e.g. Deployment replicas and Pod phases. cAdvisor (via the kubelet) provides container resource usage.

What is the recommended way to expose a constant like build version?

Answer: An info-style gauge set to 1 with the values as labels, e.g. app_build_info{version="1.2"}. Join it to other series when needed.

When instrumenting a library, which metrics are most useful to start with?

Answer: Requests, errors and duration (RED) for each operation. The RED method for services; USE (utilization, saturation, errors) for resources.

What does the _bucket{le="0.5"} series of a histogram count?

Answer: Observations less than or equal to 0.5. Buckets are cumulative; le="+Inf" equals _count.

Which alerting rule fields are required?

Answer: alert (the name); expr. for, labels and annotations are optional.

An alert with for: 5m has a true expression for 2 minutes. What is its state?

Answer: pending. It fires once the expression has stayed true for 5 minutes.

Which Alertmanager setting batches alerts that share labels into one notification?

Answer: group_by. group_wait, group_interval and repeat_interval tune the timing.

How do you mute alerts during planned maintenance?

Answer: Create a silence in Alertmanager (UI or amtool). Silences match alerts by labels for a time window.

What does an inhibition rule do?

Answer: Mutes some alerts while another, more important alert is firing. E.g. suppress all warnings for a cluster when "ClusterDown" fires.

Which Alertmanager setting resends a still-firing alert notification?

Answer: repeat_interval. Typically hours, to avoid notification fatigue.

Why does Alertmanager deduplicate alerts?

Answer: Highly available Prometheus pairs send the same alerts twice. Running Alertmanager in a cluster also deduplicates across its replicas.

Which tool lets you query and manage Alertmanager from the command line?

Answer: amtool. amtool silence add, amtool alert query, amtool check-config.

What should an alert annotation typically contain?

Answer: A human-readable summary/description and a runbook link. Labels route alerts; annotations inform the responder.

Which alert is better for paging?

Answer: High error ratio on user requests for 10 minutes. Page on symptoms users feel; causes go to dashboards or tickets.

How do you name a recording rule that aggregates request rates by job over 5 minutes?

Answer: job:http_requests:rate5m. Convention: level:metric:operations.

Where do you define which alerts go to which team?

Answer: The Alertmanager routing tree (route with matchers and receivers). Routes match on alert labels such as team or severity.

In Grafana, how do you make one dashboard work for every job?

Answer: A template variable (e.g. $job) used in the panel queries. Variables can be populated with label_values(up, job).

Which Grafana panel type suits a single current value like "requests per second now"?

Answer: Stat. Heatmaps suit histogram buckets over time.

A batch job runs for 20 seconds every night. How should it get its metrics into Prometheus?

Answer: Push them to a Pushgateway, which Prometheus scrapes. Short-lived jobs may finish before a scrape. The Pushgateway holds their last values for scraping.

Your SLO is 99.9% availability per 30 days. What is the error budget?

Answer: About 43 minutes of unavailability per 30 days. 0.1% of 30 days (43,200 min) ≈ 43.2 minutes.

Which labels does Prometheus add to every scraped series by default?

Answer: job; instance. job comes from the scrape config, instance from the target address; Kubernetes labels need service discovery + relabeling.

Why should you avoid a user_id label on a request counter?

Answer: Each value creates a new time series — unbounded cardinality blows up memory and storage. Series count = product of label values. Unbounded values belong in logs or traces.

Which query gives the per-second request rate over the last 5 minutes?

Answer: rate(http_requests_total[5m]). rate() handles counter resets and returns per-second values. delta() is meant for gauges.

Which query returns the 95th percentile latency across all instances?

Answer: histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket[5m]))). Rate the buckets, aggregate while keeping `le`, then estimate the quantile (0–1, not 0–100).

Total request rate per job, dropping every other label?

Answer: sum by (job) (rate(http_requests_total[5m])). Always rate first, then aggregate — summing raw counters breaks reset handling.

Which metric type fits "number of jobs currently waiting in the queue"?

Answer: Gauge. A queue length goes up and down, so it is a gauge. Counters only increase.

Which metric name follows Prometheus naming conventions?

Answer: http_request_duration_seconds. snake_case, base unit (seconds) as suffix, and _total reserved for counters.

You need CPU, memory and filesystem metrics from Linux hosts. What do you deploy?

Answer: node_exporter. node_exporter exposes host metrics; blackbox_exporter probes endpoints (HTTP, TCP, ICMP).

An alert must fire only after the condition has held for 10 minutes. Which field?

Answer: for: 10m. `for` keeps the alert pending until the expression has been true that long. The others are Alertmanager timing settings.

Which jobs belong to Alertmanager rather than Prometheus?

Answer: Grouping and deduplicating alerts; Silencing and inhibiting alerts. Prometheus evaluates rules and sends firing alerts; Alertmanager decides who gets notified and how.