These Prometheus interview questions cover what SRE, DevOps and platform interviews in 2026 probe: how the pull model and service discovery work, which metric type to use and why, how to write PromQL that is actually correct, and how to build alerts that wake people up only when users are hurting. Interviewers rarely stop at "what is Prometheus"; they ask you to explain why rate() handles counter resets, why you cannot average a p99, and what you would do when a single label change takes the server down. The 55 questions below run from fundamentals through PromQL, Alertmanager, Prometheus 3.x, the Grafana stack, long-term storage, SLO burn-rate alerting and eleven real-world scenarios.
How to use this guide:
- Freshers and junior engineers are usually tested on the fundamentals and metric types: pull vs push, exporters, counters vs gauges, a basic
rate()query. Be able to write a scrape config and an error-rate query from memory. - Mid-level DevOps and SRE engineers get PromQL depth (histograms, vector matching, recording rules), Alertmanager routing and Kubernetes monitoring, plus at least one debugging scenario.
- Senior and platform roles are pushed on cardinality control, long-term storage (Thanos, Mimir, Cortex), high availability, SLO design, Prometheus 3.x and OpenTelemetry, and dashboards as code.
- Answer every question in the same shape: one-line answer, the mechanism, then the trade-off. For Grafana interview questions in particular, talk about how a dashboard is used during an incident, not just which panel type it is.
Contents
- Prometheus fundamentals (Q1βQ9)
- Metric types and PromQL interview questions (Q10βQ20)
- Alerting rules and Alertmanager interview questions (Q21βQ26)
- Prometheus 3.x and OpenTelemetry (Q27βQ31)
- Grafana interview questions: Loki, Tempo, Alloy and dashboards (Q32βQ35)
- Scale, long-term storage, Kubernetes and SLOs (Q36βQ42)
- AI in observability (Q43βQ44)
- Real-world scenario questions (Q45βQ55)
- Key takeaways
- Interview preparation checklist
- FAQ
Prometheus fundamentals
1. What is Prometheus and what is its data model?
Answer: Prometheus is an open-source monitoring and alerting toolkit (a CNCF graduated project, originally built at SoundCloud) that scrapes numeric metrics from targets over HTTP, stores them in a local time-series database and lets you query and alert on them with PromQL. Its data model is simple: every time series is identified by a metric name plus a set of key-value labels, and holds a stream of (timestamp, float value) samples. http_requests_total{method="GET", status="500", service="checkout"} is one series; change any label value and you have a different series.
Know its limits too. Prometheus is built for operational metrics and deliberately favours reliability over perfect accuracy: it is not a log store, not an event database and not suitable for per-request billing, where a lost scrape would mean lost money.
Interview tip: Say "labels are dimensions, and every unique combination is a new series". That sentence sets up the cardinality question that usually follows.
2. Pull vs push: why does Prometheus pull, and when do you use the Pushgateway?
Answer: Prometheus pulls (scrapes) because it gives the monitoring system control. It decides scrape timing, so a misbehaving app cannot flood it; a failed scrape is itself a signal (the up metric drops to 0); you can curl a target's /metrics endpoint to see exactly what Prometheus sees; and running a second Prometheus against the same targets for testing or HA needs no change to the apps.
The Pushgateway exists for short-lived batch jobs that finish before a scrape could catch them: the job pushes its final metrics, and Prometheus scrapes the gateway. It is not a general push proxy. It never expires what was pushed, so a dead job's last metrics look alive forever, and it becomes a single point of failure. For service-level metrics from many instances, keep pulling. Where push is genuinely needed (edge sites, firewalled networks), use remote write from a Prometheus in agent mode or an OpenTelemetry Collector or Grafana Alloy instead.
3. What are exporters? Name the ones you have used.
Answer: An exporter is an adapter that translates a third-party system's internal stats into the Prometheus exposition format on a /metrics endpoint, for software you cannot instrument directly. Common ones: node_exporter (Linux host CPU, memory, disk, filesystem, network), blackbox_exporter (probes HTTP, TCP, ICMP and DNS from the outside), postgres_exporter and mysqld_exporter, redis_exporter, the JMX exporter for Java apps, and cloud exporters for provider metrics. In Kubernetes, kube-state-metrics and cAdvisor (built into the kubelet) play the same role.
Your own services should not need an exporter: instrument them with a client library (Go, Java, Python, .NET and others) or OpenTelemetry SDKs, so the metrics describe business behaviour, not just host stats.
4. How does service discovery work, and what is relabeling?
Answer: In dynamic environments targets come and go, so Prometheus discovers them instead of using static lists. Built-in service discovery mechanisms include Kubernetes (pods, services, endpoints, nodes, ingresses), cloud providers (EC2, Azure, GCE), Consul, DNS and file-based SD. Each discovered target arrives with metadata labels such as __meta_kubernetes_pod_label_app or __meta_ec2_tag_Environment.
Relabeling turns that metadata into the target's final labels, and decides which targets to keep. relabel_configs run before the scrape (keep only pods with a scrape annotation, set namespace and pod labels, rewrite the address); metric_relabel_configs run after the scrape on every sample (drop an expensive metric, drop a high-cardinality label). Labels starting with __ are discarded after relabeling.
5. Walk me through the Prometheus server's architecture and storage.
Answer: One Prometheus binary contains service discovery, the scrape manager, the TSDB, the rule evaluator (recording and alerting rules), the PromQL engine and an HTTP API plus UI. Alertmanager is a separate process.
The TSDB keeps the most recent data (roughly the last two hours) in an in-memory head block, protected by a write-ahead log on disk so a crash does not lose it. Head data is periodically cut into immutable two-hour blocks on disk, each with its own index and chunks, and the compactor merges blocks into larger ones over time. Retention defaults to 15 days and is set by time or size. Each Prometheus is a single node by design: no clustering, no replication. That aids reliability and is why long-term storage systems exist.
targets --scrape--> [ Prometheus ]
| SD | TSDB | rules |
| |
PromQL API alerts
| v
Grafana Alertmanager --> Slack,
PagerDuty
6. What is Grafana's role, and how does it relate to Prometheus?
Answer: Grafana is the visualisation and exploration layer. It does not store Prometheus metrics; it sends PromQL to Prometheus (or a compatible backend such as Mimir or Thanos) and renders the result as dashboards. Its value is that one UI spans many data sources: Prometheus metrics, Loki logs, Tempo traces, cloud monitoring and SQL databases, with links between them so an engineer can jump from a latency spike to the matching logs and traces. Grafana also has its own alerting engine, which can evaluate rules across data sources (see Q26). Prometheus collects, stores, evaluates and alerts on metrics; Grafana turns them into SLO views, on-call dashboards and exploration workflows.
7. What do the job, instance and up labels tell you, and what are scrape_interval and evaluation_interval?
Answer: For every scrape, Prometheus attaches job (the scrape config name) and instance (host:port of the target), and records synthetic series: up (1 if the scrape succeeded, 0 if not), scrape_duration_seconds and scrape_samples_scraped, among others. up == 0 is the most basic "target down" alert, and scrape_samples_scraped is an early warning for cardinality growth.
scrape_interval is how often targets are scraped and evaluation_interval is how often rules are evaluated; both default to one minute, and many teams set 15 or 30 seconds. Shorter intervals give finer resolution at the cost of more samples and more storage. Keep the interval consistent within a job so rate() windows behave predictably.
8. What is cardinality and why is it a risk?
Answer: Cardinality is the number of unique time series, which is the product of the distinct values across a metric's labels. Every active series costs memory in the head block, index space and query time. A metric with labels for method (5 values), status (10) and endpoint (50) is already 2,500 series per instance; add user_id or request_id and it becomes unbounded, and Prometheus runs out of memory.
The rule: labels must have a small, bounded set of values. User IDs, email addresses, full URLs with IDs, trace IDs, timestamps and raw error messages belong in logs or traces, not labels. Find offenders with the TSDB status page in the Prometheus UI, promtool tsdb analyze, or a query such as topk(10, count by (__name__) ({__name__=~".+"})) (expensive; run it carefully).
9. What are the four Prometheus metric types?
Answer:
| Type | Behaviour | Example | Query with |
|---|---|---|---|
| Counter | Only goes up; resets to zero on restart | http_requests_total | rate(), increase() |
| Gauge | Goes up and down; a current value | node_memory_MemAvailable_bytes, queue depth | Raw value, avg_over_time(), predict_linear() |
| Histogram | Counts observations into buckets, plus sum and count | http_request_duration_seconds | histogram_quantile() |
| Summary | Client-side pre-computed quantiles, plus sum and count | Per-instance latency quantiles | Read quantile series directly |
Prometheus itself stores everything as series of floats (native histograms aside, see Q28); the types are a contract between the client library and the person writing queries. The most common mistake is querying a counter's raw value, which just shows an ever-increasing line.
Metric types and PromQL interview questions
10. Why must you always apply rate() to a counter, and how are counter resets handled?
Answer: A counter's absolute value is meaningless on its own: it depends on when the process last started. What matters is how fast it grows. rate() computes the per-second increase over a window and detects resets: if a sample is lower than the previous one, it assumes the process restarted and treats the drop as a reset to zero rather than a negative change. That is why rate() and increase() must be applied before aggregation. rate(sum(http_requests_total)[5m:]) is wrong because summing across pods hides one pod's reset inside a total that merely dips, producing false negatives or spikes. The correct form is sum(rate(http_requests_total[5m])).
Interview tip: "Rate then sum, never sum then rate" is the single most commonly tested PromQL rule.
11. Histogram vs summary: which do you choose?
Answer: Choose a histogram in almost every case. A histogram exposes bucket counters (_bucket{le="0.1"}, le="0.25" and so on, cumulative), plus _sum and _count. Because buckets are counters, you can aggregate them across instances and compute any quantile at query time with histogram_quantile(). The cost is that accuracy depends on bucket boundaries, so choose them around your SLO thresholds.
A summary computes quantiles (for example the 0.99) inside the client. It is accurate for that single instance, but quantiles cannot be aggregated: averaging the p99 of ten pods does not give the fleet's p99. You also cannot choose new quantiles after the fact. Use summaries only when you need an accurate quantile from a single process and will never aggregate. Native histograms (Q28) remove most of the bucket-choice pain of classic histograms.
12. What naming conventions and instrumentation methods do you follow?
Answer: Naming: snake_case, a namespace prefix (checkout_), base units in the name (_seconds, _bytes, never milliseconds or megabytes), _total for counters, and a ratio named _ratio holding 0β1. Labels describe dimensions of the same measurement; if summing across a label makes no sense, it should be a separate metric.
Instrumentation: use RED for request-driven services (Rate, Errors, Duration) and USE for resources (Utilisation, Saturation, Errors) such as CPU, disks, connection pools and queues. Add a few business metrics (orders placed, payments failed) because they often reveal problems that infrastructure metrics miss.
Real-world example: Consider a retailer's checkout service. RED metrics show error rate and latency per endpoint; a USE view of the database connection pool shows saturation; and checkout_orders_total dropping while error rate looks normal reveals a silent failure, such as a payment provider returning success pages with no order created.
13. What is the difference between an instant vector and a range vector? Give a PromQL example.
Answer: An instant vector is a set of series with one sample each at the evaluation time: http_requests_total{status="500"}. A range vector is a set of series with all samples within a time window: http_requests_total{status="500"}[5m]. Functions like rate() take a range vector and return an instant vector, which is what graphs and alerts need.
Classic example: rate(http_requests_total{status="500"}[5m]) gives the per-second rate of 5xx responses averaged over the last five minutes, per series. For a fleet view: sum by (service) (rate(http_requests_total{status=~"5.."}[5m])).
14. rate vs irate vs increase: when do you use each?
Answer:
rate(x[5m]): average per-second increase across the whole window, extrapolated to the window edges. Smooth and robust. Use it for alerts, recording rules and most dashboards.irate(x[5m]): per-second increase using only the last two samples in the window. Highly responsive, very spiky, and it ignores everything else in the range. Use it only for zoomed-in, fast-moving graphs, never for alerts: a brief spike between two scrapes can fire, and a sustained problem can be missed depending on which two samples land last.increase(x[1h]): total increase over the window, essentiallyrate()multiplied by the window's seconds. Use it for human-readable totals ("errors in the last hour"). Because of extrapolation it can return non-integers for an integer counter; that is expected.
15. How do you calculate a p99 latency from a histogram?
Answer: Rate the bucket counters, aggregate while keeping le, then apply histogram_quantile():
histogram_quantile(0.99,
sum by (le, service) (
rate(http_request_duration_seconds_bucket[5m])
)
)
Three things interviewers check: the le label must survive aggregation (drop it and the function has nothing to work with); you rate before aggregating; and the result is an estimate via linear interpolation within a bucket, so if your buckets are 0.5s and 1s, a "p99 of 0.73s" really means "somewhere between 0.5 and 1". Average latency comes from rate(_sum[5m]) / rate(_count[5m]). For SLOs it is often better to ask "what fraction of requests were under 300ms" using the le="0.3" bucket divided by _count, which is exact if 0.3 is a bucket boundary.
16. Explain PromQL aggregation operators and by vs without.
Answer: Aggregation operators (sum, avg, min, max, count, topk, bottomk, quantile, count_values, stddev, group) collapse many series into fewer. by (service) keeps only the listed labels; without (instance, pod) keeps everything except the listed labels. without is safer in recording rules because labels added later (such as cluster or team) are preserved automatically.
Examples: sum by (namespace) (rate(container_cpu_usage_seconds_total[5m])) for CPU cores per namespace; topk(5, sum by (pod) (container_memory_working_set_bytes)) for the five heaviest pods; count by (version) (build_info) to see how many instances run each version during a rollout.
17. How does vector matching work? When do you need group_left?
Answer: Binary operators between two instant vectors match series with identical label sets by default. on (labels) matches only on the listed labels; ignoring (labels) matches on all except those. That is enough for one-to-one matching. When one side has many series per match and the other has one, you need group_left (many on the left) or group_right, and you can copy labels across from the "one" side.
The classic use is enriching metrics with metadata from an info-style metric:
sum by (namespace, pod) (
rate(container_cpu_usage_seconds_total[5m])
)
* on (namespace, pod) group_left (node)
kube_pod_info
This adds the node label to per-pod CPU. If the right side has duplicates for a match, the query fails with "many-to-many matching not allowed", usually caused by a duplicate kube-state-metrics deployment or a pod restart producing two info series.
18. How do you write an error-ratio query correctly, and how do you choose the range window?
Answer: Divide the sum of error rates by the sum of all request rates, aggregated by the same labels:
sum by (service) (
rate(http_requests_total{code=~"5.."}[5m]))
/
sum by (service) (
rate(http_requests_total[5m]))
Never average per-pod ratios (a quiet pod with one failed request out of one would dominate). For the window: it should cover at least four scrape intervals, so with a 15s scrape, [1m] is the practical minimum and [5m] a common default. A window shorter than two scrapes returns nothing because rate() needs two samples. In Grafana, use $__rate_interval, which picks a window that stays valid at the dashboard's current step and zoom level.
A subtle point since Prometheus 3.0: range selectors are left-open, so a sample exactly at the window's start is excluded. Queries whose window exactly equals the scrape interval can lose a sample; another reason to stay at four intervals or more.
19. What do absent(), offset, predict_linear() and subqueries do?
Answer:
absent(up{job="payments"})returns 1 when no matching series exists. Normal alerts cannot fire on data that is missing, soabsent()(orabsent_over_time()) catches "the exporter vanished" or "a label was renamed".offset 1wshifts a selector into the past, for week-over-week comparisons:rate(x[5m]) / rate(x[5m] offset 1w).predict_linear(node_filesystem_avail_bytes[6h], 4*3600) < 0fits a linear regression to a gauge and predicts its value four hours ahead: the standard "disk will fill" alert, better than a static threshold.- Subqueries,
max_over_time(rate(x[5m])[1h:1m]), evaluate an inner expression at a resolution over a range so you can apply a_over_timefunction to it. Useful, but expensive; move frequent ones into recording rules.
20. What are recording rules and how do you name them?
Answer: Recording rules pre-compute expensive or frequently used expressions at every evaluation interval and store the result as a new series. Benefits: dashboards load quickly, alert expressions stay simple, and SLO calculations over long windows become feasible because they run on pre-aggregated data. The convention is level:metric:operations:
groups:
- name: http
rules:
- record: service:http_requests:rate5m
expr: |
sum by (service) (
rate(http_requests_total[5m]))
Here service is the aggregation level, http_requests the metric (with _total dropped after rating) and rate5m the operation. Rules in a group run sequentially, so a rule can build on one defined earlier in the same group. Validate files with promtool check rules and unit-test them with promtool test rules in CI.
Alerting rules and Alertmanager interview questions
21. How does alerting work end to end in Prometheus?
Answer: Prometheus evaluates alerting rules on every evaluation interval. When an expression returns series, each becomes an alert in pending state; if it keeps returning results for the for duration, it becomes firing and Prometheus sends it to Alertmanager, re-sending periodically while it fires. Alertmanager then does the human-facing work: deduplicates identical alerts (for example from two HA Prometheus servers), groups related alerts into one notification, applies inhibitions and silences, and routes each group to a receiver (Slack, email, PagerDuty, Opsgenie, Microsoft Teams, webhooks). When the expression stops returning results, the alert resolves and Alertmanager can send a resolved notification.
Interview tip: Keep the split clear: Prometheus decides whether something is wrong; Alertmanager decides who hears about it, how and how often.
22. What makes a good alerting rule?
Answer: A good alert is about user-visible symptoms, actionable, and has context. In practice:
- alert: CheckoutHighErrorRate
expr: service:http_errors:ratio_rate5m{service="checkout"}
> 0.05
for: 10m
keep_firing_for: 5m
labels:
severity: page
team: payments
annotations:
summary: "Checkout 5xx ratio above 0.05"
runbook_url: "https://wiki.example.internal/rb/checkout"
for avoids paging on a single noisy evaluation; keep_firing_for stops flapping when a metric hovers at the threshold; severity and team labels drive routing; annotations carry the summary, a dashboard link and a runbook. Page on symptoms (error rate, latency, saturation that will soon hurt users); send causes (one pod restarting, high CPU) to tickets or dashboards. Every page should answer: what is broken, how badly, and what do I do first.
23. Explain Alertmanager's routing tree and its timing settings.
Answer: Alertmanager has one root route; child routes match on labels (matchers: [team="payments", severity="page"]) and the first matching child wins unless continue: true lets an alert keep matching siblings. Each route can set:
group_by: which labels define a group, for example[alertname, cluster, service]. One notification per group instead of one per pod.group_wait(default 30s): how long to wait before the first notification for a new group, so alerts arriving together are batched.group_interval(default 5m): how long to wait before notifying about new alerts added to an existing group.repeat_interval(default 4h): how often to re-send an unchanged, still-firing group.
route:
receiver: default-slack
group_by: [alertname, cluster, service]
routes:
- matchers: [severity="page"]
receiver: pagerduty-oncall
- matchers: [team="data"]
receiver: data-slack
Test routing with amtool config routes test before deploying; a wrong matcher silently sends pages to a Slack channel nobody reads.
24. What are inhibition rules and silences, and how do they differ?
Answer: Inhibition is automatic, rule-based suppression: while a source alert fires, mute target alerts that share specified labels. Typical rules: if ClusterUnreachable fires for a cluster, mute every service alert for that cluster; if a critical alert fires, mute the warning version of the same alert.
inhibit_rules:
- source_matchers: [alertname="ClusterUnreachable"]
target_matchers: [severity=~"warning|page"]
equal: [cluster]
A silence is manual and time-bounded: an engineer mutes alerts matching given labels for a maintenance window or a known incident, created in the Alertmanager UI or with amtool. Rules keep evaluating either way; only notifications stop. Silences should always have an expiry and a comment saying who and why; a forgotten wide silence is a classic way to miss a real outage.
25. How do you make Prometheus and Alertmanager highly available?
Answer: Prometheus HA is simple and a little surprising: run two identical Prometheus servers scraping the same targets with the same rules. Both fire the same alerts; there is no leader election. Deduplication happens downstream: Alertmanager deduplicates identical alerts, and for queries a layer like Thanos Querier or Mimir deduplicates the two series using an external replica label (replica="a"/"b"). Each server's data will differ slightly because scrapes happen at slightly different moments; that is accepted.
Alertmanager HA uses a gossip-based cluster (the --cluster.* flags): run several instances, point every Prometheus at all of them (not at a load balancer), and the cluster shares silences and notification state so only one notification goes out. A common mistake is putting Alertmanagers behind a load balancer, which breaks deduplication.
26. Grafana Alerting vs Prometheus plus Alertmanager: how do you choose?
Answer: Grafana Alerting evaluates Grafana-managed alert rules that can query any data source (Prometheus, Loki, SQL, cloud monitoring) and combine them with expressions. Its building blocks are alert rules, contact points (where notifications go), notification policies (a routing tree, conceptually like Alertmanager's), silences and mute timings for recurring quiet periods. It can also manage data source-managed rules that live in Mimir or Loki rulers.
Choose Prometheus rules plus Alertmanager when alerts are mostly on Prometheus metrics, you want rules version-controlled next to the service and evaluated close to the data, and you need alerting to keep working even if Grafana is down. Choose Grafana Alerting when alerts span several data sources or teams prefer a UI. Many organisations run both. Whichever you choose, avoid defining the same alert in two places.
Prometheus 3.x and OpenTelemetry
27. What changed in Prometheus 3.0?
Answer: Prometheus 3.0 (November 2024) was the first major release in years. Key changes:
- New web UI with a PromLens-style query tree view.
- UTF-8 metric and label names allowed by default (Q29).
- Native OTLP ingestion: Prometheus can receive OpenTelemetry metrics directly (Q30).
- Remote Write 2.0 support (Q31).
- Native histograms continued to mature; they later became stable in the 3.x line (Q28).
- Breaking changes: range selectors became left-open;
holt_winterswas renameddouble_exponential_smoothingand moved behind thepromql-experimental-functionsflag; scrapes now require a validContent-Typeunlessfallback_scrape_protocolis set; agent mode is enabled with--agentand the remote-write receiver with--web.enable-remote-write-receiverrather than feature flags.
Interview tip: Mentioning the scrape-protocol strictness shows you have actually upgraded: badly behaved custom exporters that send no or wrong content type stop being scraped after the upgrade.
28. What are native histograms, and what is their status?
Answer: A native histogram is a single series whose samples are whole histograms, with exponentially spaced buckets chosen automatically by a resolution "schema", instead of one series per le bucket. Benefits: much higher resolution at a fraction of the series count, no need to guess bucket boundaries up front, and buckets that remain aggregatable. Queries drop the le handling:
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds[5m])))
Plus functions such as histogram_count(), histogram_sum() and histogram_fraction(0, 0.3, ...). Status: according to the Prometheus documentation, native histograms are a stable feature from v3.8; the old native-histograms feature flag is deprecated and from v3.9 is a no-op, and scraping them must be enabled with the scrape_native_histograms configuration setting. Client libraries must expose them (usually via the protobuf exposition format), and your remote storage and Grafana version must support them, so check each layer before migrating.
29. What does UTF-8 support for metric names mean in practice?
Answer: Before 3.0, names were limited to letters, digits, underscores and colons, so OpenTelemetry names like http.server.request.duration were converted to http_server_request_duration. Prometheus 3.x accepts UTF-8 names, so dots and other characters can be preserved. In PromQL such names must be quoted inside the braces:
{"http.server.request.duration", "service.name"="cart"}
Teams that need the old behaviour can set metric_name_validation_scheme: legacy. The practical interview point: allowing UTF-8 does not mean you should rename everything. Dashboards, recording rules and downstream systems that expect underscore names will break if names change underneath them, so decide on one convention per environment and handle translation deliberately (see Q30 and Q53).
30. What is OpenTelemetry's relationship to Prometheus? How does OTLP ingestion work?
Answer: They are complementary CNCF projects. OpenTelemetry is a vendor-neutral standard for producing and transporting telemetry (APIs, SDKs, semantic conventions, the OTLP protocol and the Collector) for traces, metrics and logs. Prometheus is a metrics backend: storage, PromQL and alerting. Many organisations instrument with OpenTelemetry SDKs and store metrics in Prometheus-compatible systems.
There are two integration paths. The Collector (or Grafana Alloy) can scrape Prometheus endpoints and forward via remote write. Or Prometheus can receive OTLP directly: start it with --web.enable-otlp-receiver and send to /api/v1/otlp/v1/metrics. OTLP is push, so the Prometheus docs recommend enabling an out-of-order time window, choosing a name translation_strategy, and promoting selected resource attributes (such as service.name or k8s.namespace.name) to labels with promote_resource_attributes. Other resource attributes land in the target_info metric, which you join to when needed. Interviewers like to hear that push via OTLP loses the free up signal, so you need another way to detect a silent service.
31. What is remote write, what does Remote Write 2.0 add, and what is agent mode?
Answer: Remote write streams samples from Prometheus's WAL to an external endpoint (Mimir, Thanos Receive, Cortex, managed Prometheus services) in near real time, with queues, retries and sharding. It is how most long-term and global-view architectures receive data.
Remote Write 2.0 adds a string-interning symbols table (smaller payloads), metric metadata (type, help, unit), exemplars, native histograms and created timestamps. Its specification page still marks it as experimental (a release candidate), so check that both sender and receiver support it and expect 1.0 to stay the safe default for a while.
Agent mode (prometheus --agent) runs Prometheus as a scrape-and-forward agent: it scrapes and remote-writes but has no local querying, rules or long-term storage. Use it on edge clusters or many small Kubernetes clusters that ship everything to a central backend.
Grafana interview questions: Loki, Tempo, Alloy and dashboards
32. How do Loki and Tempo fit with Prometheus, and what are exemplars?
Answer: Grafana Labs' open-source stack is often summarised as LGTM: Loki for logs, Grafana for visualisation, Tempo for traces, Mimir for metrics (and Pyroscope for continuous profiling). Loki borrows Prometheus's label model: it indexes only a small set of labels and stores log lines compressed in object storage, queried with LogQL ({app="checkout"} |= "timeout", and metric queries such as sum by (app) (rate({app="checkout"} |= "error" [5m]))). The same cardinality discipline applies: never use a request ID as a Loki label. Tempo stores traces in object storage and is queried with TraceQL.
Exemplars tie them together: a histogram observation can carry a trace ID, so a point on a Grafana latency panel links straight to an example slow trace in Tempo. Prometheus exemplar storage is enabled with a feature flag; check the current docs for your version. Using matching labels (service, namespace) across metrics, logs and traces is what makes the jump from one to the other work.
33. What is Grafana Alloy, and what happened to Grafana Agent?
Answer: Grafana Alloy is Grafana Labs' distribution of the OpenTelemetry Collector, with native pipelines for both OpenTelemetry and Prometheus data, covering metrics, logs, traces and profiles. It uses a component-based configuration language in which you wire components together (for example discovery.kubernetes into prometheus.scrape into prometheus.remote_write). Grafana Agent (Static, Flow and Operator) reached end of life on 1 November 2025 and no longer receives security or bug fixes; Grafana's guidance is to migrate to Alloy, and Alloy includes conversion tooling for Agent, Prometheus, Promtail and OpenTelemetry Collector configs.
34. How do you manage Grafana dashboards as code?
Answer: Dashboards edited by hand in the UI drift, break silently when a metric is renamed and vanish when an instance is rebuilt. Options, from simplest:
- File provisioning: Grafana loads dashboard JSON and data source YAML from disk at startup; in Kubernetes a sidecar loads dashboards from ConfigMaps, which is how kube-prometheus-stack ships its dashboards.
- Generators: Jsonnet with Grafonnet, or the Grafana Foundation SDK (Go, Java, PHP, Python, TypeScript), to build dashboards and alert rules from code with reusable panels.
- Terraform: the Grafana provider manages folders, dashboards, data sources, alert rules and contact points alongside the rest of your infrastructure (see the Terraform interview questions guide for the workflow).
- Git Sync (Grafana 12): connects dashboards and folders to a GitHub repository so UI saves become branches and pull requests. Grafana documents it as experimental, alongside the new dashboard schema v2, so treat it as something to trial outside production for now.
Whichever method, review dashboard changes in pull requests and lint the PromQL they contain, just like application code.
35. What makes a good Grafana dashboard?
Answer: It answers one question for one audience. A service dashboard starts with the RED summary (request rate, error ratio, latency percentiles) and SLO status at the top, then drills into dependencies and resources lower down. Use template variables ($cluster, $namespace, $service) populated with label_values(), but never a variable over a high-cardinality label like pod in a large cluster without filtering first. Set units and thresholds on every panel, use $__rate_interval in rate queries, link panels to logs, traces and runbooks, and add deployment annotations so "what changed" is visible on the graph.
Scale, long-term storage, Kubernetes and SLOs
36. Thanos vs Grafana Mimir vs Cortex: how do they compare?
Answer: All three solve the same problems a single Prometheus cannot: long-term retention in cheap object storage, a global query view across many Prometheus servers, HA deduplication and multi-tenancy.
| Aspect | Thanos | Grafana Mimir | Cortex |
|---|---|---|---|
| Ingestion | Sidecar uploads Prometheus blocks to object storage; or Receive for remote write | Remote write into distributors and ingesters | Remote write into distributors and ingesters |
| Main components | Sidecar, Querier, Store Gateway, Compactor (with downsampling), Ruler, Receive | Distributor, ingester, querier, query-frontend, store-gateway, compactor, ruler, Alertmanager | Similar microservice layout to Mimir |
| Strength | Incremental adoption: keep existing Prometheus servers, add components | Horizontally scalable, multi-tenant, long-term storage for Prometheus and OpenTelemetry metrics | The original CNCF horizontally scalable Prometheus backend |
| Note | CNCF project | Started as a fork of Cortex by Grafana Labs | CNCF project, still maintained |
How to choose: Thanos suits teams that already run many Prometheus servers and want to add a global view and retention with minimal change. Mimir (or Cortex) suits a central platform team running a shared, multi-tenant metrics service that clusters remote-write into.
37. Federation vs remote write vs a Thanos sidecar: when do you use each?
Answer: Federation lets a higher-level Prometheus scrape selected series (ideally pre-aggregated recording rules) from lower-level ones via /federate. It is fine for pulling a few aggregates into a global view; it is not a way to copy all data, which overloads both sides. Remote write streams everything to a central backend in near real time; this is the default modern choice, but it puts network and backend capacity on the critical path and requires tuning queue settings. A Thanos sidecar keeps data local and uploads completed two-hour blocks to object storage, while the Querier fans out to sidecars for recent data; this means fewer moving parts per cluster but recent data is only as available as each cluster's Prometheus.
cluster A: Prometheus --remote write--+
cluster B: Prometheus (agent) --------+--> Mimir --> S3/GCS
cluster C: OTel / Alloy --------------+ |
Grafana
38. How do you monitor Kubernetes with Prometheus? kube-state-metrics vs node-exporter vs cAdvisor?
Answer: You need three different views, and interviewers want you to keep them distinct:
- node-exporter (a DaemonSet): host-level hardware and OS metrics such as
node_cpu_seconds_total,node_memory_MemAvailable_bytes,node_filesystem_avail_bytes. - cAdvisor, embedded in the kubelet: per-container resource usage such as
container_cpu_usage_seconds_totalandcontainer_memory_working_set_bytes(the value the OOM killer effectively compares with the limit). - kube-state-metrics: the state of Kubernetes objects read from the API server, not resource usage. Examples:
kube_deployment_status_replicas_available,kube_pod_status_phase,kube_pod_container_status_restarts_total,kube_pod_container_resource_requests,kube_node_status_condition.
Add control-plane metrics (API server, etcd, scheduler, CoreDNS) where your platform exposes them; managed services often hide some. Useful queries combine views: CPU usage divided by requests per namespace to find over-provisioned workloads, or increase(kube_pod_container_status_restarts_total[1h]) > 3 for crash loops. For deeper cluster topics, the Kubernetes interview questions guide is the companion to this one.
39. What does the Prometheus Operator add, and what is kube-prometheus-stack?
Answer: The Prometheus Operator manages Prometheus and Alertmanager on Kubernetes through custom resources, so teams declare monitoring alongside their apps instead of editing one central scrape config. Key CRDs: Prometheus and Alertmanager (the servers), ServiceMonitor (scrape the endpoints behind a Service), PodMonitor (scrape pods directly), Probe (blackbox probes), ScrapeConfig, PrometheusRule (recording and alerting rules) and AlertmanagerConfig (namespaced routing). kube-prometheus-stack is the Helm chart that bundles the operator, Prometheus, Alertmanager, Grafana, node-exporter, kube-state-metrics, default dashboards and alert rules.
Interview tip: The commonest "my ServiceMonitor does nothing" cause is a label selector mismatch: the Prometheus resource's serviceMonitorSelector (often release: <name>) does not match the ServiceMonitor's labels, or the ServiceMonitor's port name does not match the Service's port name.
40. What are SLIs, SLOs and error budgets, and how do you express them in Prometheus?
Answer: An SLI is a measured ratio of good events to valid events, such as successful requests over all requests, or requests faster than 300ms over all requests. An SLO is the target for that ratio over a window, for example 0.999 of requests succeed over 30 days. The error budget is what the SLO allows to fail: with a 0.999 target, 0.001 of requests in the window. The budget turns reliability into a decision tool: while budget remains, ship features; when it is exhausted, prioritise reliability work.
- record: service:sli_errors:ratio_rate5m
expr: |
sum by (service) (
rate(http_requests_total{code=~"5.."}[5m]))
/
sum by (service) (
rate(http_requests_total[5m]))
Record the same ratio at several windows (5m, 30m, 1h, 6h and longer) for burn-rate alerts. Choose SLIs measured as close to the user as possible: load balancer or ingress metrics usually beat pod metrics because they include failures the pod never saw.
41. What is a burn rate, and how do multi-window, multi-burn-rate alerts work?
Answer: Burn rate is how fast you consume the error budget relative to the SLO: a burn rate of 1 uses exactly the whole budget by the end of the window; 14.4 uses it 14.4 times faster. Alerting on a static error ratio is either too noisy or too slow; burn-rate alerting pages for fast, severe burns and tickets for slow ones. The approach popularised by the Google SRE Workbook, for a 30-day SLO, pages if the 1-hour and 5-minute burn rates both exceed 14.4, and if the 6-hour and 30-minute rates both exceed 6; and opens a ticket for a burn rate of 1 over 3 days with a 6-hour short window.
- alert: CheckoutErrorBudgetFastBurn
expr: |
service:sli_errors:ratio_rate1h{service="checkout"}
> (14.4 * 0.001)
and
service:sli_errors:ratio_rate5m{service="checkout"}
> (14.4 * 0.001)
labels:
severity: page
The long window ensures enough budget is really being lost; the short window makes the alert resolve quickly once the problem is fixed, instead of firing for another hour.
42. How do you size and protect a Prometheus server?
Answer: Memory is driven mainly by active series in the head block, and disk by ingestion rate in samples per second multiplied by retention. Watch Prometheus's own metrics: prometheus_tsdb_head_series, rate(prometheus_tsdb_head_samples_appended_total[5m]), rule evaluation duration and missed iterations, remote-write queue lag. Then build guardrails into scrape configs so one bad target cannot sink the server: sample_limit (fail the scrape if a target exposes too many samples), label_limit, label_value_length_limit and target_limit. Shard by function (one Prometheus per team or per cluster) rather than growing one server forever, and keep heavy dashboards on recording rules. Alert on Prometheus itself, and send its alerts through a path that does not depend on it, such as a "dead man's switch" watchdog alert that must always fire.
AI in observability
43. Where does AI-based anomaly detection help in observability, and where does it hurt?
Answer: Anomaly detection helps when "normal" is hard to express as a static threshold: seasonal traffic with daily and weekly patterns, or hundreds of services where nobody will hand-tune thresholds. Simple statistical baselines go a long way inside PromQL: compare against the same time last week with offset 1w, or flag values several standard deviations from a rolling mean using avg_over_time and stddev_over_time recording rules. Commercial and cloud platforms add ML-based forecasting and outlier detection, and AIOps tools correlate alerts, deployments and logs to suggest a probable cause.
Where it hurts: an anomaly is not an incident. Unusual-but-harmless changes (a marketing campaign, a batch job) generate noise, and models that page directly erode trust quickly. Use anomaly signals as context on dashboards or as low-severity tickets, keep SLO burn-rate alerts as the paging path, and require every AI-suggested root cause to be confirmed by a human. The AIOps interview guide and AI for SRE engineers go deeper; the statistical side overlaps with anomaly detection interview questions.
44. How would you monitor an LLM-powered application with Prometheus and Grafana?
Answer: Keep RED metrics for the API, then add signals specific to LLM workloads: request count and errors per model and provider (including rate-limit and timeout errors), end-to-end latency and time to first token as histograms, input and output tokens as counters (which also gives you cost per service per day), cache hit ratio, retrieval latency and the number of documents retrieved for RAG, tool-call failures for agents, and guardrail blocks. For self-hosted models, add GPU utilisation and memory plus inference-server queue depth.
The cardinality trap is acute here: prompts, user IDs and conversation IDs must never be labels. They belong in traces and logs (with privacy controls), linked to metrics through exemplars. OpenTelemetry has semantic conventions for generative AI, still evolving, that many LLM tracing tools use. Quality (groundedness, answer relevance) is not a Prometheus metric at request time; it comes from evaluation pipelines that can publish aggregate scores. The full picture is in AI observability for LLM applications, and serving-side metrics are covered in the LLM inference and serving interview questions.
If you want to practise these skills on real systems, the SRE training in Hyderabad at Cloudsoft covers Prometheus, Grafana, alerting and SLO design with hands-on labs, in Ameerpet classrooms or live online.
Real-world scenario questions
45. Scenario: Prometheus memory doubles overnight and it starts getting OOMKilled. You suspect a cardinality explosion. What do you do?
Answer: Stabilise first, then find the source, then prevent recurrence. Give Prometheus enough memory to start and replay its WAL (or temporarily drop the offending job), because a crash-looping server gives you no data at all.
What I would check:
prometheus_tsdb_head_seriesover time: when did growth start, and does it line up with a deployment?- The TSDB status page (or
promtool tsdb analyze) for the metric names and label names with the most series and label-value pairs. topk(10, scrape_series_added)orscrape_samples_scrapedby job to identify which target is producing new series.- The offending metric's labels: a new
user_id,pathwith raw IDs, apod-like label that changes on every restart, or an error message used as a label. - Whether the change came from application code, an exporter upgrade or a relabeling change.
Production consideration: Short term, drop the label or metric with metric_relabel_configs (action: labeldrop or drop). Long term, fix the instrumentation (normalise /orders/12345 to the route template /orders/{id}), add sample_limit per job so one service fails its own scrape instead of taking the server down, and add an alert on head-series growth. Review new metrics in code review just like database schema changes.
46. Scenario: the on-call team receives hundreds of alerts a week and has started ignoring them. How do you fix alert fatigue?
Answer: Treat it as a product problem: every page must be urgent, actionable and user-impacting. Start with data, not opinions.
What I would check:
- Export a few weeks of alert history (Alertmanager notifications, PagerDuty or Opsgenie reports) and rank alerts by frequency, and by how often anyone took action.
- Alerts that fire and resolve within minutes with no action: add
for, lengthen windows or delete them. - Cause-based alerts (high CPU, a single pod restart, disk at a fixed level) that paged without user impact: downgrade to tickets or dashboards.
- Grouping: is
group_byso fine that one incident produces fifty notifications? Are inhibition rules missing for "cluster down" situations? - Ownership: alerts without a
teamlabel or runbook usually end up as everyone's noise.
Production consideration: Replace threshold pages with SLO burn-rate alerts for each critical user journey, keep a small set of "the system is definitely broken" pages (watchdog, absent data for critical services) and review the alert list in every incident retrospective.
47. Scenario: after a deploy, a service's metrics disappear from Grafana. What do you check?
Answer: Work along the pipeline: is the target discovered, is it scraped, are the metrics exposed under the names the dashboard expects?
What I would check:
- Prometheus's Targets page (or Service Discovery page): is the target listed, dropped by relabeling, or shown as down with an error?
- If missing in Kubernetes: did the deploy change pod or Service labels, port names or annotations that the ServiceMonitor or relabel rules select on?
- If down: curl the
/metricsendpoint from inside the cluster. Did the port or path change, did TLS or auth get added, is the content type now invalid (Prometheus 3.x is strict), or did a network policy start blocking the scrape? - If up but the panel is empty: did a library upgrade rename metrics (for example adding a
_totalsuffix, changing units, or OpenTelemetry-style dotted names), or change a label the dashboard filters on? - Was
sample_limitexceeded, making the whole scrape fail?
Production consideration: Add an absent() or up == 0 alert for every critical job so missing metrics page someone instead of being discovered on a dashboard days later. Test metric names in CI (a smoke test that scrapes the new build and checks expected series exist) and record deploy annotations in Grafana.
48. Scenario: design SLO alerts for a bank's payment API.
Answer: Consider a bank's payment initiation API used by mobile and net banking. I would start from the user journey, not the infrastructure: "a customer submits a payment and gets a definitive answer quickly".
What I would check:
- Define the SLIs with the business: availability (non-5xx responses, excluding client errors such as invalid account numbers) and latency (share of requests under an agreed threshold), measured at the API gateway or ingress.
- Agree targets and windows with product and risk teams, for example 0.999 availability over 30 days, and confirm what counts as a valid event (health checks and synthetic traffic excluded).
- Build recording rules for the error ratio at 5m, 30m, 1h, 6h and 3d windows, and make sure histogram buckets include the latency threshold exactly.
- Create multi-window burn-rate alerts: fast burns page the payments on-call; slow burns raise a ticket. Add an
absent()alert in case the SLI itself stops reporting. - Add a business-level guard: a sudden drop in successful payments compared with the same time last week, because a downstream failure can return 200s that look healthy.
Production consideration: Publish an error-budget dashboard that product owners actually read, tie release decisions to remaining budget, and agree in advance what happens when budget runs out. Low-traffic periods (late night) make ratios noisy; consider minimum request-count conditions or synthetic probes so a single failure at 3 a.m. does not page anyone.
49. Scenario: dashboards are slow and queries time out during incidents. How do you fix it?
Answer: Incidents are exactly when everyone opens the same heavy dashboards at once, so query cost has to be designed in.
What I would check:
- Which queries are slow: Prometheus query logs, the query statistics in Grafana's query inspector, or the backend's slow-query logging.
- Queries over long ranges with raw, high-cardinality selectors, regex matchers like
{pod=~".*"}, or subqueries. - Dashboards with many panels, auto-refresh every few seconds, and template variables that load thousands of values.
Production consideration: Move common aggregations into recording rules, put a query-frontend (Thanos or Mimir) in front for caching and splitting long range queries, set sensible minimum intervals on panels, and split one giant dashboard into an overview plus drill-downs. Set query limits (timeouts, maximum samples) so one bad query cannot starve everyone else.
50. Scenario: your p99 latency panel shows a value that the application team says is impossible. Why might it be wrong?
Answer: Several classic PromQL mistakes produce plausible-looking but wrong percentiles.
What I would check:
- Is the panel averaging summary quantiles across pods (
avg(..{quantile="0.99"}))? That is not a fleet p99. - Was
ledropped orrate()skipped beforehistogram_quantile()? - Bucket layout: if most requests fall into one wide bucket, interpolation produces a number that never happened. If the true p99 exceeds the highest finite bucket, the function returns that bucket's upper bound.
- Low traffic: a 5-minute window with a handful of requests gives a meaningless p99.
- Units: milliseconds recorded into a
_secondshistogram.
Production consideration: Re-design the buckets around the SLO threshold or move to native histograms, show request count next to every percentile, and for SLOs prefer "fraction of requests under X" over a percentile, which is exact and aggregates correctly.
51. Scenario: alerts fire in Prometheus but nobody receives them on PagerDuty. Debug it.
Answer: Follow the alert hop by hop: Prometheus to Alertmanager, Alertmanager routing, Alertmanager to receiver.
What I would check:
- Prometheus Alerts page shows it firing; then
prometheus_notifications_errors_totalandprometheus_notifications_dropped_total, and the configured Alertmanager endpoints on the Status page. Is Prometheus sending at all? - Alertmanager UI: does the alert appear? Is it silenced or inhibited?
amtool config routes testwith the alert's labels: which receiver does it actually reach? Aseverity="critical"alert against a route matchingseverity="page"is a frequent culprit.alertmanager_notifications_failed_totalby integration, and Alertmanager logs for HTTP errors from PagerDuty (wrong routing key, proxy blocking egress, expired secret).- Timing: a large
group_waitor a longrepeat_intervalcan make it look like nothing was sent.
Production consideration: Run a watchdog alert that always fires and is routed to a dead man's switch service, so a broken pipeline itself raises an alarm. Test routing changes in CI with amtool.
52. Scenario: a request-rate graph shows huge spikes every time pods restart. What is wrong?
Answer: Almost certainly aggregation before rate, or the wrong function on a counter.
What I would check:
- Is the query
rate(sum(...)[5m:]),deriv()on a counter, or a raw counter with a Grafana "difference" transformation? All mishandle resets, because summing hides individual resets. - Is the counter genuinely a counter, or does the app expose a gauge-like value with a
_totalname that can decrease? - Do new pods start with a non-zero value (counters restored from somewhere), which rate interprets as a burst?
Production consideration: Fix the query to sum(rate(x[5m])), write it as a recording rule so dashboards stop re-inventing it, and add a PromQL linter to the dashboard and rules pipeline.
53. Scenario: a team migrates from Prometheus client libraries to OpenTelemetry SDKs and OTLP. What can break?
Answer: Mostly names, labels and the "is it up" signal.
What I would check:
- Metric names: OpenTelemetry names like
http.server.request.durationmay arrive with dots preserved (UTF-8) or translated with underscores and unit suffixes, depending on thetranslation_strategy. Every dashboard and rule that used the old names must be mapped. - Units and types: OTel histograms in seconds vs old milliseconds; exponential histograms becoming native histograms, which older dashboards and remote storage may not handle.
- Labels:
service.nameand Kubernetes attributes are resource attributes; unless promoted, they sit intarget_info, soby (service)queries return nothing. - Liveness: OTLP push has no
upmetric; a silently dead service just stops sending. - Out-of-order and duplicate samples if several Collectors send the same data.
Production consideration: Run old and new instrumentation in parallel for a period, compare key queries side by side, promote a small agreed set of resource attributes, and add absent_over_time() alerts for critical services. Plan the Grafana Agent to Alloy migration at the same time if Agent is still in use.
54. Scenario: a hospital group's IT team in a Hyderabad GCC needs monitoring for 20 Kubernetes clusters with one-year retention. Design it.
Answer: Consider a hospital group whose platform team runs clinical and administrative apps across several clusters, with audit requirements to retain operational metrics for a year. A single Prometheus per cluster plus a central long-term store fits.
each cluster:
node-exporter + kube-state-metrics + apps
|
Prometheus (local rules, short retention)
| remote write
v
central: Mimir or Thanos --> object storage
| (1 year)
Grafana (SSO, folders per team)
Alertmanager (HA cluster) --> on-call tools
What I would check:
- Expected active series and ingestion rate per cluster, measured from a pilot rather than guessed, to size the backend.
- Where alerts evaluate: keep critical alerts in each cluster's Prometheus so they still fire if the central system or network is down.
- Tenancy and access: per-team tenants or label-based access, SSO for Grafana, and no patient data in labels at all.
- Retention and downsampling for the year-long data, and the object storage lifecycle and encryption policy.
- Dashboards and rules as code across clusters, with a
clusterexternal label everywhere.
Production consideration: Monitor the monitoring (remote-write lag, ingester health), run a watchdog per cluster, and redact logs and traces before shipping.
55. Scenario: an interviewer gives you a vague alert, "CPU above 80 for 5 minutes", and asks you to improve it. What do you say?
Answer: I would ask what user impact the alert is meant to catch, because high CPU on its own is a cause, not a symptom: a batch node at high CPU may be working perfectly.
What I would check:
- Is there a user-facing SLI for this service? If so, page on its burn rate and keep CPU as a dashboard signal.
- In Kubernetes, is CPU throttling the real issue?
rate(container_cpu_cfs_throttled_periods_total[5m]) / rate(container_cpu_cfs_periods_total[5m])often explains latency better than utilisation. - Is the concern capacity? Then a
predict_linear()or saturation trend alert that opens a ticket days ahead is more useful than a page. - Does the expression aggregate correctly (per node or per pod, not averaged across a fleet that hides one hot node)?
Production consideration: End with the improved set: a burn-rate page for the user journey, a throttling ticket alert with a runbook, and a capacity forecast reviewed weekly. Interviewers are testing whether you design alerts around people and outcomes, not around whatever metric is easiest to threshold.
Key takeaways
- Labels are dimensions; every unique combination is a series. Cardinality discipline is the most important operational skill with Prometheus.
- Rate counters before aggregating, keep
leforhistogram_quantile(), and never average quantiles. - Prometheus decides whether something is wrong; Alertmanager decides who hears about it. Grouping, inhibition and routing tests are what prevent alert fatigue.
- Prometheus 3.x brought UTF-8 names, OTLP ingestion and Remote Write 2.0 (still experimental), and native histograms are now stable; know the upgrade gotchas.
- Grafana Alloy replaced Grafana Agent, which reached end of life on 1 November 2025.
- Long-term storage and global views come from Thanos, Mimir or Cortex; HA comes from duplicate Prometheus servers plus downstream deduplication.
- Page on SLO burn rates for user journeys; send causes to tickets and dashboards.
Interview preparation checklist
- Run Prometheus, node-exporter, Alertmanager and Grafana locally (Docker Compose or kind with kube-prometheus-stack) and scrape a small app you instrumented yourself.
- Write from memory: an error-ratio query, a p99 from a histogram, a
group_leftjoin withkube_pod_info, and a disk-fillpredict_linear()alert. - Create recording rules and test them with
promtool test rules. - Build an Alertmanager config with grouping, one inhibition rule and two receivers; test it with
amtool config routes test. - Define one SLO and implement multi-window burn-rate alerts for it.
- Break things on purpose: add a high-cardinality label and watch head series grow; rename a metric and watch the dashboard go blank; then fix both.
- Store one dashboard as code (provisioned JSON, Jsonnet, Foundation SDK or Terraform) and change it through a pull request.
- Read the Prometheus 3 migration guide and the Alloy documentation, and be ready to explain what changed.
- Prepare two incident stories in the format: symptom, what the metrics showed, root cause, the fix, and the alert or dashboard you changed afterwards.
FAQ
What skills are needed for a Prometheus and Grafana role?
Solid Linux and networking basics, Kubernetes fundamentals, PromQL, alert design with Alertmanager, Grafana dashboarding, and an understanding of SLOs. Scripting in Python or Go and infrastructure as code help, because monitoring configuration should live in Git.
Is Prometheus still relevant now that OpenTelemetry exists?
Yes. OpenTelemetry standardises how telemetry is produced and transported, while Prometheus stores, queries and alerts on metrics. Prometheus 3.x accepts OTLP directly, and PromQL remains the common query language across many metrics backends.
How should I prepare for PromQL interview questions?
Practise on a live Prometheus with real data rather than memorising syntax. Focus on rate versus irate, histogram quantiles, aggregation with by and without, vector matching, and explaining why a query is correct.
Do I need to learn Thanos or Mimir for interviews?
For mid-level roles, understand the problems they solve and their main components. For senior or platform roles, expect to compare them, explain how deduplication and object storage work, and justify a choice for a given organisation.
Is Grafana Agent still used?
Grafana Agent reached end of life on 1 November 2025 and no longer receives fixes. New deployments use Grafana Alloy, and existing Agent users are expected to migrate to it.
What is the difference between monitoring and observability?
Monitoring checks known failure conditions with predefined metrics and alerts. Observability is the ability to ask new questions about system behaviour from its telemetry, combining metrics, logs, traces and profiles to explain problems you did not predict.
Is SRE a good career path for DevOps engineers?
For engineers who enjoy reliability, incident response, automation and systems thinking, it is a natural next step. SRE roles weigh observability, SLOs and production debugging heavily, so Prometheus and Grafana skills transfer directly.
Can freshers get Prometheus and Grafana interview questions?
Yes, usually at the fundamentals level: pull versus push, metric types, exporters, a simple rate query and what Grafana does. A small home lab with an instrumented app is a strong talking point for a fresher.
How many Prometheus metric types should I know?
Know all four: counter, gauge, histogram and summary, and when to use each. Also understand native histograms, which are now a stable feature in Prometheus 3.x.
Prometheus and Grafana are learned fastest by running them against systems that break. Cloudsoft's Site Reliability Engineering course takes you through instrumentation, PromQL, Alertmanager, SLO burn-rate alerting and Kubernetes monitoring with hands-on labs, in our Ameerpet classroom beside the Metro or live online. If you also want AI, ML and cloud security in one structured track, look at the APEX AI, ML, Cloud and Cyber Security program. Book a free demo on +91 96660 19191.



