Dashboards & Alerts
Grafana Operator
observability/base/grafana-operator/ installs grafana-operator (5.25.0),
but it does not deploy its own Grafana. Its Grafana custom resource
(grafana-victoriametrics.yaml) is configured as external, pointing at
the Grafana subchart already bundled inside victoria-metrics-k8s-stack:
apiVersion: grafana.integreatly.org/v1beta1
kind: Grafana
metadata:
name: grafana-victoriametrics
labels:
dashboards: "grafana"
spec:
external:
url: http://victoria-metrics-k8s-stack-grafana
adminPassword:
name: victoria-metrics-k8s-stack-grafana-envvars
key: GF_SECURITY_ADMIN_PASSWORDgrafana-operator’s actual job here is reconciling GrafanaFolder,
GrafanaDashboard, and GrafanaDatasource CRs against that one external
instance. Every such CR in this repo carries the same two fields:
instanceSelector.matchLabels.dashboards: "grafana" (matches the Grafana
CR’s own label above) and allowCrossNamespaceImport: true (folders and
dashboards are defined across apps, infrastructure, and observability
namespaces, all targeting the one Grafana in observability).
Folder registry
A repo-wide grep for kind: GrafanaFolder finds nine CRs. Eight reconcile onto
the cluster by default, each defined in the namespace that owns its
dashboards; the ninth (llm, apps/base/ai/llm/grafana-folder.yaml) only
applies once the opt-in LLM platform’s suspended umbrella Kustomization is
resumed (see clusters/AGENTS.md) and is out of
scope for this always-on page:
| Folder | Defined in | Holds |
|---|---|---|
apps | apps | Demo “all-in-one” RED + trace/log correlation dashboard |
cilium | kube-system | Cilium agent/operator and Hubble dashboards |
databases | infrastructure | CloudNativePG query-performance and query-plan-correlation |
flux | flux-system | Flux cluster and control-plane dashboards |
kubernetes | infrastructure | Kubernetes views, node-exporter, Karpenter |
logs | observability | VictoriaLogs explorer, single/cluster overview |
runlore | observability | RunLore’s own dashboard |
traces | observability | VictoriaTraces overview |
Dashboard inventory
| Dashboard | Folder | Source |
|---|---|---|
kubernetes-views-{global,namespaces,nodes,pods} | observability | Imported — vendored by the victoria-metrics-k8s-stack chart |
kubernetes-node-exporter-full | kubernetes | Authored in-repo |
kubernetes-karpenter | kubernetes | Imported, grafana.com dashboard 20398 |
app-all-in-one | apps | Authored in-repo |
runlore | runlore | Authored in-repo |
observability-victoria-logs-explorer | logs | Authored in-repo |
observability-victoria-logs-single | logs | Imported, dashboard 22084 |
observability-victoria-traces-single | traces | Imported, dashboard 24136 |
databases-cnpg-query-performance / -query-plan-correlation | databases | Authored in-repo — see PostgreSQL |
cilium-cilium, cilium-operator, cilium-hubble{,-dns-namespace,-l7-http-metrics,-network-overview-namespace} (6 dashboards) | cilium | Imported, from the cilium/cilium upstream dashboard JSON |
flux-cluster, flux-control-plane | flux | Imported, from fluxcd/flux2-monitoring-example — pinned to commit 7ab65dc8 since 2026-08-30; tracking refs/heads/main made them mutate whenever upstream moved |
victoria-logs and victoria-traces also self-ship a dashboard via their
own chart’s dashboards.enabled: true, grafanaOperator.enabled: true values,
targeting the same logs/traces folders as the manually-authored ones
above — two sourcing paths land in the same folder, which is redundant but
not a conflict (different dashboard names).
The victoria-metrics-k8s-stack chart’s own defaultDashboards ships into an
"observability" folder (not one of the GrafanaFolder CRs above — it’s
created implicitly by the chart’s sidecar mechanism, not grafana-operator).
Its Kubernetes-views duplicates of the kubernetes folder’s dashboards are
explicitly disabled in vm-common-helm-values-configmap.yaml to avoid
showing the same dashboard twice.
Grafana itself: version and datasource plugins
The Grafana all of this lands in is the victoria-metrics-k8s-stack
subchart, its image pinned to 13.1.4 in
vm-common-helm-values-configmap.yaml — the stack constrains the subchart
to grafana: 12.7.*, so security releases need that explicit tag until the
stack bumps its dependency. The two VictoriaMetrics datasource plugins are
pinned catalog installs in the chart’s plugins: list —
victoriametrics-metrics-datasource@0.26.1 and
victoriametrics-logs-datasource@0.32.0 — which the chart maps straight to
GF_PLUGINS_PREINSTALL_SYNC. Two details of that list are load-bearing: the
pin separator is @, not a space (a space is silently split into two bare
plugin ids, so the real plugin installs unpinned), and each pin carries a
Renovate annotation, so upgrades arrive as pull requests. This mechanism
replaced two curl initContainers on 2026-08-29: one frozen at v0.14.0
for 16 months, the other fetching GitHub “latest” on every pod start.
VictoriaTraces
observability/base/victoria-traces/ runs victoria-traces-single (0.1.11),
3-day retention, wired into Grafana as a Jaeger-protocol datasource. The 3d
suffix in retentionPeriod: 3d matters: a unit-less retentionPeriod: 3
means 3 months in the VictoriaMetrics chart family, and that is what this
release silently kept — ×30 the intent — until the suffix was added on
2026-08-30. The datasource:
# observability/base/victoria-traces/grafana-datasource.yaml
datasource:
uid: VictoriaTraces # pinned, and equal to the name; see below
type: jaeger
url: http://victoria-traces-vt-single-server.observability:10428/select/jaeger
jsonData:
tracesToLogsV2:
datasourceUid: VictoriaLogs
tracesToMetrics:
datasourceUid: VictoriaMetrics
tags:
- key: service.name
nodeGraph:
enabled: truetracesToLogsV2/tracesToMetrics are what let a Grafana user pivot from a
trace span to the matching log lines (via log.trace_id, per the
LogsQL rules)
or request-rate/latency panels, without re-typing a query by hand.
Three details here are load-bearing, and all three were wrong until 2026-09-13 — every cross-link silently did nothing from the day it was written:
datasourceUid, neverdatasourceName. No Grafana schema has ever accepted the latter. There is no error: the datasource loads and the link simply never appears.tracesToLogsV2, nottracesToLogs. v1 is superseded and current Grafana reads v2, so correcting keys inside a v1 block changes nothing.- A pinned
uid, deliberately equal to thename. Without a pin the operator assigns a random uid, so a cross-reference written as the literal stringVictoriaLogsmatched no datasource. Since 2026-09-13 the uid isVictoriaLogs, following the convention the vmks chart already sets withVictoriaMetrics/VictoriaMetricsandAlertmanager/Alertmanager: adatasourceUidfield takes a uid and rejects a name outright, so a lowercase uid forces two spellings for one datasource. Note it is set atspec.datasource.uid, the field the CRD marks deprecated, and deliberately:spec.uidis immutable and the API server rejects it outright on an already-created object, which would fail every Flux reconcile.
Only part of this is caught by CI. flux schema validate sees a well-formed
map whichever key you use, so the two KEY choices above — datasourceUid
over datasourceName, tracesToLogsV2 over tracesToLogs — are gated by
nothing but review. The uid VALUES are pinned: the .doc-claims.yaml claims
victoriatraces-datasource-uid and victorialogs-datasource-uid read them
from the two manifests and fail this page when it drifts from them, which it
had already done once. The
manifest doesn’t configure an explicit receiver protocol (no OTLP toggles) —
VictoriaTraces’ chart defaults apply unmodified, so which ingest protocols
are actually enabled isn’t determinable from this repo alone.
Alerting
Two VMRule flavors coexist under the same operator.victoriametrics.com/v1beta1
API: implicit PromQL rules (the default) and explicit type: vlogs rules
evaluated as LogsQL — see loggen’s alert
for the LogsQL form. PromQL rules live mostly in
victoria-metrics-k8s-stack/vmrules/ — in the base for both clusters unless
noted:
karpenter.yaml— 3 alerts (node-registration failures, nodepool near capacity, cloud-provider errors) for a component this stack doesn’t own but scrapes. Lives in theaws-0overlay (observability/aws-0/victoria-metrics-k8s-stack/vmrules/), because thekarpenternamespace does not exist ongcp-0.openbao.yaml— 7 alerts (down, sealed, no active node, lost/at-risk Raft voter, snapshot job failed/stale), extensively commented with the incident history that motivated each one (an expired cert going unnoticed for 9 months; a stale Raft voter silently halving failure tolerance).runlore.yaml— 12 alerts covering the agent’s own health (down, no leader, split-brain), pipeline behavior (dropped/stalled/erroring investigations), and cost (token spend, model latency). Each carries arunbook_urlpointing at RunLore’s own docs. Moved from theaws-0overlay to the base on 2026-08-30, so the alerts follow the agent to both clusters. What the agent itself does with an alert is on its own page — SRE agent.
# observability/base/victoria-metrics-k8s-stack/vmrules/runlore.yaml — trimmed
- alert: RunloreAgentDown
expr: absent(runlore_build_info)
for: 5m
labels:
severity: critical
annotations:
runbook_url: "https://github.com/Smana/runlore/blob/main/docs/observability.md#runloreagentdown"Alertmanager routing
Alertmanager’s routing tree (vm-common-helm-values-configmap.yaml, applied to
the victoria-metrics-k8s-stack alertmanager.spec.config) fans every
non-blackholed alert to both RunLore and Slack, then splits by severity so
one channel can carry a page and a nag without them looking alike:
route:
receiver: "slack-monitoring"
repeat_interval: 12h
routes:
- matchers:
- alertname =~ "InfoInhibitor|Watchdog|KubeCPUOvercommit"
receiver: "blackhole"
- receiver: "runlore"
continue: true # falls through to the next route instead of stopping
- matchers: [severity = "critical"]
receiver: "slack-monitoring"
group_wait: 10s
repeat_interval: 1h
- matchers: [severity = "warning"]
receiver: "slack-monitoring"
repeat_interval: 12h
- receiver: "slack-monitoring" # no matchers — see below
repeat_interval: 24hThe final route’s absence of matchers is load-bearing. severity = "critical"
does not match an alert that has no severity label — a missing label is the
empty string. That last route is the catch-all keeping an unlabelled alert
reachable.
The dangerous edit is the well-intentioned one. Giving it an explicit matcher
like severity = "info" reads more precise and would pass review, and an alert
arriving without a severity label would then match no route and reach nobody.
No gate would fail: amtool accepts it, the schema gate sees a valid string, and
the golden files do not move, because routing changes what is delivered and never
how it is drawn. The only symptom is an alert that quietly stops arriving.
Two inhibit rules, both deliberately narrow:
inhibit_rules:
- source_matchers: [severity = "critical"]
target_matchers: [severity = "warning"]
equal: [cluster, alertname, namespace]
- source_matchers: [alertname = "OpenBaoRaftQuorumAtRisk"]
target_matchers: [alertname = "OpenBaoRaftNodeLost"]
equal: [cluster]The first catches one alertname firing at two severities. The second exists
because OpenBaoRaftQuorumAtRisk and OpenBaoRaftNodeLost are different
alertnames where one implies the other — quorum-at-risk means a peer is already
lost — and they arrived a minute apart on 2026-09-12. Add named pairs as you
observe them rather than generalising: an inhibited alert never reaches Slack at
all, so an over-broad rule silently deletes alerts, which is worse than a
duplicate.
RunLore’s webhook requires a bearer token (its v0.2.0+ fail-closed behavior —
alert labels/annotations flow into an LLM prompt, so an unauthenticated trigger
path was judged unacceptable) mirrored between its own runlore-webhook
ExternalSecret and this stack’s runlore-webhook-token ExternalSecret.
The Slack message
Rendered by repo-owned templates shipped through alertmanager.templateFiles,
with the chart’s vendored Monzo set disabled — see
ADR-0037
for why these are Alertmanager-native rather than Block Kit behind a bridge.
:fire: FIRING — OpenBaoRaftQuorumAtRisk
`aws-0 · dev`
OpenBao raft cluster cannot tolerate a node failure.
• 10.0.12.44:8200 — Failure tolerance has been below 1 for 15 minutes: losing
one more node loses quorum, and a cluster without quorum cannot issue a
certificate or read a secret. If OpenBaoRaftNodeLo…
Namespace security Severity critical
Duration 14m (since 09:06 UTC) Location aws / eu-west-3
[ Runbook ] [ Dashboard ] [ Query ] [ Silence ]The identity line (aws-0 · dev) and the Location field come from labels
vmalert stamps in externalLabels — cluster, env, cloud, region — fed
by each cluster’s flux_cluster_vars ConfigMap. They are stamped after rule
evaluation, which is the point: most upstream rules aggregate the cluster
label away, and a message once reached Slack reading on cluster .
Optional values degrade rather than disappear: a missing namespace renders —,
an empty cloud drops that half of Location, and a group of more than five
alerts shows five bullets and … showing 5 of N.
The annotation contract
./scripts/ci/validate-vmrules.sh enforces the first line of this on every
repo-authored alert:
| Annotation | Required | Rendered as |
|---|---|---|
summary | yes — one line, ≤140 chars | the headline |
description | no, any length | per-alert bullet, truncated at 180 chars |
runbook_url | no | Runbook button; falls back to this page |
dashboard | no | Dashboard button and title_link; falls back to Grafana’s home |
description is deliberately left unbounded. Slack truncates it; RunLore
receives the same annotations over its webhook and reads all of it, so trimming
operational prose out of a rule to make it fit a chat message would blind the
agent to the one thing that explains the alert.
Changing the wording
Edit templateFiles.ogenki.tmpl, then:
./scripts/ci/validate-alertmanager-templates.sh --update-golden
git diff scripts/ci/tests/alertmanager-fixtures/golden/ # read it — it IS the message
./scripts/ci/validate-alertmanager-templates.shTwo traps bind that block. Never write a literal ${ — Flux post-build
substitution expands ${var} and replaces an unknown one with an empty string,
and this applies inside comments too; a bare $var (every Go template variable)
is safe. Alertmanager ships no sprig — no default, no arithmetic. A
template calling one fails to execute, and Alertmanager then drops the
notification, which is why the gate exists at all.
The gate checks every rendered Alertmanager config with amtool check-config,
renders every templated string against five fixture payloads, and asserts every
VMAlert carries an absolute external.url. It validates structure, not
semantics: a typo’d equal label such as clustre is a syntactically valid
label name and passes. So do a shadowing route and an over-broad inhibit rule.
Grafana OnCall (removed)
The former grafana-oncall directory under observability/base/ — a complete
engine + RabbitMQ + Postgres + Valkey install that no Flux Kustomization
ever referenced — was
removed on 2026-08-29, along with the grafana-oncall-app plugin and the
provisioning ConfigMap that pointed it at an oncall-engine Service that
never existed. Upstream OnCall OSS is archived (read-only on GitHub since
2026-06-05, zero future patches), and the RunLore + Slack routing above
already carries the whole incident flow.
ADR-0029
records the decision and the full inventory of what was deleted.
CiliumNetworkPolicy coverage
Of the eight component directories under observability/base/, only
runlore defines a CiliumNetworkPolicy:
# observability/base/runlore/ciliumnetworkpolicy-ingress.yaml
spec:
endpointSelector:
matchLabels:
app.kubernetes.io/name: runlore
ingress:
- fromEntities: [ingress]
toPorts:
- ports: [{port: "8080", protocol: TCP}]It exists because runlore’s HTTPRoute parents to platform-public
(described in Private Access),
and this cluster’s shared cilium-envoy DaemonSet runs hostNetwork: true —
Gateway-originated traffic lands as Cilium’s reserved ingress entity, which
a namespaced Kubernetes NetworkPolicy can never match. grafana-operator,
kubernetes-event-exporter, loggen, metrics-server,
victoria-logs, victoria-metrics-k8s-stack, and victoria-traces have no
CiliumNetworkPolicy in these directories, and no cluster-wide default-deny
CiliumClusterwideNetworkPolicy covers the observability namespace either
— the one such policy in this repo
(infrastructure/base/gapi/allow-gateway-l7-proxy.yaml) only allows the
Gateway API L7 proxy, it does not deny anything by default.