Skip to content

Dashboards & Alerts

Grafana Operator

observability/base/grafana-operator/ installs grafana-operator (5.25.0), but it does not deploy its own Grafana. Its Grafana custom resource (grafana-victoriametrics.yaml) is configured as external, pointing at the Grafana subchart already bundled inside victoria-metrics-k8s-stack:

apiVersion: grafana.integreatly.org/v1beta1
kind: Grafana
metadata:
  name: grafana-victoriametrics
  labels:
    dashboards: "grafana"
spec:
  external:
    url: http://victoria-metrics-k8s-stack-grafana
    adminPassword:
      name: victoria-metrics-k8s-stack-grafana-envvars
      key: GF_SECURITY_ADMIN_PASSWORD

grafana-operator’s actual job here is reconciling GrafanaFolder, GrafanaDashboard, and GrafanaDatasource CRs against that one external instance. Every such CR in this repo carries the same two fields: instanceSelector.matchLabels.dashboards: "grafana" (matches the Grafana CR’s own label above) and allowCrossNamespaceImport: true (folders and dashboards are defined across apps, infrastructure, and observability namespaces, all targeting the one Grafana in observability).

Folder registry

A repo-wide grep for kind: GrafanaFolder finds nine CRs. Eight reconcile onto the cluster by default, each defined in the namespace that owns its dashboards; the ninth (llm, apps/base/ai/llm/grafana-folder.yaml) only applies once the opt-in LLM platform’s suspended umbrella Kustomization is resumed (see clusters/AGENTS.md) and is out of scope for this always-on page:

FolderDefined inHolds
appsappsDemo “all-in-one” RED + trace/log correlation dashboard
ciliumkube-systemCilium agent/operator and Hubble dashboards
databasesinfrastructureCloudNativePG query-performance and query-plan-correlation
fluxflux-systemFlux cluster and control-plane dashboards
kubernetesinfrastructureKubernetes views, node-exporter, Karpenter
logsobservabilityVictoriaLogs explorer, single/cluster overview
runloreobservabilityRunLore’s own dashboard
tracesobservabilityVictoriaTraces overview

Dashboard inventory

DashboardFolderSource
kubernetes-views-{global,namespaces,nodes,pods}observabilityImported — vendored by the victoria-metrics-k8s-stack chart
kubernetes-node-exporter-fullkubernetesAuthored in-repo
kubernetes-karpenterkubernetesImported, grafana.com dashboard 20398
app-all-in-oneappsAuthored in-repo
runlorerunloreAuthored in-repo
observability-victoria-logs-explorerlogsAuthored in-repo
observability-victoria-logs-singlelogsImported, dashboard 22084
observability-victoria-traces-singletracesImported, dashboard 24136
databases-cnpg-query-performance / -query-plan-correlationdatabasesAuthored in-repo — see PostgreSQL
cilium-cilium, cilium-operator, cilium-hubble{,-dns-namespace,-l7-http-metrics,-network-overview-namespace} (6 dashboards)ciliumImported, from the cilium/cilium upstream dashboard JSON
flux-cluster, flux-control-planefluxImported, from fluxcd/flux2-monitoring-example — pinned to commit 7ab65dc8 since 2026-08-30; tracking refs/heads/main made them mutate whenever upstream moved

victoria-logs and victoria-traces also self-ship a dashboard via their own chart’s dashboards.enabled: true, grafanaOperator.enabled: true values, targeting the same logs/traces folders as the manually-authored ones above — two sourcing paths land in the same folder, which is redundant but not a conflict (different dashboard names).

The victoria-metrics-k8s-stack chart’s own defaultDashboards ships into an "observability" folder (not one of the GrafanaFolder CRs above — it’s created implicitly by the chart’s sidecar mechanism, not grafana-operator). Its Kubernetes-views duplicates of the kubernetes folder’s dashboards are explicitly disabled in vm-common-helm-values-configmap.yaml to avoid showing the same dashboard twice.

Grafana itself: version and datasource plugins

The Grafana all of this lands in is the victoria-metrics-k8s-stack subchart, its image pinned to 13.1.4 in vm-common-helm-values-configmap.yaml — the stack constrains the subchart to grafana: 12.7.*, so security releases need that explicit tag until the stack bumps its dependency. The two VictoriaMetrics datasource plugins are pinned catalog installs in the chart’s plugins: list — victoriametrics-metrics-datasource@0.26.1 and victoriametrics-logs-datasource@0.32.0 — which the chart maps straight to GF_PLUGINS_PREINSTALL_SYNC. Two details of that list are load-bearing: the pin separator is @, not a space (a space is silently split into two bare plugin ids, so the real plugin installs unpinned), and each pin carries a Renovate annotation, so upgrades arrive as pull requests. This mechanism replaced two curl initContainers on 2026-08-29: one frozen at v0.14.0 for 16 months, the other fetching GitHub “latest” on every pod start.

VictoriaTraces

observability/base/victoria-traces/ runs victoria-traces-single (0.1.11), 3-day retention, wired into Grafana as a Jaeger-protocol datasource. The 3d suffix in retentionPeriod: 3d matters: a unit-less retentionPeriod: 3 means 3 months in the VictoriaMetrics chart family, and that is what this release silently kept — ×30 the intent — until the suffix was added on 2026-08-30. The datasource:

# observability/base/victoria-traces/grafana-datasource.yaml
datasource:
  uid: VictoriaTraces            # pinned, and equal to the name; see below
  type: jaeger
  url: http://victoria-traces-vt-single-server.observability:10428/select/jaeger
  jsonData:
    tracesToLogsV2:
      datasourceUid: VictoriaLogs
    tracesToMetrics:
      datasourceUid: VictoriaMetrics
      tags:
        - key: service.name
    nodeGraph:
      enabled: true

tracesToLogsV2/tracesToMetrics are what let a Grafana user pivot from a trace span to the matching log lines (via log.trace_id, per the LogsQL rules) or request-rate/latency panels, without re-typing a query by hand.

Three details here are load-bearing, and all three were wrong until 2026-09-13 — every cross-link silently did nothing from the day it was written:

  • datasourceUid, never datasourceName. No Grafana schema has ever accepted the latter. There is no error: the datasource loads and the link simply never appears.
  • tracesToLogsV2, not tracesToLogs. v1 is superseded and current Grafana reads v2, so correcting keys inside a v1 block changes nothing.
  • A pinned uid, deliberately equal to the name. Without a pin the operator assigns a random uid, so a cross-reference written as the literal string VictoriaLogs matched no datasource. Since 2026-09-13 the uid is VictoriaLogs, following the convention the vmks chart already sets with VictoriaMetrics/VictoriaMetrics and Alertmanager/Alertmanager: a datasourceUid field takes a uid and rejects a name outright, so a lowercase uid forces two spellings for one datasource. Note it is set at spec.datasource.uid, the field the CRD marks deprecated, and deliberately: spec.uid is immutable and the API server rejects it outright on an already-created object, which would fail every Flux reconcile.

Only part of this is caught by CI. flux schema validate sees a well-formed map whichever key you use, so the two KEY choices above — datasourceUid over datasourceName, tracesToLogsV2 over tracesToLogs — are gated by nothing but review. The uid VALUES are pinned: the .doc-claims.yaml claims victoriatraces-datasource-uid and victorialogs-datasource-uid read them from the two manifests and fail this page when it drifts from them, which it had already done once. The manifest doesn’t configure an explicit receiver protocol (no OTLP toggles) — VictoriaTraces’ chart defaults apply unmodified, so which ingest protocols are actually enabled isn’t determinable from this repo alone.

Alerting

Two VMRule flavors coexist under the same operator.victoriametrics.com/v1beta1 API: implicit PromQL rules (the default) and explicit type: vlogs rules evaluated as LogsQL — see loggen’s alert for the LogsQL form. PromQL rules live mostly in victoria-metrics-k8s-stack/vmrules/ — in the base for both clusters unless noted:

  • karpenter.yaml — 3 alerts (node-registration failures, nodepool near capacity, cloud-provider errors) for a component this stack doesn’t own but scrapes. Lives in the aws-0 overlay (observability/aws-0/victoria-metrics-k8s-stack/vmrules/), because the karpenter namespace does not exist on gcp-0.
  • openbao.yaml — 7 alerts (down, sealed, no active node, lost/at-risk Raft voter, snapshot job failed/stale), extensively commented with the incident history that motivated each one (an expired cert going unnoticed for 9 months; a stale Raft voter silently halving failure tolerance).
  • runlore.yaml — 12 alerts covering the agent’s own health (down, no leader, split-brain), pipeline behavior (dropped/stalled/erroring investigations), and cost (token spend, model latency). Each carries a runbook_url pointing at RunLore’s own docs. Moved from the aws-0 overlay to the base on 2026-08-30, so the alerts follow the agent to both clusters. What the agent itself does with an alert is on its own page — SRE agent.
# observability/base/victoria-metrics-k8s-stack/vmrules/runlore.yaml — trimmed
- alert: RunloreAgentDown
  expr: absent(runlore_build_info)
  for: 5m
  labels:
    severity: critical
  annotations:
    runbook_url: "https://github.com/Smana/runlore/blob/main/docs/observability.md#runloreagentdown"

Alertmanager routing

Alertmanager’s routing tree (vm-common-helm-values-configmap.yaml, applied to the victoria-metrics-k8s-stack alertmanager.spec.config) fans every non-blackholed alert to both RunLore and Slack, then splits by severity so one channel can carry a page and a nag without them looking alike:

route:
  receiver: "slack-monitoring"
  repeat_interval: 12h
  routes:
    - matchers:
        - alertname =~ "InfoInhibitor|Watchdog|KubeCPUOvercommit"
      receiver: "blackhole"
    - receiver: "runlore"
      continue: true   # falls through to the next route instead of stopping
    - matchers: [severity = "critical"]
      receiver: "slack-monitoring"
      group_wait: 10s
      repeat_interval: 1h
    - matchers: [severity = "warning"]
      receiver: "slack-monitoring"
      repeat_interval: 12h
    - receiver: "slack-monitoring"   # no matchers — see below
      repeat_interval: 24h

The final route’s absence of matchers is load-bearing. severity = "critical" does not match an alert that has no severity label — a missing label is the empty string. That last route is the catch-all keeping an unlabelled alert reachable.

The dangerous edit is the well-intentioned one. Giving it an explicit matcher like severity = "info" reads more precise and would pass review, and an alert arriving without a severity label would then match no route and reach nobody. No gate would fail: amtool accepts it, the schema gate sees a valid string, and the golden files do not move, because routing changes what is delivered and never how it is drawn. The only symptom is an alert that quietly stops arriving.

Two inhibit rules, both deliberately narrow:

inhibit_rules:
  - source_matchers: [severity = "critical"]
    target_matchers: [severity = "warning"]
    equal: [cluster, alertname, namespace]
  - source_matchers: [alertname = "OpenBaoRaftQuorumAtRisk"]
    target_matchers: [alertname = "OpenBaoRaftNodeLost"]
    equal: [cluster]

The first catches one alertname firing at two severities. The second exists because OpenBaoRaftQuorumAtRisk and OpenBaoRaftNodeLost are different alertnames where one implies the other — quorum-at-risk means a peer is already lost — and they arrived a minute apart on 2026-09-12. Add named pairs as you observe them rather than generalising: an inhibited alert never reaches Slack at all, so an over-broad rule silently deletes alerts, which is worse than a duplicate.

RunLore’s webhook requires a bearer token (its v0.2.0+ fail-closed behavior — alert labels/annotations flow into an LLM prompt, so an unauthenticated trigger path was judged unacceptable) mirrored between its own runlore-webhook ExternalSecret and this stack’s runlore-webhook-token ExternalSecret.

The Slack message

Rendered by repo-owned templates shipped through alertmanager.templateFiles, with the chart’s vendored Monzo set disabled — see ADR-0037 for why these are Alertmanager-native rather than Block Kit behind a bridge.

:fire: FIRING — OpenBaoRaftQuorumAtRisk
`aws-0 · dev`

OpenBao raft cluster cannot tolerate a node failure.
 • 10.0.12.44:8200 — Failure tolerance has been below 1 for 15 minutes: losing
   one more node loses quorum, and a cluster without quorum cannot issue a
   certificate or read a secret. If OpenBaoRaftNodeLo…

Namespace  security          Severity  critical
Duration   14m (since 09:06 UTC)   Location  aws / eu-west-3

[ Runbook ] [ Dashboard ] [ Query ] [ Silence ]

The identity line (aws-0 · dev) and the Location field come from labels vmalert stamps in externalLabels — cluster, env, cloud, region — fed by each cluster’s flux_cluster_vars ConfigMap. They are stamped after rule evaluation, which is the point: most upstream rules aggregate the cluster label away, and a message once reached Slack reading on cluster .

Optional values degrade rather than disappear: a missing namespace renders —, an empty cloud drops that half of Location, and a group of more than five alerts shows five bullets and … showing 5 of N.

The annotation contract

./scripts/ci/validate-vmrules.sh enforces the first line of this on every repo-authored alert:

AnnotationRequiredRendered as
summaryyes — one line, ≤140 charsthe headline
descriptionno, any lengthper-alert bullet, truncated at 180 chars
runbook_urlnoRunbook button; falls back to this page
dashboardnoDashboard button and title_link; falls back to Grafana’s home

description is deliberately left unbounded. Slack truncates it; RunLore receives the same annotations over its webhook and reads all of it, so trimming operational prose out of a rule to make it fit a chat message would blind the agent to the one thing that explains the alert.

Changing the wording

Edit templateFiles.ogenki.tmpl, then:

./scripts/ci/validate-alertmanager-templates.sh --update-golden
git diff scripts/ci/tests/alertmanager-fixtures/golden/   # read it — it IS the message
./scripts/ci/validate-alertmanager-templates.sh

Two traps bind that block. Never write a literal ${ — Flux post-build substitution expands ${var} and replaces an unknown one with an empty string, and this applies inside comments too; a bare $var (every Go template variable) is safe. Alertmanager ships no sprig — no default, no arithmetic. A template calling one fails to execute, and Alertmanager then drops the notification, which is why the gate exists at all.

The gate checks every rendered Alertmanager config with amtool check-config, renders every templated string against five fixture payloads, and asserts every VMAlert carries an absolute external.url. It validates structure, not semantics: a typo’d equal label such as clustre is a syntactically valid label name and passes. So do a shadowing route and an over-broad inhibit rule.

Grafana OnCall (removed)

The former grafana-oncall directory under observability/base/ — a complete engine + RabbitMQ + Postgres + Valkey install that no Flux Kustomization ever referenced — was removed on 2026-08-29, along with the grafana-oncall-app plugin and the provisioning ConfigMap that pointed it at an oncall-engine Service that never existed. Upstream OnCall OSS is archived (read-only on GitHub since 2026-06-05, zero future patches), and the RunLore + Slack routing above already carries the whole incident flow. ADR-0029 records the decision and the full inventory of what was deleted.

CiliumNetworkPolicy coverage

Of the eight component directories under observability/base/, only runlore defines a CiliumNetworkPolicy:

# observability/base/runlore/ciliumnetworkpolicy-ingress.yaml
spec:
  endpointSelector:
    matchLabels:
      app.kubernetes.io/name: runlore
  ingress:
    - fromEntities: [ingress]
      toPorts:
        - ports: [{port: "8080", protocol: TCP}]

It exists because runlore’s HTTPRoute parents to platform-public (described in Private Access), and this cluster’s shared cilium-envoy DaemonSet runs hostNetwork: true — Gateway-originated traffic lands as Cilium’s reserved ingress entity, which a namespaced Kubernetes NetworkPolicy can never match. grafana-operator, kubernetes-event-exporter, loggen, metrics-server, victoria-logs, victoria-metrics-k8s-stack, and victoria-traces have no CiliumNetworkPolicy in these directories, and no cluster-wide default-deny CiliumClusterwideNetworkPolicy covers the observability namespace either — the one such policy in this repo (infrastructure/base/gapi/allow-gateway-l7-proxy.yaml) only allows the Gateway API L7 proxy, it does not deny anything by default.