diff --git a/alloy/README.md b/alloy/README.md index 963336e..29b17e3 100644 --- a/alloy/README.md +++ b/alloy/README.md @@ -64,6 +64,10 @@ Runs `privileged: true` + `network_mode: host`, matching the upstream Grafana Cl | `/etc/machine-id:ro` | stable host id for the journal reader | | `alloy-data` (named volume) | WAL + remotecfg cache | +## Design + +- [Upstream sources of truth](docs/upstream-sources-of-truth.md) — what we follow, what's in scope, how to audit dashboard metric needs. + ## Known noise (special cases) - [Coolify SSH session spam](docs/known-noise-coolify-ssh-sessions.md) — only relevant if Coolify manages the host. Safe to ignore otherwise. diff --git a/alloy/docs/upstream-sources-of-truth.md b/alloy/docs/upstream-sources-of-truth.md new file mode 100644 index 0000000..0f77436 --- /dev/null +++ b/alloy/docs/upstream-sources-of-truth.md @@ -0,0 +1,95 @@ +# Upstream Sources of Truth + +## Hard rule: official sources only + +> **Only use official sources** when deciding what metrics, labels, log pipelines, dashboards, or rules this repo ships. **Never** consult or copy from third-party repos, blog posts, community Gists, or personal exports — even if the names match. +> +> If an answer can't be found in an official source, document the gap rather than fill it with unofficial data. **Non-negotiable.** + +### What counts as an official source + +| Tier | Source | Use for | +|---|---|---| +| 1 | `https://grafana.com/docs/grafana-cloud/...` (Grafana Cloud integration reference pages) | Alloy snippet, scrape config, label conventions, drop/keep rules | +| 2 | `github.com/grafana/...` (Grafana Labs org repos) | Mixin libsonnet/jsonnet, `dashboards_out/*.json`, alerting rules | +| 3 | `github.com/prometheus/...` (Prometheus org repos) | Upstream node_exporter mixin, prom-mixin libraries | +| 4 | The Grafana Cloud stack's own running dashboards (via authenticated API) | Authoritative answer for "what does my live integration use" | + +### What doesn't count (do not use) + +- Personal forks, Terraform exports, ops-team copies of integration dashboards. +- `grafana.com/grafana/dashboards/` community submissions, unless explicitly published by the `grafana` user. +- AI summaries of dashboards. +- Stale snapshots in unrelated repos that happen to contain `job=integrations/...` strings. + +## Principle + +Alloy here ships **only what the integration dashboards need** — not the full firehose of every metric/log line node_exporter and cadvisor can produce. The goal is dashboards that work out-of-the-box on Grafana Cloud, without paying for unused active series. + +When the upstream integration changes (adds a metric, drops a panel, renames a label), this repo follows. Local additions outside that scope go into a separate, named, documented overlay. + +## Upstream references + +The two reference pages this repo's `docker-compose.yml` mirrors: + +- **Linux Node integration** — +- **Docker integration** — + +Where official mixin dashboards exist they're tier-2 corroboration: + +- Docker mixin (Grafana Labs): +- Node-exporter mixin (Prometheus project): + +## Mapping our config to upstream + +| Block in `docker-compose.yml` | Upstream source | +|---|---| +| `prometheus.exporter.unix` (collectors, mounts, fs/net excludes) | Linux Node integration page → "Configure Alloy" | +| `prometheus.relabel "integrations_node_exporter"` (drop `node_scrape_collector_*`) | Linux Node integration page → drop rule snippet | +| `loki.source.journal "default"` + relabel rules (`unit`, `boot_id`, `transport`, `level`) | Linux Node integration page → log scraping | +| `loki.source.file` for `/var/log/{syslog,messages,*.log}` | Linux Node integration page → file log scraping | +| `prometheus.exporter.cadvisor` (`docker_only = true`) | Docker integration page | +| `prometheus.relabel "integrations_cadvisor"` keep allowlist | Docker integration page → metric list | +| `discovery.docker` + `loki.source.docker` (job/instance/container/stream) | Docker integration page → log scraping | + +## What "follow upstream" means in practice + +1. **Don't expand the metric set unilaterally.** If a panel in a Grafana Cloud dashboard requires a metric we don't ship, we'd add it — but the trigger is "the integration dashboard needs it, verified from a tier 1–4 source", not "node_exporter exposes it". +2. **Don't filter further than upstream does.** We ship at least what the upstream config does. Tightening (e.g. for cost) goes in a clearly-named overlay or a downstream env-specific config. +3. **Re-check on Alloy/integration major versions.** When bumping `grafana/alloy` image or when the Grafana Cloud integration revs, diff the upstream Alloy snippet against `docker-compose.yml` and update. + +## Audit (2026-04-26) + +Verified with tier 1 + tier 2 sources only. The Grafana Cloud integration's full dashboard set (the 7 Linux-Node dashboards + 2 Docker dashboards) is **not** publicly hosted; auditing them requires tier 4 access (an authenticated Grafana Cloud API call against your own stack). That step was not performed and is the only known gap. + +- **Linux-Node** — tier 1 specifies `drop "node_scrape_collector_.+"`, tier 3 (`prometheus/node_exporter` mixin) confirms that the metrics that mixin's dashboards reference all pass through the drop rule. Config matches. **No action.** +- **Docker** — tier 1 lists 14 metrics + `up` / `machine_scrape_error` (16 total). Tier 2 (`grafana/jsonnet-libs` `docker.json`) confirms the same 14 panel-referenced metrics. Config matches. **No action.** + +### Re-running the audit + +```bash +# Tier 2 — Grafana Labs docker-mixin +curl -sL https://raw.githubusercontent.com/grafana/jsonnet-libs/master/docker-mixin/dashboards_out/docker.json \ + | jq -r '.. | .expr? // empty' \ + | grep -oE 'container_[a-zA-Z0-9_]+|machine_[a-zA-Z0-9_]+' | sort -u + +# Tier 3 — Prometheus node-exporter mixin (libsonnet, grep for metric names) +curl -sL https://raw.githubusercontent.com/prometheus/node_exporter/master/docs/node-mixin/lib/prom-mixin.libsonnet \ + | grep -oE 'node_[a-zA-Z0-9_]+|process_[a-zA-Z0-9_]+' | sort -u + +# Tier 4 — your own Grafana Cloud stack +GRAFANA_TOKEN=... # service-account token, dashboards:read +STACK=.grafana.net +curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \ + "https://$STACK/api/search?query=Linux+Node&type=dash-db" \ + | jq -r '.[].uid' | while read uid; do + curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \ + "https://$STACK/api/dashboards/uid/$uid" > "$uid.json" + done +# then jq + grep as above +``` + +## Out of scope + +- Custom dashboards or metrics for app-level monitoring (deploy a separate Alloy config for that). +- Host-specific filtering (e.g. dropping noisy log lines from a particular daemon — see `docs/known-noise-coolify-ssh-sessions.md` for an example).