docs: codify upstream-sources-only rule for config decisions

Adds docs/upstream-sources-of-truth.md as the binding policy for what
this repo follows when deciding metrics, labels, log pipelines, and
dashboards to ship.

Hard rule: only tier 1-4 official sources (Grafana Cloud integration
docs, github.com/grafana/*, github.com/prometheus/*, the user's own
authenticated Grafana Cloud API). No third-party Terraform exports,
community gists, blog posts, or AI summaries — even when names match.

Records a tier-1+2 audit confirming the current cadvisor allowlist
matches both the Docker integration page and grafana/jsonnet-libs
docker-mixin/docker.json. Notes that tier-4 verification against the
live stack's full integration dashboard set was not performed and is
the only known gap.
This commit is contained in:
tiennm99 committed 2026-04-26 09:50:20 +07:00
1 parent 54ab1fed27
commit c1db79a359
2 files changed
+99

No files matched your search

+4
View File
@@ -64,6 +64,10 @@ Runs `privileged: true` + `network_mode: host`, matching the upstream Grafana Cl
| `/etc/machine-id:ro` | stable host id for the journal reader |
| `alloy-data` (named volume) | WAL + remotecfg cache |
## Design
- [Upstream sources of truth](docs/upstream-sources-of-truth.md) — what we follow, what's in scope, how to audit dashboard metric needs.
## Known noise (special cases)
- [Coolify SSH session spam](docs/known-noise-coolify-ssh-sessions.md) — only relevant if Coolify manages the host. Safe to ignore otherwise.
+95
View File
@@ -0,0 +1,95 @@
# Upstream Sources of Truth
## Hard rule: official sources only
> **Only use official sources** when deciding what metrics, labels, log pipelines, dashboards, or rules this repo ships. **Never** consult or copy from third-party repos, blog posts, community Gists, or personal exports — even if the names match.
>
> If an answer can't be found in an official source, document the gap rather than fill it with unofficial data. **Non-negotiable.**
### What counts as an official source
| Tier | Source | Use for |
|---|---|---|
| 1 | `https://grafana.com/docs/grafana-cloud/...` (Grafana Cloud integration reference pages) | Alloy snippet, scrape config, label conventions, drop/keep rules |
| 2 | `github.com/grafana/...` (Grafana Labs org repos) | Mixin libsonnet/jsonnet, `dashboards_out/*.json`, alerting rules |
| 3 | `github.com/prometheus/...` (Prometheus org repos) | Upstream node_exporter mixin, prom-mixin libraries |
| 4 | The Grafana Cloud stack's own running dashboards (via authenticated API) | Authoritative answer for "what does my live integration use" |
### What doesn't count (do not use)
- Personal forks, Terraform exports, ops-team copies of integration dashboards.
- `grafana.com/grafana/dashboards/<id>` community submissions, unless explicitly published by the `grafana` user.
- AI summaries of dashboards.
- Stale snapshots in unrelated repos that happen to contain `job=integrations/...` strings.
## Principle
Alloy here ships **only what the integration dashboards need** — not the full firehose of every metric/log line node_exporter and cadvisor can produce. The goal is dashboards that work out-of-the-box on Grafana Cloud, without paying for unused active series.
When the upstream integration changes (adds a metric, drops a panel, renames a label), this repo follows. Local additions outside that scope go into a separate, named, documented overlay.
## Upstream references
The two reference pages this repo's `docker-compose.yml` mirrors:
- **Linux Node integration** — <https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-linux-node/>
- **Docker integration** — <https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-docker/>
Where official mixin dashboards exist they're tier-2 corroboration:
- Docker mixin (Grafana Labs): <https://github.com/grafana/jsonnet-libs/blob/master/docker-mixin/dashboards_out/docker.json>
- Node-exporter mixin (Prometheus project): <https://github.com/prometheus/node_exporter/tree/master/docs/node-mixin>
## Mapping our config to upstream
| Block in `docker-compose.yml` | Upstream source |
|---|---|
| `prometheus.exporter.unix` (collectors, mounts, fs/net excludes) | Linux Node integration page → "Configure Alloy" |
| `prometheus.relabel "integrations_node_exporter"` (drop `node_scrape_collector_*`) | Linux Node integration page → drop rule snippet |
| `loki.source.journal "default"` + relabel rules (`unit`, `boot_id`, `transport`, `level`) | Linux Node integration page → log scraping |
| `loki.source.file` for `/var/log/{syslog,messages,*.log}` | Linux Node integration page → file log scraping |
| `prometheus.exporter.cadvisor` (`docker_only = true`) | Docker integration page |
| `prometheus.relabel "integrations_cadvisor"` keep allowlist | Docker integration page → metric list |
| `discovery.docker` + `loki.source.docker` (job/instance/container/stream) | Docker integration page → log scraping |
## What "follow upstream" means in practice
1. **Don't expand the metric set unilaterally.** If a panel in a Grafana Cloud dashboard requires a metric we don't ship, we'd add it — but the trigger is "the integration dashboard needs it, verified from a tier 1–4 source", not "node_exporter exposes it".
2. **Don't filter further than upstream does.** We ship at least what the upstream config does. Tightening (e.g. for cost) goes in a clearly-named overlay or a downstream env-specific config.
3. **Re-check on Alloy/integration major versions.** When bumping `grafana/alloy` image or when the Grafana Cloud integration revs, diff the upstream Alloy snippet against `docker-compose.yml` and update.
## Audit (2026-04-26)
Verified with tier 1 + tier 2 sources only. The Grafana Cloud integration's full dashboard set (the 7 Linux-Node dashboards + 2 Docker dashboards) is **not** publicly hosted; auditing them requires tier 4 access (an authenticated Grafana Cloud API call against your own stack). That step was not performed and is the only known gap.
- **Linux-Node** — tier 1 specifies `drop "node_scrape_collector_.+"`, tier 3 (`prometheus/node_exporter` mixin) confirms that the metrics that mixin's dashboards reference all pass through the drop rule. Config matches. **No action.**
- **Docker** — tier 1 lists 14 metrics + `up` / `machine_scrape_error` (16 total). Tier 2 (`grafana/jsonnet-libs` `docker.json`) confirms the same 14 panel-referenced metrics. Config matches. **No action.**
### Re-running the audit
```bash
# Tier 2 — Grafana Labs docker-mixin
curl -sL https://raw.githubusercontent.com/grafana/jsonnet-libs/master/docker-mixin/dashboards_out/docker.json \
| jq -r '.. | .expr? // empty' \
| grep -oE 'container_[a-zA-Z0-9_]+|machine_[a-zA-Z0-9_]+' | sort -u
# Tier 3 — Prometheus node-exporter mixin (libsonnet, grep for metric names)
curl -sL https://raw.githubusercontent.com/prometheus/node_exporter/master/docs/node-mixin/lib/prom-mixin.libsonnet \
| grep -oE 'node_[a-zA-Z0-9_]+|process_[a-zA-Z0-9_]+' | sort -u
# Tier 4 — your own Grafana Cloud stack
GRAFANA_TOKEN=... # service-account token, dashboards:read
STACK=<your-stack>.grafana.net
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"https://$STACK/api/search?query=Linux+Node&type=dash-db" \
| jq -r '.[].uid' | while read uid; do
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"https://$STACK/api/dashboards/uid/$uid" > "$uid.json"
done
# then jq + grep as above
```
## Out of scope
- Custom dashboards or metrics for app-level monitoring (deploy a separate Alloy config for that).
- Host-specific filtering (e.g. dropping noisy log lines from a particular daemon — see `docs/known-noise-coolify-ssh-sessions.md` for an example).