Document the compose services, published ports, volumes, and first-run
setup. Fill .env.example with sample values and note that compose.yml
does not yet consume them.
container_memory_usage_bytes counts page cache attributed to the cgroup,
which makes disk-heavy containers (e.g. gitea) appear to use ~all host RAM.
Add container_memory_working_set_bytes (the metric Grafana's Docker
integration dashboard expects), plus container_memory_rss,
container_memory_cache, and container_spec_memory_limit_bytes for
breakdown and limit-percentage panels.
`loki.source.journal "default"` no longer pins `path = "/var/log/journal"`.
Upstream omits the field, which lets Alloy default to BOTH
`/var/log/journal` (persistent) and `/run/log/journal` (volatile). The
explicit path silently dropped journal logs on hosts with volatile-only
storage. Matches the canonical Linux Node integration template.
Also bumps the image six minor versions to current stable. Doc records
the 2026-04-26 re-audit.
Copies the exact 157-metric list from the Linux Node integration's
Metrics anchor as the cadvisor keep-list already does for Docker
(16 metrics). Replaces the earlier `drop node_scrape_collector_.+`
rule, which was the integration page's alternate snippet but didn't
ship the explicit allowlist users see in the docs.
`instance:node_num_cpu:sum` from the Metrics section is intentionally
omitted — it's a recording-rule output computed server-side by
Grafana Cloud's ruler, not produced by the agent.
Doc + README updated to point at the Metrics anchors directly so the
source of truth is unambiguous.
Adds docs/upstream-sources-of-truth.md as the binding policy for what
this repo follows when deciding metrics, labels, log pipelines, and
dashboards to ship.
Hard rule: only tier 1-4 official sources (Grafana Cloud integration
docs, github.com/grafana/*, github.com/prometheus/*, the user's own
authenticated Grafana Cloud API). No third-party Terraform exports,
community gists, blog posts, or AI summaries — even when names match.
Records a tier-1+2 audit confirming the current cadvisor allowlist
matches both the Docker integration page and grafana/jsonnet-libs
docker-mixin/docker.json. Notes that tier-4 verification against the
live stack's full integration dashboard set was not performed and is
the only known gap.
Linux-Node integration:
- replace curated keep-list of ~140 node_* metrics with the upstream
drop rule (drops only node_scrape_collector_*); ships the full
~130+ metric set the integration dashboards expect.
- add loki.source.file for /var/log/{syslog,messages,*.log} alongside
the existing journal scrape, matching the upstream config.
- broaden the /var/log mount to cover both pipelines (was journal only).
Docker integration:
- drop container_memory_working_set_bytes from the cadvisor allowlist;
not part of the documented metric set.
README: refresh "What it collects" + "Mounts" tables, document the
syslog-vs-journald duplication caveat for rsyslog hosts.
Document a Coolify-specific noise pattern observed in the journal
pipeline: ~300 root sessions/hour from the Coolify host's connection
checks. Verified against coollabsio/coolify v4.x source (Kernel.php,
ServerManagerJob, ServerCheckJob, SshMultiplexingHelper).
Includes:
- exact call flow and skip conditions per Coolify source
- triage commands and key-fingerprint matcher
- two mitigations: drop at Alloy (loki.process stage.drop) or enable
Sentinel server-side to bypass the SSH polling entirely
- framing: Coolify-only, base setup unchanged
Previous step failed in CI because the bare alloy container had no
/var/log/journal, no docker.sock, etc., so loki.source.journal +
discovery.docker exited the process — and --rm wiped the container
before logs could be inspected.
Now: drop --rm, mount the same host paths the prod compose uses, plus
an empty /var/log/journal stand-in. Capture logs unconditionally and
fail only on config-level patterns ('unknown component',
'undefined reference', 'syntax error', etc.). Runtime/component
failures against dummy endpoints are tolerated.
- network_mode: host so prometheus.exporter.unix reports real host
interfaces (eth0...) rather than the alloy container's veth pair.
- loki.source.journal: set path = "/var/log/journal" explicitly so it
doesn't silently fall through to /run/log/journal on volatile-journal
hosts.
- cadvisor keep-list: add container_memory_working_set_bytes (drives
several panels on the standard Docker dashboard).
- Drop /dev/kmsg device + extra_hosts:host.docker.internal — neither is
needed by the current keep-lists, and host-network mode makes the
extra_hosts entry meaningless.
- CI: extend Alloy validation beyond `fmt` (syntax-only) by booting
alloy with the embedded config and asserting it stays running, which
catches bad component refs / wrong arg names that fmt accepts.
- README: refresh Mounts table + Security note to match.
prometheus.exporter.unix now reads /rootproc and /rootfs (the existing
host bind mounts) instead of the container's own namespace, so the
filesystem + process metrics actually describe the host. This makes
pid:host unnecessary, so remove it — privileged is enough and pid:host
exposes every host process inside the container.
cadvisor regex: drop fs_inodes/fs_limit/network_tcp_usage; add fs_reads
+ fs_writes + network_(receive|transmit)_(errors|packets_dropped)_total
to match the standard Grafana Cloud docker integration dashboard.
Compose:
- pid: host so prometheus.exporter.unix sees host /proc (cpu, mem,
load, processes) instead of the container's namespace.
- mount /etc/machine-id so loki.source.journal has a stable host id.
bind-mounting /dev/kmsg under volumes: creates the node but leaves the
device-cgroup controller blocking the read (EPERM). cap_drop=[ALL]
clears the default device allow-list, so even with CAP_SYSLOG the
kernel refuses. moving it under devices: adds the cgroup allow rule
alongside the bind-mount, which is what cadvisor actually needs.
clears two startup warnings:
- cadvisor "Could not configure a source for OOM detection" — needs
/dev/kmsg bind-mount and CAP_SYSLOG (kernel.dmesg_restrict=1 default)
- node-exporter "Failed to open /run/udev/data" — diskstats collector
enriches node_disk_* with model/serial/WWN labels from udev
both mounted read-only. CAP_SYSLOG added alongside DAC_OVERRIDE.
grafana cloud loki rejects entries older than 7 days with HTTP 400.
loki.source.docker has no "tail since" option, so on first start it
replays logs from each container's start time — long-running
containers (coolify-proxy, traffmonetizer) produced weeks of backlog
that loki refused to ingest.
insert a loki.process drop stage (older_than=144h) between the docker
source and loki.write.gc so stale entries are filtered before egress.
journal source already bounded by max_age=12h — no filter needed there.
container runs as root (0:0) but cap_drop=[ALL] stripped DAC_OVERRIDE,
so mkdir on /var/lib/alloy/data (alloy-owned in the image) failed with
permission denied. add it back (only DAC_OVERRIDE, nothing else) and
mount /etc/machine-id ro for a stable journal host id. annotate every
volume with the component that uses it.
- merge discovery.relabel "metrics" into prometheus.relabel "filter"
(instance label + keep-filter in one component)
- move logs 'instance' label to loki.write external_labels (DRY, one source)
- drop alloy container log exclusion (~500kb/day trivial; adds visibility
into alloy's own health via loki)
net: -5 lines, one fewer component, same functionality
correctness:
- drop dead ALLOY_HOSTNAME env passthrough (alloy uses constants.hostname)
- tighten netdev regex so iface names like 'lore0' no longer match 'lo'
- widen alloy-container drop to /alloy.* (catches renamed variants)
more info for same cost (~80 extra series, still ~1.6k of 10k):
- add MemFree, disk IOPS (reads/writes_completed), disk saturation
(io_time_weighted), inode tracking (filesystem_files, _files_free),
network drops, container working_set memory, container_last_seen,
machine_scrape_error
- add journal unit + level labels (log querying by service/severity)
- docker discovery refresh 60s -> 5s (fast pickup of new containers)
efficiency:
- loki batch_wait=5s, batch_size=1MiB (~5x fewer push requests)
layout:
- drop redundant expose: [12345] (other containers can reach it anyway)
- node_exporter (host) + cadvisor (containers) via one scrape job
- docker container logs + systemd journal to loki
- filter regex keeps only metrics needed for standard dashboards
- sized for grafana cloud free tier (~10k active series cap)
- shell env driven (no .env file); fail-fast on missing vars
- least-privilege: cap_drop all, no privileged, no pid host, ui not published