cadvisor regex: drop fs_inodes/fs_limit/network_tcp_usage; add fs_reads
+ fs_writes + network_(receive|transmit)_(errors|packets_dropped)_total
to match the standard Grafana Cloud docker integration dashboard.
Compose:
- pid: host so prometheus.exporter.unix sees host /proc (cpu, mem,
load, processes) instead of the container's namespace.
- mount /etc/machine-id so loki.source.journal has a stable host id.
bind-mounting /dev/kmsg under volumes: creates the node but leaves the
device-cgroup controller blocking the read (EPERM). cap_drop=[ALL]
clears the default device allow-list, so even with CAP_SYSLOG the
kernel refuses. moving it under devices: adds the cgroup allow rule
alongside the bind-mount, which is what cadvisor actually needs.
clears two startup warnings:
- cadvisor "Could not configure a source for OOM detection" — needs
/dev/kmsg bind-mount and CAP_SYSLOG (kernel.dmesg_restrict=1 default)
- node-exporter "Failed to open /run/udev/data" — diskstats collector
enriches node_disk_* with model/serial/WWN labels from udev
both mounted read-only. CAP_SYSLOG added alongside DAC_OVERRIDE.
grafana cloud loki rejects entries older than 7 days with HTTP 400.
loki.source.docker has no "tail since" option, so on first start it
replays logs from each container's start time — long-running
containers (coolify-proxy, traffmonetizer) produced weeks of backlog
that loki refused to ingest.
insert a loki.process drop stage (older_than=144h) between the docker
source and loki.write.gc so stale entries are filtered before egress.
journal source already bounded by max_age=12h — no filter needed there.
container runs as root (0:0) but cap_drop=[ALL] stripped DAC_OVERRIDE,
so mkdir on /var/lib/alloy/data (alloy-owned in the image) failed with
permission denied. add it back (only DAC_OVERRIDE, nothing else) and
mount /etc/machine-id ro for a stable journal host id. annotate every
volume with the component that uses it.
- merge discovery.relabel "metrics" into prometheus.relabel "filter"
(instance label + keep-filter in one component)
- move logs 'instance' label to loki.write external_labels (DRY, one source)
- drop alloy container log exclusion (~500kb/day trivial; adds visibility
into alloy's own health via loki)
net: -5 lines, one fewer component, same functionality
correctness:
- drop dead ALLOY_HOSTNAME env passthrough (alloy uses constants.hostname)
- tighten netdev regex so iface names like 'lore0' no longer match 'lo'
- widen alloy-container drop to /alloy.* (catches renamed variants)
more info for same cost (~80 extra series, still ~1.6k of 10k):
- add MemFree, disk IOPS (reads/writes_completed), disk saturation
(io_time_weighted), inode tracking (filesystem_files, _files_free),
network drops, container working_set memory, container_last_seen,
machine_scrape_error
- add journal unit + level labels (log querying by service/severity)
- docker discovery refresh 60s -> 5s (fast pickup of new containers)
efficiency:
- loki batch_wait=5s, batch_size=1MiB (~5x fewer push requests)
layout:
- drop redundant expose: [12345] (other containers can reach it anyway)
- node_exporter (host) + cadvisor (containers) via one scrape job
- docker container logs + systemd journal to loki
- filter regex keeps only metrics needed for standard dashboards
- sized for grafana cloud free tier (~10k active series cap)
- shell env driven (no .env file); fail-fast on missing vars
- least-privilege: cap_drop all, no privileged, no pid host, ui not published