Files
tiennm99 d889338306 fix(alloy): drop docker log lines older than Loki accepts
loki.source.docker resumes from a one-second position and Docker's since
is inclusive, so each restart re-reads a container's last logged second.
For idle containers those lines are older than 7 days and Grafana Cloud
Loki rejects the batch with 400 timestamp too old. Drop them before
loki.write; Loki would not have stored them anyway.
2026-10-06 16:56:46 +07:00

8.2 KiB
Raw Permalink Blame History

alloy

Grafana Alloy shipping host and container telemetry to Grafana Cloud, with remote config from Grafana Fleet Management.

One container runs both the node_exporter (host) and cadvisor (container) collectors. A second, tiny container proxies a read-only slice of the Docker API to it. The Alloy config is embedded inline via Compose configs:, so there is no config.alloy on disk.

dockerproxy publishes 127.0.0.1:2375 — Alloy runs with host networking, so it has no compose network to reach the proxy over.

What it collects

Source Component Notes
Host metrics prometheus.exporter.unix CPU, memory, load, disk I/O, filesystem, network, uname, boot time, systemd, vmstat, sockstat — the default collector set minus ipvs/btrfs/infiniband/xfs/zfs
Container metrics prometheus.exporter.cadvisor CPU, memory, fs usage/limit, network, last_seen. Docker labels are not copied onto the series (store_container_labels = false): Coolify gives its containers 50–80 labels, which pushes series past Grafana Cloud's 60-label limit and gets them rejected
Container logs loki.source.docker All running containers, labeled container, stream, instance
Journal logs loki.source.journal systemd journal, labeled unit, boot_id, transport, level
File logs loki.source.file /var/log/syslog, /var/log/messages, /var/log/*.log
Remote config remotecfg Polls Grafana Fleet Management every 60s

Metric filtering copies the keep-lists from the upstream Grafana Cloud integrations verbatim (Linux Node, Docker). Logs are unfiltered, except that container log lines older than 168h are dropped before they are sent. Grafana Cloud Loki rejects anything older than 7 days, and loki.source.docker re-reads a container's last second of logs on every restart (its saved position has one-second precision), so an idle container's last lines would otherwise come back as a 400 timestamp too old error on every redeploy. Nothing Loki would have accepted is dropped.

Environment

All nine are required; docker compose up fails fast if any is unset.

Variable Purpose
ALLOY_HOSTNAME Container hostname, and the Loki/Prometheus instance label
REMOTECFG_URL Fleet Management endpoint
REMOTECFG_ID Fleet Management agent id
REMOTECFG_USER Fleet Management user id
PROM_URL / PROM_USER Prometheus remote-write endpoint and user id
LOKI_URL / LOKI_USER Loki push endpoint and user id
GRAFANA_TOKEN One Cloud Access Policy token, scopes metrics:write + logs:write + fleet-management:read

GRAFANA_TOKEN is also passed into the container as GCLOUD_RW_API_KEY. Fleet Management's auto-generated self_monitoring_* pipelines read the token from that fixed name; without it their remote-write gets an empty password and Grafana Cloud answers 401 invalid token, so the collector shows no health data in Fleet Management.

Find the values under Grafana Cloud → your stack → Details on each data source, and under Fleet Management. The same token serves remotecfg, Prometheus and Loki basic-auth.

cp .env.example .env   # then fill in the values
docker compose up -d

Run the same file on every host, changing ALLOY_HOSTNAME and REMOTECFG_ID per host. Filter in Grafana with instance=~"...".

Privileges

The upstream Grafana Cloud docker integration runs privileged: true with the raw Docker socket mounted read-write. Nothing in this config needs that, and the combination is host-root-equivalent: privileged grants every capability and unmasks /proc and /sys, and a container that can talk to the Docker socket can start another container that mounts / writable. The socket's :ro flag does not help — it stops the socket file being replaced, not the API being used.

What it runs instead:

Setting Reason
cap_drop: [ALL] + cap_add: [DAC_OVERRIDE] The image's entrypoint runs as uid 0 and reads host files owned by other users — the journal, paths under /rootfs, /var/log. Dropping every capability leaves it unable to open them, and unable to create its own storage directory. DAC_OVERRIDE restores exactly that and nothing else; SYS_ADMIN, NET_ADMIN, SYS_PTRACE, MKNOD and the rest stay dropped.
no-new-privileges:true No setuid binary in the image can regain what was dropped.
mem_limit: 2g, pids_limit: 512 Steady state is around 900 MB; the limit stops a leak taking the host down with it.
dockerproxy instead of /var/run/docker.sock See below.

read_only: true is deliberately absent. Compose materialises an inline configs: entry by writing it into the container, and refuses to do that on a read-only service — cannot create config ... : \file` is the sole supported option. Keeping the config inline is worth more than the read-only rootfs here; adding it back means moving the config to a config.alloyfile on disk and switching theconfigs:entry tofile:`.

network_mode: host stays. /proc/net is a symlink to /proc/self/net and resolves against the reading process's network namespace, so bind-mounting the host's /proc to /rootproc is not enough — without host networking the netdev, netstat, sockstat and conntrack families would describe the container's veth pair instead of eth0. That is roughly 45 of the 157 kept metrics.

/:/rootfs:ro also stays, and is the widest remaining exposure: the container can read every file on the host. The filesystem collector needs to statfs() each mount point, and a bind mount cannot grant that without granting reads. Dropping it would cost disk-space monitoring. Treat the container as secret-bearing — it holds GRAFANA_TOKEN regardless.

Docker API access

prometheus.exporter.cadvisor and loki.source.docker both need the Docker API, so it cannot simply be removed. dockerproxy runs tecnativa/docker-socket-proxy with the socket mounted read-only and publishes it on 127.0.0.1:2375, which Alloy reaches over host networking.

POST is revoked by default in that image, which is the point: container create, exec, start and kill are all refused with 403, so a compromise of Alloy can no longer become root on the host. The granted sections are the ones the two components actually call:

Variable Called by
CONTAINERS discovery.docker listing, loki.source.docker inspect and log read, cadvisor metadata
NETWORKS discovery.docker — Prometheus' Docker SD resolves network names per container, and returns zero targets without it
IMAGES, INFO, VERSION cadvisor
EVENTS cadvisor container watch (granted by default)

What this does not do: GET /containers/{id}/json still returns every container's environment variables. The proxy closes the escalation path, not the disclosure one.

Mounts

Mount Why
/proc:/rootproc:ro node-exporter cpu/mem/load, via procfs_path
/sys:/sys:ro node-exporter and cadvisor cgroups
/:/rootfs:ro filesystem collector, via rootfs_path
/dev/disk/:/dev/disk:ro node-exporter diskstats device labels
/run/udev:/run/udev:ro node-exporter diskstats udev device properties (/run/udev/data)
/var/lib/docker:ro cadvisor container metadata
/var/log:/var/log:ro loki.source.journal and loki.source.file
/etc/machine-id:ro Stable host id for the journal reader
alloy-data WAL and remotecfg cache — the only writable path

dockerproxy mounts /var/run/docker.sock:ro and nothing else.

Version pinning

Both images use latest. Neither project publishes a moving major tag: grafana/alloy ships only exact v1.x.y tags, and tecnativa/docker-socket-proxy only exact v0.x.y tags (its bare 0 tag is a stale leftover). latest is the closest equivalent, so new releases arrive on the next redeploy.