Files
composes/alloy/README.md
T

148 lines
7.5 KiB
Markdown

# alloy
[Grafana Alloy](https://grafana.com/docs/alloy/latest/) shipping host and
container telemetry to Grafana Cloud, with remote config from Grafana Fleet
Management.
One container runs both the `node_exporter` (host) and `cadvisor` (container)
collectors. A second, tiny container proxies a read-only slice of the Docker
API to it. The Alloy config is embedded inline via Compose `configs:`, so
there is no `config.alloy` on disk. There is no `.env.example` either: the
nine variables are exported before `docker compose up`.
Both containers have fixed names,
`container_name: alloy` and `alloy-dockerproxy`, and `dockerproxy` publishes
`127.0.0.1:2375` — Alloy runs with host networking, so it has no compose
network to reach the proxy over.
## What it collects
| Source | Component | Notes |
|---|---|---|
| Host metrics | `prometheus.exporter.unix` | CPU, memory, load, disk I/O, filesystem, network, uname, boot time, systemd, vmstat, sockstat — the default collector set minus `ipvs/btrfs/infiniband/xfs/zfs` |
| Container metrics | `prometheus.exporter.cadvisor` | CPU, memory, fs usage/limit, network, `last_seen` |
| Container logs | `loki.source.docker` | All running containers, labeled `container`, `stream`, `instance` |
| Journal logs | `loki.source.journal` | systemd journal, labeled `unit`, `boot_id`, `transport`, `level` |
| File logs | `loki.source.file` | `/var/log/syslog`, `/var/log/messages`, `/var/log/*.log` |
| Remote config | `remotecfg` | Polls Grafana Fleet Management every 60s |
Metric filtering copies the `keep`-lists from the upstream Grafana Cloud
integrations verbatim ([Linux Node](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-linux-node/#metrics),
[Docker](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-docker/#metrics)).
Logs are unfiltered.
Where rsyslog mirrors journald into `/var/log/syslog` — the Debian and Ubuntu
default — the journal and file pipelines double-ship the same lines. Drop one
source on those hosts; the file-based one is the redundant one on systemd-only
stacks.
## Environment
All nine are required; `docker compose up` fails fast if any is unset.
| Variable | Purpose |
| --- | --- |
| `ALLOY_HOSTNAME` | Container hostname, and the Loki/Prometheus `instance` label |
| `REMOTECFG_URL` | Fleet Management endpoint |
| `REMOTECFG_ID` | Fleet Management agent id |
| `REMOTECFG_USER` | Fleet Management user id |
| `PROM_URL` / `PROM_USER` | Prometheus remote-write endpoint and user id |
| `LOKI_URL` / `LOKI_USER` | Loki push endpoint and user id |
| `GRAFANA_TOKEN` | One Cloud Access Policy token, scopes `metrics:write` + `logs:write` + `fleet-management:read` |
Find the values under Grafana Cloud → your stack → **Details** on each data
source, and under Fleet Management. The same token serves `remotecfg`,
Prometheus and Loki basic-auth.
```bash
export ALLOY_HOSTNAME=example-host REMOTECFG_ID=example-host ...
docker compose up -d
```
Run the same file on every host, changing `ALLOY_HOSTNAME` and `REMOTECFG_ID`
per host. Filter in Grafana with `instance=~"..."`.
## Privileges
The upstream Grafana Cloud docker integration runs `privileged: true` with the
raw Docker socket mounted read-write. Nothing in this config needs that, and
the combination is host-root-equivalent: `privileged` grants every capability
and unmasks `/proc` and `/sys`, and a container that can talk to the Docker
socket can start another container that mounts `/` writable. The socket's `:ro`
flag does not help — it stops the socket *file* being replaced, not the API
being used.
What it runs instead:
| Setting | Reason |
| --- | --- |
| `cap_drop: [ALL]` + `cap_add: [DAC_OVERRIDE]` | The image's entrypoint runs as uid 0 and reads host files owned by other users — the journal, paths under `/rootfs`, `/var/log`. Dropping every capability leaves it unable to open them, and unable to create its own storage directory. `DAC_OVERRIDE` restores exactly that and nothing else; `SYS_ADMIN`, `NET_ADMIN`, `SYS_PTRACE`, `MKNOD` and the rest stay dropped. |
| `no-new-privileges:true` | No setuid binary in the image can regain what was dropped. |
| `mem_limit: 2g`, `pids_limit: 512` | Steady state is around 900 MB; the limit stops a leak taking the host down with it. |
| `dockerproxy` instead of `/var/run/docker.sock` | See below. |
`read_only: true` is deliberately absent. Compose materialises an inline
`configs:` entry by writing it into the container, and refuses to do that on a
read-only service — `cannot create config ... : \`file\` is the sole supported
option`. Keeping the config inline is worth more than the read-only rootfs
here; adding it back means moving the config to a `config.alloy` file on disk
and switching the `configs:` entry to `file:`.
`network_mode: host` stays. `/proc/net` is a symlink to `/proc/self/net` and
resolves against the reading process's network namespace, so bind-mounting the
host's `/proc` to `/rootproc` is not enough — without host networking the
netdev, netstat, sockstat and conntrack families would describe the container's
veth pair instead of `eth0`. That is roughly 45 of the 157 kept metrics.
`/:/rootfs:ro` also stays, and is the widest remaining exposure: the container
can read every file on the host. The filesystem collector needs to `statfs()`
each mount point, and a bind mount cannot grant that without granting reads.
Dropping it would cost disk-space monitoring. Treat the container as
secret-bearing — it holds `GRAFANA_TOKEN` regardless.
## Docker API access
`prometheus.exporter.cadvisor` and `loki.source.docker` both need the Docker
API, so it cannot simply be removed. `dockerproxy` runs
[tecnativa/docker-socket-proxy](https://github.com/Tecnativa/docker-socket-proxy)
with the socket mounted read-only and publishes it on `127.0.0.1:2375`, which
Alloy reaches over host networking.
`POST` is revoked by default in that image, which is the point: container
create, `exec`, start and kill are all refused with 403, so a compromise of
Alloy can no longer become root on the host. The granted sections are the ones
the two components actually call:
| Variable | Called by |
| --- | --- |
| `CONTAINERS` | `discovery.docker` listing, `loki.source.docker` inspect and log read, cadvisor metadata |
| `NETWORKS` | `discovery.docker` — Prometheus' Docker SD resolves network names per container, and returns **zero targets** without it |
| `IMAGES`, `INFO`, `VERSION` | cadvisor |
| `EVENTS` | cadvisor container watch (granted by default) |
What this does not do: `GET /containers/{id}/json` still returns every
container's environment variables. The proxy closes the escalation path, not
the disclosure one.
## Mounts
| Mount | Why |
|---|---|
| `/proc:/rootproc:ro` | node-exporter cpu/mem/load, via `procfs_path` |
| `/sys:/sys:ro` | node-exporter and cadvisor cgroups |
| `/:/rootfs:ro` | filesystem collector, via `rootfs_path` |
| `/dev/disk/:/dev/disk:ro` | node-exporter diskstats device labels |
| `/var/lib/docker:ro` | cadvisor container metadata |
| `/var/log:/var/log:ro` | `loki.source.journal` and `loki.source.file` |
| `/etc/machine-id:ro` | Stable host id for the journal reader |
| `alloy-data` | WAL and remotecfg cache — the only writable path |
`dockerproxy` mounts `/var/run/docker.sock:ro` and nothing else.
## Notes
- [Upstream sources of truth](docs/upstream-sources-of-truth.md) — what this
follows, what is in scope, how to audit dashboard metric needs.
- [Coolify SSH session noise](docs/known-noise-coolify-ssh-sessions.md) — only
relevant on Coolify-managed hosts.