mirror of
https://github.com/tiennm99/composes.git
synced 2026-10-11 03:13:16 +00:00
148 lines
7.5 KiB
Markdown
148 lines
7.5 KiB
Markdown
# alloy
|
|
|
|
[Grafana Alloy](https://grafana.com/docs/alloy/latest/) shipping host and
|
|
container telemetry to Grafana Cloud, with remote config from Grafana Fleet
|
|
Management.
|
|
|
|
One container runs both the `node_exporter` (host) and `cadvisor` (container)
|
|
collectors. A second, tiny container proxies a read-only slice of the Docker
|
|
API to it. The Alloy config is embedded inline via Compose `configs:`, so
|
|
there is no `config.alloy` on disk. There is no `.env.example` either: the
|
|
nine variables are exported before `docker compose up`.
|
|
|
|
Both containers have fixed names,
|
|
`container_name: alloy` and `alloy-dockerproxy`, and `dockerproxy` publishes
|
|
`127.0.0.1:2375` — Alloy runs with host networking, so it has no compose
|
|
network to reach the proxy over.
|
|
|
|
## What it collects
|
|
|
|
| Source | Component | Notes |
|
|
|---|---|---|
|
|
| Host metrics | `prometheus.exporter.unix` | CPU, memory, load, disk I/O, filesystem, network, uname, boot time, systemd, vmstat, sockstat — the default collector set minus `ipvs/btrfs/infiniband/xfs/zfs` |
|
|
| Container metrics | `prometheus.exporter.cadvisor` | CPU, memory, fs usage/limit, network, `last_seen` |
|
|
| Container logs | `loki.source.docker` | All running containers, labeled `container`, `stream`, `instance` |
|
|
| Journal logs | `loki.source.journal` | systemd journal, labeled `unit`, `boot_id`, `transport`, `level` |
|
|
| File logs | `loki.source.file` | `/var/log/syslog`, `/var/log/messages`, `/var/log/*.log` |
|
|
| Remote config | `remotecfg` | Polls Grafana Fleet Management every 60s |
|
|
|
|
Metric filtering copies the `keep`-lists from the upstream Grafana Cloud
|
|
integrations verbatim ([Linux Node](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-linux-node/#metrics),
|
|
[Docker](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-docker/#metrics)).
|
|
Logs are unfiltered.
|
|
|
|
Where rsyslog mirrors journald into `/var/log/syslog` — the Debian and Ubuntu
|
|
default — the journal and file pipelines double-ship the same lines. Drop one
|
|
source on those hosts; the file-based one is the redundant one on systemd-only
|
|
stacks.
|
|
|
|
## Environment
|
|
|
|
All nine are required; `docker compose up` fails fast if any is unset.
|
|
|
|
| Variable | Purpose |
|
|
| --- | --- |
|
|
| `ALLOY_HOSTNAME` | Container hostname, and the Loki/Prometheus `instance` label |
|
|
| `REMOTECFG_URL` | Fleet Management endpoint |
|
|
| `REMOTECFG_ID` | Fleet Management agent id |
|
|
| `REMOTECFG_USER` | Fleet Management user id |
|
|
| `PROM_URL` / `PROM_USER` | Prometheus remote-write endpoint and user id |
|
|
| `LOKI_URL` / `LOKI_USER` | Loki push endpoint and user id |
|
|
| `GRAFANA_TOKEN` | One Cloud Access Policy token, scopes `metrics:write` + `logs:write` + `fleet-management:read` |
|
|
|
|
Find the values under Grafana Cloud → your stack → **Details** on each data
|
|
source, and under Fleet Management. The same token serves `remotecfg`,
|
|
Prometheus and Loki basic-auth.
|
|
|
|
```bash
|
|
export ALLOY_HOSTNAME=example-host REMOTECFG_ID=example-host ...
|
|
docker compose up -d
|
|
```
|
|
|
|
Run the same file on every host, changing `ALLOY_HOSTNAME` and `REMOTECFG_ID`
|
|
per host. Filter in Grafana with `instance=~"..."`.
|
|
|
|
## Privileges
|
|
|
|
The upstream Grafana Cloud docker integration runs `privileged: true` with the
|
|
raw Docker socket mounted read-write. Nothing in this config needs that, and
|
|
the combination is host-root-equivalent: `privileged` grants every capability
|
|
and unmasks `/proc` and `/sys`, and a container that can talk to the Docker
|
|
socket can start another container that mounts `/` writable. The socket's `:ro`
|
|
flag does not help — it stops the socket *file* being replaced, not the API
|
|
being used.
|
|
|
|
What it runs instead:
|
|
|
|
| Setting | Reason |
|
|
| --- | --- |
|
|
| `cap_drop: [ALL]` + `cap_add: [DAC_OVERRIDE]` | The image's entrypoint runs as uid 0 and reads host files owned by other users — the journal, paths under `/rootfs`, `/var/log`. Dropping every capability leaves it unable to open them, and unable to create its own storage directory. `DAC_OVERRIDE` restores exactly that and nothing else; `SYS_ADMIN`, `NET_ADMIN`, `SYS_PTRACE`, `MKNOD` and the rest stay dropped. |
|
|
| `no-new-privileges:true` | No setuid binary in the image can regain what was dropped. |
|
|
| `mem_limit: 2g`, `pids_limit: 512` | Steady state is around 900 MB; the limit stops a leak taking the host down with it. |
|
|
| `dockerproxy` instead of `/var/run/docker.sock` | See below. |
|
|
|
|
`read_only: true` is deliberately absent. Compose materialises an inline
|
|
`configs:` entry by writing it into the container, and refuses to do that on a
|
|
read-only service — `cannot create config ... : \`file\` is the sole supported
|
|
option`. Keeping the config inline is worth more than the read-only rootfs
|
|
here; adding it back means moving the config to a `config.alloy` file on disk
|
|
and switching the `configs:` entry to `file:`.
|
|
|
|
`network_mode: host` stays. `/proc/net` is a symlink to `/proc/self/net` and
|
|
resolves against the reading process's network namespace, so bind-mounting the
|
|
host's `/proc` to `/rootproc` is not enough — without host networking the
|
|
netdev, netstat, sockstat and conntrack families would describe the container's
|
|
veth pair instead of `eth0`. That is roughly 45 of the 157 kept metrics.
|
|
|
|
`/:/rootfs:ro` also stays, and is the widest remaining exposure: the container
|
|
can read every file on the host. The filesystem collector needs to `statfs()`
|
|
each mount point, and a bind mount cannot grant that without granting reads.
|
|
Dropping it would cost disk-space monitoring. Treat the container as
|
|
secret-bearing — it holds `GRAFANA_TOKEN` regardless.
|
|
|
|
## Docker API access
|
|
|
|
`prometheus.exporter.cadvisor` and `loki.source.docker` both need the Docker
|
|
API, so it cannot simply be removed. `dockerproxy` runs
|
|
[tecnativa/docker-socket-proxy](https://github.com/Tecnativa/docker-socket-proxy)
|
|
with the socket mounted read-only and publishes it on `127.0.0.1:2375`, which
|
|
Alloy reaches over host networking.
|
|
|
|
`POST` is revoked by default in that image, which is the point: container
|
|
create, `exec`, start and kill are all refused with 403, so a compromise of
|
|
Alloy can no longer become root on the host. The granted sections are the ones
|
|
the two components actually call:
|
|
|
|
| Variable | Called by |
|
|
| --- | --- |
|
|
| `CONTAINERS` | `discovery.docker` listing, `loki.source.docker` inspect and log read, cadvisor metadata |
|
|
| `NETWORKS` | `discovery.docker` — Prometheus' Docker SD resolves network names per container, and returns **zero targets** without it |
|
|
| `IMAGES`, `INFO`, `VERSION` | cadvisor |
|
|
| `EVENTS` | cadvisor container watch (granted by default) |
|
|
|
|
What this does not do: `GET /containers/{id}/json` still returns every
|
|
container's environment variables. The proxy closes the escalation path, not
|
|
the disclosure one.
|
|
|
|
## Mounts
|
|
|
|
| Mount | Why |
|
|
|---|---|
|
|
| `/proc:/rootproc:ro` | node-exporter cpu/mem/load, via `procfs_path` |
|
|
| `/sys:/sys:ro` | node-exporter and cadvisor cgroups |
|
|
| `/:/rootfs:ro` | filesystem collector, via `rootfs_path` |
|
|
| `/dev/disk/:/dev/disk:ro` | node-exporter diskstats device labels |
|
|
| `/var/lib/docker:ro` | cadvisor container metadata |
|
|
| `/var/log:/var/log:ro` | `loki.source.journal` and `loki.source.file` |
|
|
| `/etc/machine-id:ro` | Stable host id for the journal reader |
|
|
| `alloy-data` | WAL and remotecfg cache — the only writable path |
|
|
|
|
`dockerproxy` mounts `/var/run/docker.sock:ro` and nothing else.
|
|
|
|
## Notes
|
|
|
|
- [Upstream sources of truth](docs/upstream-sources-of-truth.md) — what this
|
|
follows, what is in scope, how to audit dashboard metric needs.
|
|
- [Coolify SSH session noise](docs/known-noise-coolify-ssh-sessions.md) — only
|
|
relevant on Coolify-managed hosts.
|