From d889338306529fc9c5c75602498fe56818629563 Mon Sep 17 00:00:00 2001 From: tiennm99 Date: Tue, 6 Oct 2026 16:56:46 +0700 Subject: [PATCH] fix(alloy): drop docker log lines older than Loki accepts loki.source.docker resumes from a one-second position and Docker's since is inclusive, so each restart re-reads a container's last logged second. For idle containers those lines are older than 7 days and Grafana Cloud Loki rejects the batch with 400 timestamp too old. Drop them before loki.write; Loki would not have stored them anyway. --- alloy/README.md | 7 +++- alloy/compose.yml | 12 ++++++- docs/alloy/docker-logs-resent-on-restart.md | 37 +++++++++++++++++++++ docs/alloy/upstream-sources-of-truth.md | 1 + 4 files changed, 55 insertions(+), 2 deletions(-) create mode 100644 docs/alloy/docker-logs-resent-on-restart.md diff --git a/alloy/README.md b/alloy/README.md index 52008fa..e6648a4 100644 --- a/alloy/README.md +++ b/alloy/README.md @@ -27,7 +27,12 @@ network to reach the proxy over. Metric filtering copies the `keep`-lists from the upstream Grafana Cloud integrations verbatim ([Linux Node](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-linux-node/#metrics), [Docker](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-docker/#metrics)). -Logs are unfiltered. +Logs are unfiltered, except that container log lines older than 168h are +dropped before they are sent. Grafana Cloud Loki rejects anything older than +7 days, and `loki.source.docker` re-reads a container's last second of logs +on every restart (its saved position has one-second precision), so an idle +container's last lines would otherwise come back as a `400 timestamp too old` +error on every redeploy. Nothing Loki would have accepted is dropped. ## Environment diff --git a/alloy/compose.yml b/alloy/compose.yml index 435040d..867056a 100644 --- a/alloy/compose.yml +++ b/alloy/compose.yml @@ -276,7 +276,17 @@ configs: loki.source.docker "logs_integrations_docker" { host = "tcp://127.0.0.1:2375" targets = discovery.docker.logs_integrations_docker.targets - forward_to = [loki.write.grafana_cloud_loki.receiver] + forward_to = [loki.process.logs_integrations_docker.receiver] relabel_rules = discovery.relabel.logs_integrations_docker.rules refresh_interval = "5s" } + + // Drop container log lines older than Grafana Cloud Loki accepts. + loki.process "logs_integrations_docker" { + forward_to = [loki.write.grafana_cloud_loki.receiver] + + stage.drop { + older_than = "168h" + drop_counter_reason = "too_old" + } + } diff --git a/docs/alloy/docker-logs-resent-on-restart.md b/docs/alloy/docker-logs-resent-on-restart.md new file mode 100644 index 0000000..ff45d8f --- /dev/null +++ b/docs/alloy/docker-logs-resent-on-restart.md @@ -0,0 +1,37 @@ +# Docker logs re-sent on restart + +Applies to `alloy/compose.yml`. Checked against Alloy v1.20.1. + +## Symptom + +Every time Alloy restarts, `loki.write` logs one rejected batch: + +``` +final error sending batch, no retries left, dropping data ... status=400 +13 errors like: entry for stream '{container="coolify-realtime", ...}' +has timestamp too old: 2026-09-18T17:00:51Z, oldest acceptable timestamp is: ... +``` + +The same containers, the same timestamps and the same counts show up on every +restart. + +## Cause + +`loki.source.docker` saves each container's read position as a Unix timestamp +in **seconds** (`internal/component/loki/source/docker/tailer.go`, +`t.positions.Put(..., ts.Unix())`). On start it asks the Docker API for logs +`since` that second, and Docker includes the second itself. So every restart +re-reads every line written in the container's last logged second. + +For a busy container those lines are recent, and Loki silently discards them +as exact duplicates. For an idle container whose last line is more than 7 days +old, Grafana Cloud Loki rejects them with `timestamp too old` instead. +Positions are kept in the `alloy-data` volume, so they are not being lost. +Re-reading that last second is simply how the component works. + +## Fix + +`loki.process "logs_integrations_docker"` sits between `loki.source.docker` and +`loki.write` and drops lines older than `168h` with `stage.drop`. Those lines +would be rejected by Loki anyway, so nothing that would have been stored is +lost. The drops are counted under `loki_process_dropped_lines_total{reason="too_old"}`. diff --git a/docs/alloy/upstream-sources-of-truth.md b/docs/alloy/upstream-sources-of-truth.md index 61bb950..4e7b794 100644 --- a/docs/alloy/upstream-sources-of-truth.md +++ b/docs/alloy/upstream-sources-of-truth.md @@ -51,6 +51,7 @@ Where official mixin dashboards exist they're tier-2 corroboration: | `prometheus.exporter.cadvisor` (`docker_only = true`) | Docker integration page | | `prometheus.relabel "integrations_cadvisor"` (`keep` allowlist of 16 metrics) | Docker integration page → [Metrics](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-docker/#metrics) section, verbatim | | `discovery.docker` + `loki.source.docker` (job/instance/container/stream) | Docker integration page → log scraping | +| `loki.process "logs_integrations_docker"` (drop lines older than 168h) | Local addition, not upstream. Drops only what Loki rejects anyway; see [docker-logs-resent-on-restart.md](docker-logs-resent-on-restart.md) | ## What "follow upstream" means in practice