mirror of
https://github.com/tiennm99/composes.git
synced 2026-10-11 12:09:25 +00:00
docs: move service issue notes from READMEs into docs/<service>
Service READMEs now cover only what the service is and how to deploy it. Known issues, log noise and troubleshooting move to docs/<service>/, named after the service directory, so editing them never redeploys the service. Drop alloy's validate workflow, which never ran from a subdirectory.
This commit is contained in:
1 parent
b6952d84df
commit
198329c008
13 files changed
+78
-126
No files matched your search
@@ -0,0 +1,18 @@
|
||||
# Duplicate journal and syslog lines
|
||||
|
||||
Applies to `alloy/compose.yml`.
|
||||
|
||||
## Symptom
|
||||
|
||||
The same log line arrives in Loki twice: once from `loki.source.journal`, once
|
||||
from `loki.source.file`.
|
||||
|
||||
## Cause
|
||||
|
||||
Where rsyslog mirrors journald into `/var/log/syslog` — the Debian and Ubuntu
|
||||
default — the journal and file pipelines both ship it.
|
||||
|
||||
## Fix
|
||||
|
||||
Drop one source on those hosts. On systemd-only stacks the file-based one is
|
||||
the redundant one.
|
||||
@@ -0,0 +1,103 @@
|
||||
# Known Noise: Coolify SSH Session Spam
|
||||
|
||||
> **Skip this doc** if you don't use [Coolify](https://coolify.io/) to manage the host. This is a special case for our setup, not a general issue.
|
||||
|
||||
## Symptom
|
||||
|
||||
Grafana Loki shows a constant stream of `info`-level lines on the **Linux Node** dashboard:
|
||||
|
||||
```
|
||||
Session closed
|
||||
New session NNNNN of user root.
|
||||
```
|
||||
|
||||
Labels: `instance=<host>`, `job=integrations/node_exporter`, `level=info`, `unit=systemd-logind.service` (or `session-NNN.scope`).
|
||||
|
||||
Rate observed on a Coolify-managed host: **~300 sessions/hour ≈ 5/min**, in bursts of 2–3 within the same second.
|
||||
|
||||
## Cause
|
||||
|
||||
The Coolify host SSHes into each managed server every minute to check reachability and collect Docker container state. Each SSH connect → PAM `session opened` + `session closed` → systemd-logind writes it to the journal → `loki.source.journal` ships it to Grafana Cloud.
|
||||
|
||||
The **misleading part**: the label `job=integrations/node_exporter` makes it look like node_exporter emitted the log. It didn't. node_exporter only produces metrics — the journal-log pipeline reuses that label so the logs land on the same dashboard.
|
||||
|
||||
## Exact Coolify call flow (verified against source)
|
||||
|
||||
`coollabsio/coolify` v4.x — `app/Console/Kernel.php` and `app/Jobs/ServerManagerJob.php`:
|
||||
|
||||
1. **`ServerManagerJob` runs every minute** (self-hosted) or every 5 minutes (Cloud).
|
||||
2. Per managed server, per cycle, it dispatches:
|
||||
|
||||
| Job | What it does over SSH | Skip condition |
|
||||
|---|---|---|
|
||||
| `ServerConnectionCheckJob` | Opens SSH, runs reachability probe | `isSentinelEnabled() && isSentinelLive()` |
|
||||
| `ServerCheckJob` | Opens SSH, runs `docker container ls --format json` + proxy/log-drain checks | Sentinel "in sync" — last push within `sentinel_push_interval_seconds × 3` (min 120s) |
|
||||
| `ServerStorageCheckJob` | Opens SSH, reads filesystem usage | Daily cron only (`0 23 * * *`), and only when Sentinel out of sync |
|
||||
| `ServerPatchCheckJob` | Opens SSH, checks patch info | Weekly only (`0 0 * * 0`) |
|
||||
| `CheckAndStartSentinelJob` | Opens SSH to (re)start the Sentinel container | Daily only |
|
||||
|
||||
= **roughly 2–3 SSH connects per minute per server** when Sentinel is OFF (matches the 5/min × bursts pattern in our journal).
|
||||
|
||||
3. The only built-in throttle is `shouldSkipDueToBackoff` — it backs off to every 3 / 6 / 12 minutes only **after the server has been marked unreachable** several times. There is **no UI setting to slow checks on a healthy server**.
|
||||
|
||||
## How to confirm on your host
|
||||
|
||||
```bash
|
||||
# Top sources of SSH sessions in the last hour
|
||||
sudo journalctl _COMM=sshd --since "1 hour ago" \
|
||||
| grep "Accepted" | grep -oE "from [0-9.]+" | sort | uniq -c | sort -rn | head
|
||||
|
||||
# Match the SSH key fingerprint to its owner — replace fingerprint
|
||||
while read -r line; do
|
||||
fp=$(echo "$line" | ssh-keygen -lf - 2>/dev/null | awk '{print $2}')
|
||||
[[ "$fp" == "SHA256:<paste-fingerprint-here>" ]] && echo "MATCH: $line"
|
||||
done < /root/.ssh/authorized_keys
|
||||
```
|
||||
|
||||
If the matching key's comment is `coolify` → this doc applies.
|
||||
|
||||
## Fixes (pick one)
|
||||
|
||||
### Option A — drop the noise at Alloy (host stays as-is)
|
||||
|
||||
Edit the `loki.process "default"` block inside `journal_module` in `alloy/compose.yml`:
|
||||
|
||||
```alloy
|
||||
loki.process "default" {
|
||||
forward_to = argument.forward_to.value
|
||||
|
||||
stage.drop {
|
||||
expression = "(session opened|session closed|New session|Removed session) .*"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Restart: `docker compose up -d --force-recreate alloy`.
|
||||
|
||||
Trade-off: also loses visibility of legitimate human SSH logins. Narrow to `unit="systemd-logind.service"` or to a specific source IP if you want to keep auditing real users.
|
||||
|
||||
### Option B — eliminate the SSH polling at Coolify (recommended root-cause fix)
|
||||
|
||||
Enable **Sentinel** on the managed server. Sentinel is a small `coolify-sentinel` container that *pushes* metrics to the Coolify API. When it's healthy, `ServerManagerJob` **skips both `ServerConnectionCheckJob` and `ServerCheckJob`** for that server — SSH polling stops.
|
||||
|
||||
**Steps in the Coolify UI:**
|
||||
|
||||
1. Go to **Servers → `<your server>` → Configurations → General**.
|
||||
2. Toggle **Sentinel** on. Optionally toggle **Metrics** in the same section if you want CPU/mem/disk pushed too.
|
||||
3. Save. Coolify will deploy the `coolify-sentinel` container on the target server.
|
||||
4. Wait ~2 minutes. Verify on the server: `docker ps | grep coolify-sentinel`.
|
||||
5. Re-check the journal — SSH session rate should drop to occasional (daily Sentinel restart, weekly patch check, on-demand deploys), not ~5/min.
|
||||
|
||||
**Tunable:** `sentinel_push_interval_seconds` in the server's settings controls push cadence and the SSH-skip window (skip if last push within `× 3`, min 120s). Lower = fresher metrics, slightly more API traffic from Sentinel.
|
||||
|
||||
**Caveats:** Sentinel is flagged "experimental" in the Coolify docs. Metrics collection is **not** available for Docker Compose / Service-Template-based deployments — but Sentinel itself (connectivity + container status) still works.
|
||||
|
||||
## Why we don't ship a filter by default
|
||||
|
||||
The base setup is meant to be a generic Grafana Cloud Linux/Docker integration. Coolify-specific filtering belongs in a host-specific overlay, not in the shared compose file.
|
||||
|
||||
## References
|
||||
|
||||
- Coolify Sentinel docs: <https://coolify.io/docs/knowledge-base/server/sentinel>
|
||||
- DeepWiki — Server Monitoring: <https://deepwiki.com/coollabsio/coolify/3.5-server-monitoring-(sentinel-and-metrics)>
|
||||
- Source: [`app/Console/Kernel.php`](https://github.com/coollabsio/coolify/blob/v4.x/app/Console/Kernel.php), [`app/Jobs/ServerManagerJob.php`](https://github.com/coollabsio/coolify/blob/v4.x/app/Jobs/ServerManagerJob.php), [`app/Jobs/ServerCheckJob.php`](https://github.com/coollabsio/coolify/blob/v4.x/app/Jobs/ServerCheckJob.php)
|
||||
@@ -0,0 +1,102 @@
|
||||
# Upstream Sources of Truth
|
||||
|
||||
## Hard rule: official sources only
|
||||
|
||||
> **Only use official sources** when deciding what metrics, labels, log pipelines, dashboards, or rules this repo ships. **Never** consult or copy from third-party repos, blog posts, community Gists, or personal exports — even if the names match.
|
||||
>
|
||||
> If an answer can't be found in an official source, document the gap rather than fill it with unofficial data. **Non-negotiable.**
|
||||
|
||||
### What counts as an official source
|
||||
|
||||
| Tier | Source | Use for |
|
||||
|---|---|---|
|
||||
| 1 | `https://grafana.com/docs/grafana-cloud/...` (Grafana Cloud integration reference pages) | Alloy snippet, scrape config, label conventions, drop/keep rules |
|
||||
| 2 | `github.com/grafana/...` (Grafana Labs org repos) | Mixin libsonnet/jsonnet, `dashboards_out/*.json`, alerting rules |
|
||||
| 3 | `github.com/prometheus/...` (Prometheus org repos) | Upstream node_exporter mixin, prom-mixin libraries |
|
||||
| 4 | The Grafana Cloud stack's own running dashboards (via authenticated API) | Authoritative answer for "what does my live integration use" |
|
||||
|
||||
### What doesn't count (do not use)
|
||||
|
||||
- Personal forks, Terraform exports, ops-team copies of integration dashboards.
|
||||
- `grafana.com/grafana/dashboards/<id>` community submissions, unless explicitly published by the `grafana` user.
|
||||
- AI summaries of dashboards.
|
||||
- Stale snapshots in unrelated repos that happen to contain `job=integrations/...` strings.
|
||||
|
||||
## Principle
|
||||
|
||||
Alloy here ships **only what the integration dashboards need** — not the full firehose of every metric/log line node_exporter and cadvisor can produce. The goal is dashboards that work out-of-the-box on Grafana Cloud, without paying for unused active series.
|
||||
|
||||
When the upstream integration changes (adds a metric, drops a panel, renames a label), this repo follows. Local additions outside that scope go into a separate, named, documented overlay.
|
||||
|
||||
## Upstream references
|
||||
|
||||
The two reference pages `alloy/compose.yml` mirrors:
|
||||
|
||||
- **Linux Node integration** — <https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-linux-node/>
|
||||
- **Docker integration** — <https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-docker/>
|
||||
|
||||
Where official mixin dashboards exist they're tier-2 corroboration:
|
||||
|
||||
- Docker mixin (Grafana Labs): <https://github.com/grafana/jsonnet-libs/blob/master/docker-mixin/dashboards_out/docker.json>
|
||||
- Node-exporter mixin (Prometheus project): <https://github.com/prometheus/node_exporter/tree/master/docs/node-mixin>
|
||||
|
||||
## Mapping our config to upstream
|
||||
|
||||
| Block in `alloy/compose.yml` | Upstream source |
|
||||
|---|---|
|
||||
| `prometheus.exporter.unix` (collectors, mounts, fs/net excludes) | Linux Node integration page → "Configure Alloy" |
|
||||
| `prometheus.relabel "integrations_node_exporter"` (`keep` allowlist of 157 metrics) | Linux Node integration page → [Metrics](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-linux-node/#metrics) section, verbatim |
|
||||
| `loki.source.journal "default"` + relabel rules (`unit`, `boot_id`, `transport`, `level`) | Linux Node integration page → log scraping |
|
||||
| `loki.source.file` for `/var/log/{syslog,messages,*.log}` | Linux Node integration page → file log scraping |
|
||||
| `prometheus.exporter.cadvisor` (`docker_only = true`) | Docker integration page |
|
||||
| `prometheus.relabel "integrations_cadvisor"` (`keep` allowlist of 16 metrics) | Docker integration page → [Metrics](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-docker/#metrics) section, verbatim |
|
||||
| `discovery.docker` + `loki.source.docker` (job/instance/container/stream) | Docker integration page → log scraping |
|
||||
|
||||
## What "follow upstream" means in practice
|
||||
|
||||
1. **Don't expand the metric set unilaterally.** If a panel in a Grafana Cloud dashboard requires a metric we don't ship, we'd add it — but the trigger is "the integration dashboard needs it, verified from a tier 1–4 source", not "node_exporter exposes it".
|
||||
2. **Don't filter further than upstream does.** We ship at least what the upstream config does. Tightening (e.g. for cost) goes in a clearly-named overlay or a downstream env-specific config.
|
||||
3. **Re-check on Alloy/integration major versions.** When bumping `grafana/alloy` image or when the Grafana Cloud integration revs, diff the upstream Alloy snippet against `alloy/compose.yml` and update.
|
||||
|
||||
## Audit (2026-04-26)
|
||||
|
||||
Both keep-lists in `alloy/compose.yml` are copied verbatim from the **Metrics** section of each integration page (tier 1):
|
||||
|
||||
- **Linux-Node** — 157 raw metrics (`node_*`, `process_max_fds`, `process_open_fds`, `up`). The list also contains `instance:node_num_cpu:sum`, which is a recording-rule output computed server-side by Grafana Cloud's ruler — it's intentionally **not** in the keep-list because the agent doesn't produce it.
|
||||
- **Docker** — 16 metrics (`container_*`, `machine_memory_bytes`, `machine_scrape_error`, `up`).
|
||||
|
||||
Re-verification on 2026-04-26 also confirmed:
|
||||
|
||||
- **Alloy image** bumped `v1.10.0` → `v1.16.1` (latest stable; v1.16.0 released 2026-04-23, patched to v1.16.1 on 2026-05-05).
|
||||
- **`loki.source.journal` `path`** was previously pinned to `/var/log/journal`; upstream omits the field, letting Alloy default to **both** `/var/log/journal` (persistent) and `/run/log/journal` (volatile). Local now matches — `path` removed.
|
||||
|
||||
The Grafana Cloud integration's full dashboard set (the 7 Linux-Node + 2 Docker dashboards) is not publicly hosted. Tier-4 verification (against the live stack via authenticated API) was **not** performed and is the only known gap.
|
||||
|
||||
### Re-running the audit
|
||||
|
||||
```bash
|
||||
# Tier 2 — Grafana Labs docker-mixin
|
||||
curl -sL https://raw.githubusercontent.com/grafana/jsonnet-libs/master/docker-mixin/dashboards_out/docker.json \
|
||||
| jq -r '.. | .expr? // empty' \
|
||||
| grep -oE 'container_[a-zA-Z0-9_]+|machine_[a-zA-Z0-9_]+' | sort -u
|
||||
|
||||
# Tier 3 — Prometheus node-exporter mixin (libsonnet, grep for metric names)
|
||||
curl -sL https://raw.githubusercontent.com/prometheus/node_exporter/master/docs/node-mixin/lib/prom-mixin.libsonnet \
|
||||
| grep -oE 'node_[a-zA-Z0-9_]+|process_[a-zA-Z0-9_]+' | sort -u
|
||||
|
||||
# Tier 4 — your own Grafana Cloud stack
|
||||
GRAFANA_TOKEN=... # service-account token, dashboards:read
|
||||
STACK=<your-stack>.grafana.net
|
||||
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
|
||||
"https://$STACK/api/search?query=Linux+Node&type=dash-db" \
|
||||
| jq -r '.[].uid' | while read uid; do
|
||||
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
|
||||
"https://$STACK/api/dashboards/uid/$uid" > "$uid.json"
|
||||
done
|
||||
# then jq + grep as above
|
||||
```
|
||||
|
||||
## Out of scope
|
||||
|
||||
- Custom dashboards or metrics for app-level monitoring (deploy a separate Alloy config for that).
|
||||
- Host-specific filtering (e.g. dropping noisy log lines from a particular daemon — see `docs/known-noise-coolify-ssh-sessions.md` for an example).
|
||||
@@ -0,0 +1,18 @@
|
||||
# gitea-mirror troubleshooting
|
||||
|
||||
Applies to `gitea-mirror/compose.yml`.
|
||||
|
||||
## Gitea cannot connect after changing `POSTGRES_PASSWORD`
|
||||
|
||||
Postgres sets the password only when it first initialises `db-data`. Changing
|
||||
`POSTGRES_PASSWORD` later breaks Gitea's connection until the role is altered
|
||||
to match:
|
||||
|
||||
```sh
|
||||
docker compose exec db psql -U gitea -c "ALTER USER gitea PASSWORD '<new>';"
|
||||
```
|
||||
|
||||
## A redeploy kills clones in progress
|
||||
|
||||
Gitea restarts and the clone dies with it. Avoid pushing to `gitea-mirror/`
|
||||
while a large first mirror runs.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Known noise: `xdg-open` stack trace on start
|
||||
|
||||
Applies to `opencode/compose.yml`.
|
||||
|
||||
The container logs a Bun stack trace ending in
|
||||
`Executable not found in $PATH: "xdg-open"` on every start. `opencode web`
|
||||
tries to open the UI in a local browser; there isn't one. It is noise — the
|
||||
server is already listening by then, and the container keeps running.
|
||||
|
||||
Installing `xdg-utils` would silence it at the cost of pulling X11 in and
|
||||
tripling the image, which is not worth it for a log line.
|
||||
@@ -0,0 +1,9 @@
|
||||
# paseo troubleshooting
|
||||
|
||||
Applies to `paseo/compose.yml`.
|
||||
|
||||
## The pairing screen stays on `localhost:6767`
|
||||
|
||||
After entering the address with its port, the UI may still show
|
||||
`localhost:6767`: the old entry is cached in `localStorage`. Clear the site
|
||||
data for the domain and enter it again.
|
||||
Reference in new issue
Block a user