docs: move service issue notes from READMEs into docs/<service>

Service READMEs now cover only what the service is and how to deploy it.
Known issues, log noise and troubleshooting move to docs/<service>/, named
after the service directory, so editing them never redeploys the service.
Drop alloy's validate workflow, which never ran from a subdirectory.
This commit is contained in:
tiennm99 committed 2026-10-04 09:46:08 +07:00
1 parent b6952d84df
commit 198329c008
13 files changed
+78 -126

No files matched your search

@@ -0,0 +1,18 @@
# Duplicate journal and syslog lines
Applies to `alloy/compose.yml`.
## Symptom
The same log line arrives in Loki twice: once from `loki.source.journal`, once
from `loki.source.file`.
## Cause
Where rsyslog mirrors journald into `/var/log/syslog` — the Debian and Ubuntu
default — the journal and file pipelines both ship it.
## Fix
Drop one source on those hosts. On systemd-only stacks the file-based one is
the redundant one.
@@ -0,0 +1,103 @@
# Known Noise: Coolify SSH Session Spam
> **Skip this doc** if you don't use [Coolify](https://coolify.io/) to manage the host. This is a special case for our setup, not a general issue.
## Symptom
Grafana Loki shows a constant stream of `info`-level lines on the **Linux Node** dashboard:
```
Session closed
New session NNNNN of user root.
```
Labels: `instance=<host>`, `job=integrations/node_exporter`, `level=info`, `unit=systemd-logind.service` (or `session-NNN.scope`).
Rate observed on a Coolify-managed host: **~300 sessions/hour ≈ 5/min**, in bursts of 2–3 within the same second.
## Cause
The Coolify host SSHes into each managed server every minute to check reachability and collect Docker container state. Each SSH connect → PAM `session opened` + `session closed` → systemd-logind writes it to the journal → `loki.source.journal` ships it to Grafana Cloud.
The **misleading part**: the label `job=integrations/node_exporter` makes it look like node_exporter emitted the log. It didn't. node_exporter only produces metrics — the journal-log pipeline reuses that label so the logs land on the same dashboard.
## Exact Coolify call flow (verified against source)
`coollabsio/coolify` v4.x — `app/Console/Kernel.php` and `app/Jobs/ServerManagerJob.php`:
1. **`ServerManagerJob` runs every minute** (self-hosted) or every 5 minutes (Cloud).
2. Per managed server, per cycle, it dispatches:
| Job | What it does over SSH | Skip condition |
|---|---|---|
| `ServerConnectionCheckJob` | Opens SSH, runs reachability probe | `isSentinelEnabled() && isSentinelLive()` |
| `ServerCheckJob` | Opens SSH, runs `docker container ls --format json` + proxy/log-drain checks | Sentinel "in sync" — last push within `sentinel_push_interval_seconds × 3` (min 120s) |
| `ServerStorageCheckJob` | Opens SSH, reads filesystem usage | Daily cron only (`0 23 * * *`), and only when Sentinel out of sync |
| `ServerPatchCheckJob` | Opens SSH, checks patch info | Weekly only (`0 0 * * 0`) |
| `CheckAndStartSentinelJob` | Opens SSH to (re)start the Sentinel container | Daily only |
= **roughly 2–3 SSH connects per minute per server** when Sentinel is OFF (matches the 5/min × bursts pattern in our journal).
3. The only built-in throttle is `shouldSkipDueToBackoff` — it backs off to every 3 / 6 / 12 minutes only **after the server has been marked unreachable** several times. There is **no UI setting to slow checks on a healthy server**.
## How to confirm on your host
```bash
# Top sources of SSH sessions in the last hour
sudo journalctl _COMM=sshd --since "1 hour ago" \
| grep "Accepted" | grep -oE "from [0-9.]+" | sort | uniq -c | sort -rn | head
# Match the SSH key fingerprint to its owner — replace fingerprint
while read -r line; do
fp=$(echo "$line" | ssh-keygen -lf - 2>/dev/null | awk '{print $2}')
[[ "$fp" == "SHA256:<paste-fingerprint-here>" ]] && echo "MATCH: $line"
done < /root/.ssh/authorized_keys
```
If the matching key's comment is `coolify` → this doc applies.
## Fixes (pick one)
### Option A — drop the noise at Alloy (host stays as-is)
Edit the `loki.process "default"` block inside `journal_module` in `alloy/compose.yml`:
```alloy
loki.process "default" {
forward_to = argument.forward_to.value
stage.drop {
expression = "(session opened|session closed|New session|Removed session) .*"
}
}
```
Restart: `docker compose up -d --force-recreate alloy`.
Trade-off: also loses visibility of legitimate human SSH logins. Narrow to `unit="systemd-logind.service"` or to a specific source IP if you want to keep auditing real users.
### Option B — eliminate the SSH polling at Coolify (recommended root-cause fix)
Enable **Sentinel** on the managed server. Sentinel is a small `coolify-sentinel` container that *pushes* metrics to the Coolify API. When it's healthy, `ServerManagerJob` **skips both `ServerConnectionCheckJob` and `ServerCheckJob`** for that server — SSH polling stops.
**Steps in the Coolify UI:**
1. Go to **Servers → `<your server>` → Configurations → General**.
2. Toggle **Sentinel** on. Optionally toggle **Metrics** in the same section if you want CPU/mem/disk pushed too.
3. Save. Coolify will deploy the `coolify-sentinel` container on the target server.
4. Wait ~2 minutes. Verify on the server: `docker ps | grep coolify-sentinel`.
5. Re-check the journal — SSH session rate should drop to occasional (daily Sentinel restart, weekly patch check, on-demand deploys), not ~5/min.
**Tunable:** `sentinel_push_interval_seconds` in the server's settings controls push cadence and the SSH-skip window (skip if last push within `× 3`, min 120s). Lower = fresher metrics, slightly more API traffic from Sentinel.
**Caveats:** Sentinel is flagged "experimental" in the Coolify docs. Metrics collection is **not** available for Docker Compose / Service-Template-based deployments — but Sentinel itself (connectivity + container status) still works.
## Why we don't ship a filter by default
The base setup is meant to be a generic Grafana Cloud Linux/Docker integration. Coolify-specific filtering belongs in a host-specific overlay, not in the shared compose file.
## References
- Coolify Sentinel docs: <https://coolify.io/docs/knowledge-base/server/sentinel>
- DeepWiki — Server Monitoring: <https://deepwiki.com/coollabsio/coolify/3.5-server-monitoring-(sentinel-and-metrics)>
- Source: [`app/Console/Kernel.php`](https://github.com/coollabsio/coolify/blob/v4.x/app/Console/Kernel.php), [`app/Jobs/ServerManagerJob.php`](https://github.com/coollabsio/coolify/blob/v4.x/app/Jobs/ServerManagerJob.php), [`app/Jobs/ServerCheckJob.php`](https://github.com/coollabsio/coolify/blob/v4.x/app/Jobs/ServerCheckJob.php)
+102
View File
@@ -0,0 +1,102 @@
# Upstream Sources of Truth
## Hard rule: official sources only
> **Only use official sources** when deciding what metrics, labels, log pipelines, dashboards, or rules this repo ships. **Never** consult or copy from third-party repos, blog posts, community Gists, or personal exports — even if the names match.
>
> If an answer can't be found in an official source, document the gap rather than fill it with unofficial data. **Non-negotiable.**
### What counts as an official source
| Tier | Source | Use for |
|---|---|---|
| 1 | `https://grafana.com/docs/grafana-cloud/...` (Grafana Cloud integration reference pages) | Alloy snippet, scrape config, label conventions, drop/keep rules |
| 2 | `github.com/grafana/...` (Grafana Labs org repos) | Mixin libsonnet/jsonnet, `dashboards_out/*.json`, alerting rules |
| 3 | `github.com/prometheus/...` (Prometheus org repos) | Upstream node_exporter mixin, prom-mixin libraries |
| 4 | The Grafana Cloud stack's own running dashboards (via authenticated API) | Authoritative answer for "what does my live integration use" |
### What doesn't count (do not use)
- Personal forks, Terraform exports, ops-team copies of integration dashboards.
- `grafana.com/grafana/dashboards/<id>` community submissions, unless explicitly published by the `grafana` user.
- AI summaries of dashboards.
- Stale snapshots in unrelated repos that happen to contain `job=integrations/...` strings.
## Principle
Alloy here ships **only what the integration dashboards need** — not the full firehose of every metric/log line node_exporter and cadvisor can produce. The goal is dashboards that work out-of-the-box on Grafana Cloud, without paying for unused active series.
When the upstream integration changes (adds a metric, drops a panel, renames a label), this repo follows. Local additions outside that scope go into a separate, named, documented overlay.
## Upstream references
The two reference pages `alloy/compose.yml` mirrors:
- **Linux Node integration** — <https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-linux-node/>
- **Docker integration** — <https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-docker/>
Where official mixin dashboards exist they're tier-2 corroboration:
- Docker mixin (Grafana Labs): <https://github.com/grafana/jsonnet-libs/blob/master/docker-mixin/dashboards_out/docker.json>
- Node-exporter mixin (Prometheus project): <https://github.com/prometheus/node_exporter/tree/master/docs/node-mixin>
## Mapping our config to upstream
| Block in `alloy/compose.yml` | Upstream source |
|---|---|
| `prometheus.exporter.unix` (collectors, mounts, fs/net excludes) | Linux Node integration page → "Configure Alloy" |
| `prometheus.relabel "integrations_node_exporter"` (`keep` allowlist of 157 metrics) | Linux Node integration page → [Metrics](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-linux-node/#metrics) section, verbatim |
| `loki.source.journal "default"` + relabel rules (`unit`, `boot_id`, `transport`, `level`) | Linux Node integration page → log scraping |
| `loki.source.file` for `/var/log/{syslog,messages,*.log}` | Linux Node integration page → file log scraping |
| `prometheus.exporter.cadvisor` (`docker_only = true`) | Docker integration page |
| `prometheus.relabel "integrations_cadvisor"` (`keep` allowlist of 16 metrics) | Docker integration page → [Metrics](https://grafana.com/docs/grafana-cloud/monitor-infrastructure/integrations/integration-reference/integration-docker/#metrics) section, verbatim |
| `discovery.docker` + `loki.source.docker` (job/instance/container/stream) | Docker integration page → log scraping |
## What "follow upstream" means in practice
1. **Don't expand the metric set unilaterally.** If a panel in a Grafana Cloud dashboard requires a metric we don't ship, we'd add it — but the trigger is "the integration dashboard needs it, verified from a tier 1–4 source", not "node_exporter exposes it".
2. **Don't filter further than upstream does.** We ship at least what the upstream config does. Tightening (e.g. for cost) goes in a clearly-named overlay or a downstream env-specific config.
3. **Re-check on Alloy/integration major versions.** When bumping `grafana/alloy` image or when the Grafana Cloud integration revs, diff the upstream Alloy snippet against `alloy/compose.yml` and update.
## Audit (2026-04-26)
Both keep-lists in `alloy/compose.yml` are copied verbatim from the **Metrics** section of each integration page (tier 1):
- **Linux-Node** — 157 raw metrics (`node_*`, `process_max_fds`, `process_open_fds`, `up`). The list also contains `instance:node_num_cpu:sum`, which is a recording-rule output computed server-side by Grafana Cloud's ruler — it's intentionally **not** in the keep-list because the agent doesn't produce it.
- **Docker** — 16 metrics (`container_*`, `machine_memory_bytes`, `machine_scrape_error`, `up`).
Re-verification on 2026-04-26 also confirmed:
- **Alloy image** bumped `v1.10.0` → `v1.16.1` (latest stable; v1.16.0 released 2026-04-23, patched to v1.16.1 on 2026-05-05).
- **`loki.source.journal` `path`** was previously pinned to `/var/log/journal`; upstream omits the field, letting Alloy default to **both** `/var/log/journal` (persistent) and `/run/log/journal` (volatile). Local now matches — `path` removed.
The Grafana Cloud integration's full dashboard set (the 7 Linux-Node + 2 Docker dashboards) is not publicly hosted. Tier-4 verification (against the live stack via authenticated API) was **not** performed and is the only known gap.
### Re-running the audit
```bash
# Tier 2 — Grafana Labs docker-mixin
curl -sL https://raw.githubusercontent.com/grafana/jsonnet-libs/master/docker-mixin/dashboards_out/docker.json \
| jq -r '.. | .expr? // empty' \
| grep -oE 'container_[a-zA-Z0-9_]+|machine_[a-zA-Z0-9_]+' | sort -u
# Tier 3 — Prometheus node-exporter mixin (libsonnet, grep for metric names)
curl -sL https://raw.githubusercontent.com/prometheus/node_exporter/master/docs/node-mixin/lib/prom-mixin.libsonnet \
| grep -oE 'node_[a-zA-Z0-9_]+|process_[a-zA-Z0-9_]+' | sort -u
# Tier 4 — your own Grafana Cloud stack
GRAFANA_TOKEN=... # service-account token, dashboards:read
STACK=<your-stack>.grafana.net
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"https://$STACK/api/search?query=Linux+Node&type=dash-db" \
| jq -r '.[].uid' | while read uid; do
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"https://$STACK/api/dashboards/uid/$uid" > "$uid.json"
done
# then jq + grep as above
```
## Out of scope
- Custom dashboards or metrics for app-level monitoring (deploy a separate Alloy config for that).
- Host-specific filtering (e.g. dropping noisy log lines from a particular daemon — see `docs/known-noise-coolify-ssh-sessions.md` for an example).
+18
View File
@@ -0,0 +1,18 @@
# gitea-mirror troubleshooting
Applies to `gitea-mirror/compose.yml`.
## Gitea cannot connect after changing `POSTGRES_PASSWORD`
Postgres sets the password only when it first initialises `db-data`. Changing
`POSTGRES_PASSWORD` later breaks Gitea's connection until the role is altered
to match:
```sh
docker compose exec db psql -U gitea -c "ALTER USER gitea PASSWORD '<new>';"
```
## A redeploy kills clones in progress
Gitea restarts and the clone dies with it. Avoid pushing to `gitea-mirror/`
while a large first mirror runs.
+11
View File
@@ -0,0 +1,11 @@
# Known noise: `xdg-open` stack trace on start
Applies to `opencode/compose.yml`.
The container logs a Bun stack trace ending in
`Executable not found in $PATH: "xdg-open"` on every start. `opencode web`
tries to open the UI in a local browser; there isn't one. It is noise — the
server is already listening by then, and the container keeps running.
Installing `xdg-utils` would silence it at the cost of pulling X11 in and
tripling the image, which is not worth it for a log line.
+9
View File
@@ -0,0 +1,9 @@
# paseo troubleshooting
Applies to `paseo/compose.yml`.
## The pairing screen stays on `localhost:6767`
After entering the address with its port, the UI may still show
`localhost:6767`: the old entry is cached in `localStorage`. Clear the site
data for the domain and enter it again.