Commit Graph
102 Commits
Author SHA1 Message Date
tiennm99 21de313dec docs: add README and sample environment values
Document the compose services, published ports, volumes, and first-run
setup. Fill .env.example with sample values and note that compose.yml
does not yet consume them.
2026-07-25 18:45:18 +07:00
tiennm99 25aba88d0c chore: add gitea mirror docker compose setup
Compose stack for Gitea with Postgres and gitea-mirror, plus an
environment template listing the required configuration keys.
2026-07-25 18:08:56 +07:00
tiennm99 872d9ba5ea fix: bump alloy to v1.16.1 and align CI validator image 2026-05-31 10:41:57 +07:00
tiennm99 bad7082ac3 docs(readme): add customization section and related cluster links 2026-05-11 21:45:31 +07:00
tiennm99 2752b8fb5a docs: expand README — features table, GPU notes, smoke test 2026-05-11 20:44:53 +07:00
tiennm99 28cf2bd1f9 docs: flesh out README 2026-05-11 20:15:20 +07:00
tiennm99 6e79c6851d docs: add README 2026-05-11 17:04:17 +07:00
tiennm99 40d7b0e3fc fix: ship cadvisor working_set/rss metrics so memory panels reflect real usage
container_memory_usage_bytes counts page cache attributed to the cgroup,
which makes disk-heavy containers (e.g. gitea) appear to use ~all host RAM.
Add container_memory_working_set_bytes (the metric Grafana's Docker
integration dashboard expects), plus container_memory_rss,
container_memory_cache, and container_spec_memory_limit_bytes for
breakdown and limit-percentage panels.
2026-04-29 09:59:19 +07:00
tiennm99 2a24eaf10b fix: drop explicit journal path and bump alloy to v1.16.0
`loki.source.journal "default"` no longer pins `path = "/var/log/journal"`.
Upstream omits the field, which lets Alloy default to BOTH
`/var/log/journal` (persistent) and `/run/log/journal` (volatile). The
explicit path silently dropped journal logs on hosts with volatile-only
storage. Matches the canonical Linux Node integration template.

Also bumps the image six minor versions to current stable. Doc records
the 2026-04-26 re-audit.
2026-04-26 10:15:29 +07:00
tiennm99 9399a8b115 feat: keep-list verbatim from each integration's Metrics section
Copies the exact 157-metric list from the Linux Node integration's
Metrics anchor as the cadvisor keep-list already does for Docker
(16 metrics). Replaces the earlier `drop node_scrape_collector_.+`
rule, which was the integration page's alternate snippet but didn't
ship the explicit allowlist users see in the docs.

`instance:node_num_cpu:sum` from the Metrics section is intentionally
omitted — it's a recording-rule output computed server-side by
Grafana Cloud's ruler, not produced by the agent.

Doc + README updated to point at the Metrics anchors directly so the
source of truth is unambiguous.
2026-04-26 09:54:43 +07:00
tiennm99 c1db79a359 docs: codify upstream-sources-only rule for config decisions
Adds docs/upstream-sources-of-truth.md as the binding policy for what
this repo follows when deciding metrics, labels, log pipelines, and
dashboards to ship.

Hard rule: only tier 1-4 official sources (Grafana Cloud integration
docs, github.com/grafana/*, github.com/prometheus/*, the user's own
authenticated Grafana Cloud API). No third-party Terraform exports,
community gists, blog posts, or AI summaries — even when names match.

Records a tier-1+2 audit confirming the current cadvisor allowlist
matches both the Docker integration page and grafana/jsonnet-libs
docker-mixin/docker.json. Notes that tier-4 verification against the
live stack's full integration dashboard set was not performed and is
the only known gap.
2026-04-26 09:50:20 +07:00
tiennm99 54ab1fed27 feat: align metric/log collection with upstream Grafana Cloud integrations
Linux-Node integration:
- replace curated keep-list of ~140 node_* metrics with the upstream
  drop rule (drops only node_scrape_collector_*); ships the full
  ~130+ metric set the integration dashboards expect.
- add loki.source.file for /var/log/{syslog,messages,*.log} alongside
  the existing journal scrape, matching the upstream config.
- broaden the /var/log mount to cover both pipelines (was journal only).

Docker integration:
- drop container_memory_working_set_bytes from the cadvisor allowlist;
  not part of the documented metric set.

README: refresh "What it collects" + "Mounts" tables, document the
syslog-vs-journald duplication caveat for rsyslog hosts.
2026-04-26 09:24:44 +07:00
tiennm99 fd8cf058b2 docs: add Coolify SSH session spam runbook
Document a Coolify-specific noise pattern observed in the journal
pipeline: ~300 root sessions/hour from the Coolify host's connection
checks. Verified against coollabsio/coolify v4.x source (Kernel.php,
ServerManagerJob, ServerCheckJob, SshMultiplexingHelper).

Includes:
- exact call flow and skip conditions per Coolify source
- triage commands and key-fingerprint matcher
- two mitigations: drop at Alloy (loki.process stage.drop) or enable
  Sentinel server-side to bypass the SSH polling entirely
- framing: Coolify-only, base setup unchanged
2026-04-26 09:21:11 +07:00
tiennm99 513f688ea0 ci: bind alloy smoke test to real port (clustering needs non-zero) 2026-04-25 19:13:22 +07:00
tiennm99 b1911c0538 ci: make alloy smoke-test robust to runtime errors
Previous step failed in CI because the bare alloy container had no
/var/log/journal, no docker.sock, etc., so loki.source.journal +
discovery.docker exited the process — and --rm wiped the container
before logs could be inspected.

Now: drop --rm, mount the same host paths the prod compose uses, plus
an empty /var/log/journal stand-in. Capture logs unconditionally and
fail only on config-level patterns ('unknown component',
'undefined reference', 'syntax error', etc.). Runtime/component
failures against dummy endpoints are tolerated.
2026-04-25 19:12:05 +07:00
tiennm99 8a7f156487 fix: address review findings (network ns, journal path, CI semantics)
- network_mode: host so prometheus.exporter.unix reports real host
  interfaces (eth0...) rather than the alloy container's veth pair.
- loki.source.journal: set path = "/var/log/journal" explicitly so it
  doesn't silently fall through to /run/log/journal on volatile-journal
  hosts.
- cadvisor keep-list: add container_memory_working_set_bytes (drives
  several panels on the standard Docker dashboard).
- Drop /dev/kmsg device + extra_hosts:host.docker.internal — neither is
  needed by the current keep-lists, and host-network mode makes the
  extra_hosts entry meaningless.
- CI: extend Alloy validation beyond `fmt` (syntax-only) by booting
  alloy with the embedded config and asserting it stays running, which
  catches bad component refs / wrong arg names that fmt accepts.
- README: refresh Mounts table + Security note to match.
2026-04-25 19:09:49 +07:00
tiennm99 fed5d6f8c7 fix: point node_exporter at host bind mounts; drop pid:host
prometheus.exporter.unix now reads /rootproc and /rootfs (the existing
host bind mounts) instead of the container's own namespace, so the
filesystem + process metrics actually describe the host. This makes
pid:host unnecessary, so remove it — privileged is enough and pid:host
exposes every host process inside the container.
2026-04-25 11:20:03 +07:00
tiennm99 a63d142b53 feat: update cadvisor keep-list and add pid:host + machine-id mount
cadvisor regex: drop fs_inodes/fs_limit/network_tcp_usage; add fs_reads
+ fs_writes + network_(receive|transmit)_(errors|packets_dropped)_total
to match the standard Grafana Cloud docker integration dashboard.

Compose:
- pid: host so prometheus.exporter.unix sees host /proc (cpu, mem,
  load, processes) instead of the container's namespace.
- mount /etc/machine-id so loki.source.journal has a stable host id.
2026-04-25 11:16:13 +07:00
tiennm99 ac4c785053 ci: add REMOTECFG_* dummy vars so compose config passes
The compose file gained three new :?required guards
(REMOTECFG_URL/ID/USER) that the validate workflow was not exporting.
2026-04-25 10:51:16 +07:00
tiennm99 94c63452eb feat: merge linux+docker alloy configs and source creds from env
Embed unified config.alloy via compose configs, combining node_exporter
+ journal (linux) and cadvisor + docker logs (docker) collectors into one
container. Add remotecfg block for grafana fleet-management. Replace
hardcoded credentials with sys.env() reads of nine shell variables
(ALLOY_HOSTNAME, REMOTECFG_*, PROM_*, LOKI_*, GRAFANA_TOKEN).
2026-04-25 10:49:07 +07:00
tiennm99 f009b37ca0 fix: pass /dev/kmsg via devices: so cgroup allows the read
bind-mounting /dev/kmsg under volumes: creates the node but leaves the
device-cgroup controller blocking the read (EPERM). cap_drop=[ALL]
clears the default device allow-list, so even with CAP_SYSLOG the
kernel refuses. moving it under devices: adds the cgroup allow rule
alongside the bind-mount, which is what cadvisor actually needs.
2026-04-24 14:35:40 +07:00
tiennm99 9ca8c2a65b fix: mount /dev/kmsg and /run/udev/data for cadvisor + diskstats
clears two startup warnings:
 - cadvisor "Could not configure a source for OOM detection" — needs
   /dev/kmsg bind-mount and CAP_SYSLOG (kernel.dmesg_restrict=1 default)
 - node-exporter "Failed to open /run/udev/data" — diskstats collector
   enriches node_disk_* with model/serial/WWN labels from udev

both mounted read-only. CAP_SYSLOG added alongside DAC_OVERRIDE.
2026-04-24 14:24:00 +07:00
tiennm99 01261b18b0 fix: drop docker logs older than 6d before sending to loki
grafana cloud loki rejects entries older than 7 days with HTTP 400.
loki.source.docker has no "tail since" option, so on first start it
replays logs from each container's start time — long-running
containers (coolify-proxy, traffmonetizer) produced weeks of backlog
that loki refused to ingest.

insert a loki.process drop stage (older_than=144h) between the docker
source and loki.write.gc so stale entries are filtered before egress.
journal source already bounded by max_age=12h — no filter needed there.
2026-04-24 14:20:50 +07:00
tiennm99 59bf58a412 fix: grant DAC_OVERRIDE so alloy can write its data dir
container runs as root (0:0) but cap_drop=[ALL] stripped DAC_OVERRIDE,
so mkdir on /var/lib/alloy/data (alloy-owned in the image) failed with
permission denied. add it back (only DAC_OVERRIDE, nothing else) and
mount /etc/machine-id ro for a stable journal host id. annotate every
volume with the component that uses it.
2026-04-24 13:56:37 +07:00
tiennm99 fe7e48bf19 chore: switch license from mit to apache 2.0
fetched via gh api /licenses/apache-2.0 (github-official template).
2026-04-24 09:59:09 +07:00
tiennm99 1a894935d1 fix: expand one-line rule blocks (river requires newline-separated attrs)
caught by new ci — alloy fmt rejected `rule { a = b c = d }` on one line.
river/flow config language requires each attribute on its own line.
2026-04-24 09:55:15 +07:00
tiennm99 2ff943190c ci: add mit license + validate workflow
workflow runs on push to main and all prs:
- docker compose config -q (yaml + env interpolation)
- yq extract embedded alloy config
- alloy fmt (syntax check)
2026-04-24 09:54:03 +07:00
tiennm99 ca22737dfb docs: add readme with quick-start, multi-host pattern, security posture, free-tier budget 2026-04-24 09:49:34 +07:00
tiennm99 f5955a82f0 refactor: simplify for standalone single-host deployment
- merge discovery.relabel "metrics" into prometheus.relabel "filter"
  (instance label + keep-filter in one component)
- move logs 'instance' label to loki.write external_labels (DRY, one source)
- drop alloy container log exclusion (~500kb/day trivial; adds visibility
  into alloy's own health via loki)

net: -5 lines, one fewer component, same functionality
2026-04-24 09:43:32 +07:00
tiennm99 b055d4f66b refactor: tighten filters, add missing metrics + journal labels
correctness:
- drop dead ALLOY_HOSTNAME env passthrough (alloy uses constants.hostname)
- tighten netdev regex so iface names like 'lore0' no longer match 'lo'
- widen alloy-container drop to /alloy.* (catches renamed variants)

more info for same cost (~80 extra series, still ~1.6k of 10k):
- add MemFree, disk IOPS (reads/writes_completed), disk saturation
  (io_time_weighted), inode tracking (filesystem_files, _files_free),
  network drops, container working_set memory, container_last_seen,
  machine_scrape_error
- add journal unit + level labels (log querying by service/severity)
- docker discovery refresh 60s -> 5s (fast pickup of new containers)

efficiency:
- loki batch_wait=5s, batch_size=1MiB (~5x fewer push requests)

layout:
- drop redundant expose: [12345] (other containers can reach it anyway)
2026-04-24 09:27:57 +07:00
tiennm99 d0fb74ca58 feat: initial single-container alloy setup for grafana cloud free
- node_exporter (host) + cadvisor (containers) via one scrape job
- docker container logs + systemd journal to loki
- filter regex keeps only metrics needed for standard dashboards
- sized for grafana cloud free tier (~10k active series cap)
- shell env driven (no .env file); fail-fast on missing vars
- least-privilege: cap_drop all, no privileged, no pid host, ui not published
2026-04-24 09:21:29 +07:00
tiennm99 baac61cd70 feature: init 2025-04-07 12:44:18 +07:00
Tien Nguyen Minh b90541ac8b Initial commit 2025-04-07 12:39:46 +07:00
tiennm99 a420279f2b Update docker-compose.yml 2025-03-25 21:58:42 +07:00
tiennm99 7efe3c00ce Update docker-compose.yml 2025-03-25 20:58:52 +07:00
tiennm99 10c8824b4d feat: use docker stable image 2025-03-24 21:18:44 +07:00
tiennm99 fc766b2c49 feat: remove config 2025-03-24 21:15:41 +07:00
tiennm99 a3c1ea7cc3 feat: update config 2025-03-24 21:06:03 +07:00
tiennm99 292770e716 feat: init 2025-03-24 20:56:09 +07:00
Tien Nguyen Minh 9a10d42a43 Initial commit 2025-03-24 20:51:11 +07:00
tiennm99 03fbd4d3a2 Create docker-compose.yml 2025-03-15 20:43:33 +07:00
Tien Nguyen Minh cb4052d445 Initial commit 2025-03-15 20:37:50 +07:00
tiennm99 c1b902e455 feat: init 2025-03-07 00:42:01 +07:00
Tien Nguyen Minh ad4db2e29a Initial commit 2025-03-07 00:40:52 +07:00
Tien Nguyen Minh 2d0982ec55 Update README.md 2024-05-03 16:38:43 +07:00
tiennm99 c5567c623e [Init] 2023-12-23 08:15:16 +07:00
Tien Nguyen Minh 3426cf1da6 Initial commit 2023-12-23 07:29:50 +07:00
tiennm99 0fad91affd [Fix] run script 2023-11-13 14:34:08 +07:00
tiennm99 1320a67a15 Update README.md 2023-11-11 18:22:57 +07:00
tiennm99 416407d8ac chmod +x run.sh 2023-11-11 18:07:53 +07:00