Gitea's migrate API stays silent until the clone ends, and Bun's fetch
drops a connection idle for 5 minutes. gitea-mirror then marked large
repositories failed while Gitea kept cloning, and a later retry took the
half-made repository for a finished mirror. GITEA_CLONE_TIMEOUT now also
sets BUN_CONFIG_HTTP_IDLE_TIMEOUT.
Deletes the archived-* Gitea copies gitea-mirror keeps when a source in the
given owners disappears, and removes their tracking rows so they are not
re-mirrored. Dry run by default.
The skill reached Gitea on localhost and the mirror database inside the
container, which only worked on a local host. It now uses tea for Gitea
and the gitea-mirror API key for repository status and retries, in bash.
Private upstreams are no longer probed anonymously, which reported them
as deleted.
sources/ holds gitignored upstream checkouts for debugging a service
against its real code; the debug-service skill walks through it. Coolify
is now the primary deployment target and Dokploy optional.
Drop the localhost-bound ports, read the database password and both
public URLs from the environment, pin gitea to its major tag, add a gitea
health check and disable gitea's SSH server.
A service README now describes only its own service: no links to other
services or to the root, and no restating of the shared conventions that
the root README and CLAUDE.md already carry. Each service is a separate
Coolify app on a <service>/** watch path, so a cross-link made editing one
service redeploy another. The rule is recorded at the root; the alloy
compose comment now points at a heading that exists.
Coolify injects the same value when a compose service omits one, but Dokploy
runs the file as written, so in its default compose mode an omitted policy
leaves the container down after a crash or a host reboot.
Remove the directory and its row from the root table, the related-services link
in opencode-web, and the workspace-volume convention's reference to it.
Diun needs the Docker API to enumerate containers and inspect each one's
image. Mounting the socket into it directly is host-root-equivalent, and :ro on
a socket mount is cosmetic, so the socket goes into a docker-socket-proxy
sidecar and Diun reaches it at tcp://dockerproxy:2375. POST is revoked there,
so container create and exec return 403.
CONTAINERS and IMAGES are both required: with CONTAINERS alone the provider
loads and enumerates containers, then every ImageInspect returns 403 and
nothing is analysed. Verified against a live watch cycle — 25 images analysed,
no errors.
Pin crazymax/diun:4.33 rather than :latest, since this is the service whose job
is to talk to the daemon.
The workspace moves off the config volume onto code-server-workspace, so
wiping editor state and wiping code are separate acts.
The image only ever chowns the literal path /config/workspace, and reads
DEFAULT_WORKSPACE to pick the folder to open, so a named volume on /workspace
would come up root-owned and unwritable. Creating the directory in a local
Dockerfile seeds the volume with the right ownership instead.
Compose materialises a configs: entry with inline content by writing it into
the container and refuses to do so on a read-only service: "cannot create
config ... : `file` is the sole supported option". The container was created
without /etc/alloy/config.alloy and the deployment failed at start.
Keeping the config inline matters more than the read-only rootfs, so the flag
and its tmpfs go. Every other control stays: no privileged, cap_drop ALL with
only DAC_OVERRIDE added, no-new-privileges, the socket proxy, and the limits.
Replace privileged: true with cap_drop ALL plus DAC_OVERRIDE, no-new-privileges,
a read-only rootfs and memory/pid limits. DAC_OVERRIDE is what lets the uid-0
entrypoint create its storage directory and read the journal and /rootfs; every
other capability stays dropped.
Route prometheus.exporter.cadvisor, discovery.docker and loki.source.docker
through a docker-socket-proxy sidecar on 127.0.0.1:2375 instead of bind-mounting
the socket. POST is refused there, so container create and exec are no longer
reachable. NETWORKS is granted because Docker SD resolves network names per
container and returns no targets without it.
Add an alloy validate step to CI and boot the test container with the shipped
capability set, read-only rootfs and proxy rather than --privileged.
openhands runs each agent session in a container it spawns through the host
docker socket, so the socket is mounted read-write and host.docker.internal is
resolved. No host workspace is exposed.
opencode-web serves the opencode agent as a browser UI. The vendor image ships
only the opencode binary on bare Alpine, so a local Dockerfile adds bash, git,
curl and an ssh client.
The base image ships /home/paseo owned by uid 1000 and declares it a
volume, so a fresh named volume is seeded with that ownership, and the
base entrypoint chowns it and the agent config directories when they are
not.
The image installs python, gh, glab, build-essential, sudo, zsh and nano
only; go or a jvm goes into $HOME from a terminal instead.
SHELL and TZ are written into compose.yml rather than read from .env, so
the deploying shell's own values can no longer win over them.
The entrypoint calls chpasswd, chown and gosu by absolute path, tolerates
a chpasswd failure instead of taking the start down with it, turns
globbing off around the agent loop, and counts an agent as installed only
when its binary actually runs.
environment: blocks were in no particular order. They now run must-have ->
should-have -> optional, with related variables kept adjacent as a group that
takes the tier of its most important member: PUID/PGID, PASSWORD with
SUDO_PASSWORD, DOCKER_MODS ahead of the INSTALL_PACKAGES and
NODEJS_MOD_VERSION that configure it, the four GIT_* entries, the PASEO_*
daemon settings.
Each .env.example is reordered to match its compose file. The names do not map
one to one -- PASSWORD feeds both PASSWORD and SUDO_PASSWORD, SERVICE_HOSTNAME
feeds HOST -- so an entry sits where the first compose entry reading it sits.
The HOST comment in both compose files is dropped; the READMEs already carry
that explanation in full. CLAUDE.md records the ordering convention.
alloy and gitea-mirror-local are untouched: every variable there is required,
so the tiers collapse and the existing grouping is the better one.
pnpm is not used anywhere -- npm is the package manager everywhere -- so the
mod that installs it is dead weight on every container start.
INSTALL_PACKAGES loses apache2-utils, bfs, ffmpeg, imagemagick, librsvg2-bin,
lsof, psmisc and ugrep, and gains glab. Each entry is re-resolved by apt on
every start, so the list is kept to what is actually reached for.
Record the convention in CLAUDE.md: .env.example is a template, so its values
stay generic and the real ones are set per deployment in Coolify or Dokploy.
Note the naming trap alongside it -- Compose interpolation reads the deploying
shell's environment before the .env file, so a variable must not collide with
one the shell already exports.
The alloy README's example host is made generic to match.
Coolify injects HOST=0.0.0.0 into every compose app, and zsh seeds $HOST and
the %m/%M prompt escapes from that variable rather than calling gethostname().
The prompt read "0", the first dot-separated field of 0.0.0.0, even though the
container hostname itself was set correctly.
Pass the hostname in as HOST alongside the hostname: key, both from a single
SERVICE_HOSTNAME variable. Neither service reads HOST itself -- code-server
binds [::]:8443, Paseo binds PASEO_LISTEN -- so this only affects the prompt.
Not named HOSTNAME: Compose interpolation lets the deploying shell's
environment win over the .env file, and HOSTNAME is set in every container,
including the one Coolify runs in.
PASEO_LABEL is renamed to SERVICE_HOSTNAME; it was never a Paseo variable.
The example git identity is blanked out along with it.
SDKMAN goes into ~/.sdkman whenever that directory is missing, the same way
the agent CLIs arrive: as paseo, through gosu, onto the volume, where `sdk
install` can write and the candidates persist. No variable gates it.
The chown of /home/paseo moves out of the AGENT_CLIS guard, since SDKMAN now
needs it even when no agent is named, and the agent loop reads AGENT_CLIS into
a local first so `set -u` does not trip on it being unset.
SDKMAN_DIR is set in the image, and the current/bin of java, scala, gradle,
maven and sbt joins PATH: `sdk` itself is a shell function from the rc hook and
so exists in terminals only, while the daemon and an agent's non-interactive
commands read no rc file and need the binaries on PATH.
paseo-sudo-entrypoint was named for the one thing it used to do. It is
/usr/local/bin/entrypoint now, matching the source file.
Bind-mount /var/run/docker.sock so the universal-docker mod's CLI has a
daemon to talk to, and document the sibling-container and permission
caveats in the service README.
The entrypoint runs each vendor's installer through gosu paseo for every name
in AGENT_CLIS whose command is not already on PATH, so a fresh paseo-home
volume comes up with agents ready. Defaults to claude and codex; the other
four are opt-in.
On start rather than in the Dockerfile: Docker seeds a named volume from the
image once, at creation, so a build-time install into /home/paseo would only
ever reach a volume that did not exist yet. The check also chowns /home/paseo
first, since a freshly created volume can arrive owned by root and the base
entrypoint has not run its own chown by then.
Unknown names and failed installs are logged and skipped rather than taking
the container down with them.
entrypoint.sh loses its explanations to the README along the way.
The daemon probes each provider's binary with `which` in its own
environment, which comes from the image and never sources a shell rc, so
agents installed into $HOME were reported unavailable.
The paseo user cannot write /usr/local, so a baked-in agent CLI can never
apply its own update — every one of them ships an updater that expects to
rewrite its own binary, and Claude Code nags about the failure at startup.
Installed under $HOME they update themselves and still persist, since that is
the paseo-home volume. The README now lists each vendor's installer.
Bun goes with them: it was only there because omp is compiled against it.
glab arrives as the .deb from GitLab's releases page, pinned by GLAB_VERSION
because the URL carries the version. GitLab runs no apt repository and calls
Homebrew its only officially supported Linux package manager.
The Dockerfile drops its explanations along the way; they are in the README.
Comments say what a section installs or configures. Why not the distro
package, why that directory, why a version is pinned — that belongs in the
service README, where it can be read in full.
Replaces the line under "Installing software in an image" that asked for the
opposite.
The image had no compiler at all -- gcc, g++, make, cc and ld were all
missing -- so cgo, npm's node-gyp addons and Python C extensions could
not build. Goes in the existing apt layer to keep a single apt-get
update.
Drop the history and roadmap asides: which services predate the collection's
conventions, the unwired Open Web UI plan, the generic clone-and-troubleshoot
boilerplate. Services that publish ports or set `restart:` now simply say so.
couchbase, openvpn-as and traffmonetizer had two-line READMEs; give them the
ports, variables and storage the root README promises. traffmonetizer reads
${TOKEN} and had no .env.example, so add one.
Six agents now ship in the image. All come from npm; omp additionally needs
Bun, since its bin is Bun-compiled and opens with `#!/usr/bin/env bun`.
BUN_INSTALL puts Bun in /usr/local, clear of the paseo-home volume.
GIT_NAME and GIT_EMAIL expand into the four GIT_AUTHOR_*/GIT_COMMITTER_*
variables, the same shape code-server already uses. Passing the identity as
environment rather than running `git config` means agents and terminals commit
correctly with no setup step, and it does not depend on ~/.gitconfig surviving
in the /home/paseo volume.
The base image installs no sudo and leaves paseo out of the sudo group, so
add both. The password cannot be baked in at build time -- it is a secret and
would land in a layer -- and it cannot be set after the base entrypoint, which
ends in `exec gosu paseo` and never returns. Wrap that entrypoint instead and
set the password while still root, on every start: /etc/shadow lives in the
image, not on the /home/paseo volume, so it reverts on each recreate.
Paseo spawns terminals with process.env.SHELL and falls back to /bin/sh,
which is dash. It never consults the login shell, so chsh has no effect --
and would be reverted by the next rebuild regardless.
System packages, Python, Go, gh and the agent CLIs each get their own
block. Ordered least-changing first, so a Claude Code bump no longer
re-runs the toolchain installs.
Go 1.26.8 from the official tarball and Python 3.12 via uv, since Debian
12 carries 1.19 and 3.11. Both land outside $HOME, which the paseo-home
volume would otherwise mask.
Hostname and timezone come from the environment rather than being fixed in
the compose file. Paseo uses the container hostname as the host label in
its web UI, so without one the UI shows a random container ID.
Debian does not package gh, so use GitHub's signed apt repo. Config
lands in /home/paseo/.config/gh, inside the paseo-home volume, so the
login survives a redeploy.
The pairing screen needs an explicit port, which is the least obvious
part of getting connected. Lead with the setup steps and cut the
explanation around them down to what a reader has to act on.
The daemon trusts X-Forwarded-Proto from loopback only by default, but
Coolify's Traefik reaches it from the Docker bridge network. It therefore
reported the request as plain HTTP and handed the UI useTls: false, so the
UI built a ws:// URL on an https:// page. The browser blocked it as mixed
content and the UI fell back to its built-in localhost:6767 default.
uniquelocal covers the private ranges Docker uses. An exact CIDR is
tighter but Coolify assigns a fresh subnet per project.
The upstream image ships no agent CLIs, so build from a local Dockerfile
that layers Claude Code on top. npm delivers the same native binary as
the standalone installer, which cannot be used here: it writes to
$HOME/.local, and $HOME is /home/paseo, a volume mount that masks
anything baked in at build time.
Runs as root by design -- the entrypoint chowns the mounted volumes and
then drops to the unprivileged paseo user with gosu.
Set up the per-service layout: each service directory holds compose.yml
with a committed .env.example and a gitignored .env.
Add code-server as the first service, plus a README and CLAUDE.md
documenting that these files target Coolify/Dokploy and deliberately
omit ports and restart policies.
Invoke-Tea started its job in the caller's directory. When that is a git
work tree, tea infers the target from the local remote, and a remote that
matches no configured login makes it discard --login, fall back to the
first login, and fail with "remote repository required". Pin the job to a
neutral directory so the requested login always resolves.
Adds a skill to audit the local Gitea mirror stack for repos whose pull
failed, then clean up only what is safe to delete.
Detection combines four signals, since none is sufficient alone: the Gitea
API (empty repos with no completed initial pull), an upstream reachability
probe, the mirror app database, and the gitea container log. The container
log reflects a retention window rather than history, so it is never the
sole basis for deletion.
Deletion is gated on a contradiction between Gitea and the mirror app:
a repo that is empty while the app records the pull as finished. Repos the
app is still cloning or has queued look identical by API fields alone
(empty, zero size, no mirror timestamp), so they are excluded to avoid
destroying work in progress. Repos that still hold content are reported
for retry and never deleted, so a transient fetch error cannot cost a
mirror.
Deleting a broken repo also resets its mirror-app row to pending;
without that the app never re-pulls it and the mirror is lost instead of
restored. Cleanup is dry-run by default and warns on a stale plan.
Document the compose services, published ports, volumes, and first-run
setup. Fill .env.example with sample values and note that compose.yml
does not yet consume them.
container_memory_usage_bytes counts page cache attributed to the cgroup,
which makes disk-heavy containers (e.g. gitea) appear to use ~all host RAM.
Add container_memory_working_set_bytes (the metric Grafana's Docker
integration dashboard expects), plus container_memory_rss,
container_memory_cache, and container_spec_memory_limit_bytes for
breakdown and limit-percentage panels.
`loki.source.journal "default"` no longer pins `path = "/var/log/journal"`.
Upstream omits the field, which lets Alloy default to BOTH
`/var/log/journal` (persistent) and `/run/log/journal` (volatile). The
explicit path silently dropped journal logs on hosts with volatile-only
storage. Matches the canonical Linux Node integration template.
Also bumps the image six minor versions to current stable. Doc records
the 2026-04-26 re-audit.
Copies the exact 157-metric list from the Linux Node integration's
Metrics anchor as the cadvisor keep-list already does for Docker
(16 metrics). Replaces the earlier `drop node_scrape_collector_.+`
rule, which was the integration page's alternate snippet but didn't
ship the explicit allowlist users see in the docs.
`instance:node_num_cpu:sum` from the Metrics section is intentionally
omitted — it's a recording-rule output computed server-side by
Grafana Cloud's ruler, not produced by the agent.
Doc + README updated to point at the Metrics anchors directly so the
source of truth is unambiguous.
Adds docs/upstream-sources-of-truth.md as the binding policy for what
this repo follows when deciding metrics, labels, log pipelines, and
dashboards to ship.
Hard rule: only tier 1-4 official sources (Grafana Cloud integration
docs, github.com/grafana/*, github.com/prometheus/*, the user's own
authenticated Grafana Cloud API). No third-party Terraform exports,
community gists, blog posts, or AI summaries — even when names match.
Records a tier-1+2 audit confirming the current cadvisor allowlist
matches both the Docker integration page and grafana/jsonnet-libs
docker-mixin/docker.json. Notes that tier-4 verification against the
live stack's full integration dashboard set was not performed and is
the only known gap.
Linux-Node integration:
- replace curated keep-list of ~140 node_* metrics with the upstream
drop rule (drops only node_scrape_collector_*); ships the full
~130+ metric set the integration dashboards expect.
- add loki.source.file for /var/log/{syslog,messages,*.log} alongside
the existing journal scrape, matching the upstream config.
- broaden the /var/log mount to cover both pipelines (was journal only).
Docker integration:
- drop container_memory_working_set_bytes from the cadvisor allowlist;
not part of the documented metric set.
README: refresh "What it collects" + "Mounts" tables, document the
syslog-vs-journald duplication caveat for rsyslog hosts.
Document a Coolify-specific noise pattern observed in the journal
pipeline: ~300 root sessions/hour from the Coolify host's connection
checks. Verified against coollabsio/coolify v4.x source (Kernel.php,
ServerManagerJob, ServerCheckJob, SshMultiplexingHelper).
Includes:
- exact call flow and skip conditions per Coolify source
- triage commands and key-fingerprint matcher
- two mitigations: drop at Alloy (loki.process stage.drop) or enable
Sentinel server-side to bypass the SSH polling entirely
- framing: Coolify-only, base setup unchanged
Previous step failed in CI because the bare alloy container had no
/var/log/journal, no docker.sock, etc., so loki.source.journal +
discovery.docker exited the process — and --rm wiped the container
before logs could be inspected.
Now: drop --rm, mount the same host paths the prod compose uses, plus
an empty /var/log/journal stand-in. Capture logs unconditionally and
fail only on config-level patterns ('unknown component',
'undefined reference', 'syntax error', etc.). Runtime/component
failures against dummy endpoints are tolerated.
- network_mode: host so prometheus.exporter.unix reports real host
interfaces (eth0...) rather than the alloy container's veth pair.
- loki.source.journal: set path = "/var/log/journal" explicitly so it
doesn't silently fall through to /run/log/journal on volatile-journal
hosts.
- cadvisor keep-list: add container_memory_working_set_bytes (drives
several panels on the standard Docker dashboard).
- Drop /dev/kmsg device + extra_hosts:host.docker.internal — neither is
needed by the current keep-lists, and host-network mode makes the
extra_hosts entry meaningless.
- CI: extend Alloy validation beyond `fmt` (syntax-only) by booting
alloy with the embedded config and asserting it stays running, which
catches bad component refs / wrong arg names that fmt accepts.
- README: refresh Mounts table + Security note to match.
prometheus.exporter.unix now reads /rootproc and /rootfs (the existing
host bind mounts) instead of the container's own namespace, so the
filesystem + process metrics actually describe the host. This makes
pid:host unnecessary, so remove it — privileged is enough and pid:host
exposes every host process inside the container.
cadvisor regex: drop fs_inodes/fs_limit/network_tcp_usage; add fs_reads
+ fs_writes + network_(receive|transmit)_(errors|packets_dropped)_total
to match the standard Grafana Cloud docker integration dashboard.
Compose:
- pid: host so prometheus.exporter.unix sees host /proc (cpu, mem,
load, processes) instead of the container's namespace.
- mount /etc/machine-id so loki.source.journal has a stable host id.
bind-mounting /dev/kmsg under volumes: creates the node but leaves the
device-cgroup controller blocking the read (EPERM). cap_drop=[ALL]
clears the default device allow-list, so even with CAP_SYSLOG the
kernel refuses. moving it under devices: adds the cgroup allow rule
alongside the bind-mount, which is what cadvisor actually needs.
clears two startup warnings:
- cadvisor "Could not configure a source for OOM detection" — needs
/dev/kmsg bind-mount and CAP_SYSLOG (kernel.dmesg_restrict=1 default)
- node-exporter "Failed to open /run/udev/data" — diskstats collector
enriches node_disk_* with model/serial/WWN labels from udev
both mounted read-only. CAP_SYSLOG added alongside DAC_OVERRIDE.
grafana cloud loki rejects entries older than 7 days with HTTP 400.
loki.source.docker has no "tail since" option, so on first start it
replays logs from each container's start time — long-running
containers (coolify-proxy, traffmonetizer) produced weeks of backlog
that loki refused to ingest.
insert a loki.process drop stage (older_than=144h) between the docker
source and loki.write.gc so stale entries are filtered before egress.
journal source already bounded by max_age=12h — no filter needed there.
container runs as root (0:0) but cap_drop=[ALL] stripped DAC_OVERRIDE,
so mkdir on /var/lib/alloy/data (alloy-owned in the image) failed with
permission denied. add it back (only DAC_OVERRIDE, nothing else) and
mount /etc/machine-id ro for a stable journal host id. annotate every
volume with the component that uses it.