build: container health check, licence in the image, CI hardening

HEALTHCHECK via noitu-server -healthcheck; the root LICENSE ships next
to NOTICE and the CI image check requires it; .claude and .env files
stay out of the build context. CI uses go-version stable, runs
govulncheck and npm audit, checks out without persisted credentials, and
proto.yml moves off the archived buf-setup-action. dependabot.yml is
dropped; the audit steps are the dependency signal. deployment.md gains
the Coolify/Traefik recipe, the stop-grace rule, the hello deadline and
per-address room budget, and the suppressed-log counter.
This commit is contained in:
tiennm99 committed 2026-09-29 20:33:16 +07:00
1 parent 1f2624c0e3
commit 00d3dadad4
6 files changed
+148 -15

No files matched your search

+4
View File
@@ -21,5 +21,9 @@ build-dictionary
build-dictionary.exe build-dictionary.exe
plans plans
docs docs
.claude
**/.claude
.env*
**/.env*
*.md *.md
!data/ATTRIBUTION.md !data/ATTRIBUTION.md
+28 -3
View File
@@ -24,9 +24,11 @@ jobs:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@v4
with:
persist-credentials: false
- uses: actions/setup-go@v5 - uses: actions/setup-go@v5
with: with:
go-version-file: server/go.mod go-version: stable
cache-dependency-path: server/go.sum cache-dependency-path: server/go.sum
- name: Vet - name: Vet
@@ -55,6 +57,14 @@ jobs:
with: with:
working-directory: server working-directory: server
# The module's go directive is only the minimum. CI runs the newest
# toolchain, the way the image's golang:1 does, so vet and race results
# come from the compiler that ships. govulncheck reads the call graph
# against the current advisory database.
- name: Vulnerability scan
run: go run golang.org/x/vuln/cmd/govulncheck@latest ./...
working-directory: server
# -race because the whole transport layer is goroutines and timers, and a # -race because the whole transport layer is goroutines and timers, and a
# data race there is exactly the kind of defect that passes without it. # data race there is exactly the kind of defect that passes without it.
- name: Test - name: Test
@@ -66,6 +76,8 @@ jobs:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@v4
with:
persist-credentials: false
- uses: actions/setup-node@v4 - uses: actions/setup-node@v4
with: with:
node-version: 24 node-version: 24
@@ -83,6 +95,12 @@ jobs:
run: npm run lint run: npm run lint
working-directory: web working-directory: web
# Runtime dependencies only: the frontend ships as static files, so a
# dev-tool advisory cannot reach production.
- name: Audit
run: npm audit --omit=dev --audit-level=high
working-directory: web
# npm test builds first, so this also proves the bundle compiles and # npm test builds first, so this also proves the bundle compiles and
# carries no wordlist. # carries no wordlist.
- name: Test - name: Test
@@ -94,9 +112,11 @@ jobs:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@v4
with:
persist-credentials: false
- uses: actions/setup-go@v5 - uses: actions/setup-go@v5
with: with:
go-version-file: server/go.mod go-version: stable
cache-dependency-path: server/go.sum cache-dependency-path: server/go.sum
- uses: actions/setup-node@v4 - uses: actions/setup-node@v4
with: with:
@@ -136,6 +156,8 @@ jobs:
needs: [go, web, e2e] needs: [go, web, e2e]
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@v4
with:
persist-credentials: false
# On a release the image is built from the real upstream release, which # On a release the image is built from the real upstream release, which
# is the only job in this file that downloads it. Every other run builds # is the only job in this file that downloads it. Every other run builds
@@ -159,7 +181,7 @@ jobs:
docker export check | tar -t > files.txt docker export check | tar -t > files.txt
docker rm check docker rm check
for required in app/data/LICENSE app/data/ATTRIBUTION.md app/NOTICE app/data/noitu.db; do for required in app/LICENSE app/data/LICENSE app/data/ATTRIBUTION.md app/NOTICE app/data/noitu.db; do
grep -qx "$required" files.txt || { echo "missing from the image: $required"; exit 1; } grep -qx "$required" files.txt || { echo "missing from the image: $required"; exit 1; }
done done
@@ -179,6 +201,9 @@ jobs:
sleep 1 sleep 1
done done
curl -fsS http://localhost:8080/healthz curl -fsS http://localhost:8080/healthz
# The image has no curl, so the HEALTHCHECK command is the binary
# itself; run it the way the container runtime does.
docker exec noitu /app/noitu-server -healthcheck
# A deep link is a client route, so the binary has to answer it with # A deep link is a client route, so the binary has to answer it with
# the app shell rather than a 404. # the app shell rather than a 404.
curl -fsS -o /dev/null -w '%{http_code}\n' 'http://localhost:8080/play?difficulty=2' | grep -qx 200 curl -fsS -o /dev/null -w '%{http_code}\n' 'http://localhost:8080/play?difficulty=2' | grep -qx 200
+8 -6
View File
@@ -19,19 +19,21 @@ jobs:
steps: steps:
- uses: actions/checkout@v4 - uses: actions/checkout@v4
with: with:
persist-credentials: false
# buf breaking compares against main, which needs real history. # buf breaking compares against main, which needs real history.
fetch-depth: 0 fetch-depth: 0
# A moving version rather than an exact pin, per house rule: updates # No version input, so the newest buf is used rather than an exact pin,
# arrive on the next run, and a breaking major would surface here before # per house rule: updates arrive on the next run, and a breaking major
# it could surprise a contributor's own machine. # would surface here before it could surprise a contributor's own
- uses: bufbuild/buf-setup-action@v1 # machine.
- uses: bufbuild/buf-action@v1
with: with:
version: latest setup_only: true
- uses: actions/setup-go@v5 - uses: actions/setup-go@v5
with: with:
go-version-file: server/go.mod go-version: stable
cache-dependency-path: server/go.sum cache-dependency-path: server/go.sum
- uses: actions/setup-node@v4 - uses: actions/setup-node@v4
+7
View File
@@ -77,6 +77,7 @@ COPY --from=dict /out/noitu.db /app/data/noitu.db
COPY data/LICENSE /app/data/LICENSE COPY data/LICENSE /app/data/LICENSE
COPY data/ATTRIBUTION.md /app/data/ATTRIBUTION.md COPY data/ATTRIBUTION.md /app/data/ATTRIBUTION.md
COPY NOTICE /app/NOTICE COPY NOTICE /app/NOTICE
COPY LICENSE /app/LICENSE
ENV NOITU_ADDR=:8080 \ ENV NOITU_ADDR=:8080 \
NOITU_DB_PATH=/app/data/noitu.db \ NOITU_DB_PATH=/app/data/noitu.db \
@@ -84,4 +85,10 @@ ENV NOITU_ADDR=:8080 \
EXPOSE 8080 EXPOSE 8080
USER nonroot:nonroot USER nonroot:nonroot
# The image has no curl or wget, so the health check is the binary itself:
# -healthcheck GETs /healthz on NOITU_ADDR and exits 0 or 1. It is liveness
# only; /readyz is the drain signal a load balancer polls.
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
CMD ["/app/noitu-server", "-healthcheck"]
ENTRYPOINT ["/app/noitu-server"] ENTRYPOINT ["/app/noitu-server"]
+2 -1
View File
@@ -218,7 +218,8 @@ Endpoints: `GET /ws` (Protobuf over binary WebSocket frames), `GET /healthz`
(liveness), `GET /readyz` (readiness — 503 while draining), `GET /version` (liveness), `GET /readyz` (readiness — 503 while draining), `GET /version`
(plain text), and — when `NOITU_WEB_DIR` is set — the frontend on everything (plain text), and — when `NOITU_WEB_DIR` is set — the frontend on everything
else, with unknown paths falling back to `index.html` because deep links are else, with unknown paths falling back to `index.html` because deep links are
client routes. client routes. `noitu-server -healthcheck` probes `/healthz` on `NOITU_ADDR` and exits 0 or 1;
it is what the image's `HEALTHCHECK` runs, since the image has no `curl`.
### Smoke-testing without a frontend ### Smoke-testing without a frontend
+99 -5
View File
@@ -86,9 +86,14 @@ one.
### What travels with the data ### What travels with the data
The derived wordlist is CC BY-SA 4.0 while the code is Apache-2.0, so the image The derived wordlist is CC BY-SA 4.0 while the code is Apache-2.0, so the image
carries `data/LICENSE`, `data/ATTRIBUTION.md` and `NOTICE` alongside it. CI carries `data/LICENSE`, `data/ATTRIBUTION.md`, `NOTICE` and the Apache-2.0
asserts all three are present, and that the upstream file is not. Removing them `LICENSE` that `NOTICE` refers to. CI asserts all four are present, and that
would put the image out of compliance. the upstream file is not. Removing them would put the image out of compliance.
The image does not carry a notice file for the third-party Go and JavaScript
dependencies. That is accepted while the image is only built and run by the
operator; it must be generated before an image is ever published or handed to
someone else.
## Behind a reverse proxy ## Behind a reverse proxy
@@ -135,6 +140,39 @@ noitu.example {
} }
``` ```
### Coolify and Traefik
Coolify fronts every app with Traefik. With its default forwarded-headers
settings Traefik discards any `X-Forwarded-For` the client sent and appends the
real peer, so it is safe to trust. Left at the defaults of this server it is
not trusted, and the whole player base shares one client address: one join
bucket, and no way to switch the per-address connection cap on. Set these on
the app's environment variables:
| Variable | Value | What it changes |
|---|---|---|
| `NOITU_TRUSTED_PROXIES` | the subnet Traefik reaches the app on, for example `10.0.1.0/24` | `X-Forwarded-For` is believed when the socket peer is inside this range, so join limiting and the per-address cap key on the real player. Find the range with `docker network inspect coolify` (or the app's own network, if Coolify put it on one) and use the narrowest CIDR that contains the Traefik container. Never a range the public can connect from |
| `NOITU_MAX_CONNECTIONS_PER_IP` | `32` as a starting point | One address may hold at most this many open sockets. Off (`0`) by default because it is only meaningful once the address above is the real client; 32 leaves room for a household or campus NAT. Without it one host can hold every socket up to `NOITU_MAX_CONNECTIONS` |
| `NOITU_DRAIN_TIMEOUT` | below the container stop grace; see "Draining on deploy" | Live games get this long to finish on redeploy instead of ending at once |
Variables take effect on the next deploy. An entry in `NOITU_TRUSTED_PROXIES`
that does not parse is logged as a warning at startup and skipped, so read the
startup log after redeploying; a range that is too narrow to contain Traefik
fails quietly, back to one shared address.
If Cloudflare or another CDN sits in front of Traefik, Traefik itself must
trust that CDN's ranges (`forwardedHeaders.trustedIPs` on the entrypoint),
otherwise it overwrites the header with the CDN edge address and every player
behind the same edge shares one bucket again.
To make Traefik stop routing to an instance that has started draining, give the
service a load-balancer health check on `/readyz` through Coolify's custom
labels: `traefik.http.services.<service>.loadbalancer.healthcheck.path=/readyz`
and `...healthcheck.interval=5s`, using the service name Coolify generated for
the app (visible in the container's labels). Without it Traefik keeps sending
new players to the old container until it is removed, and they are refused with
`server_restarting`.
### The client's own address ### The client's own address
Rate limiting keys on the client's address, and by default that is the Rate limiting keys on the client's address, and by default that is the
@@ -169,6 +207,14 @@ client past it is disconnected rather than throttled. The defaults are
generous for one binary on a small host; lower them if memory is tight, generous for one binary on a small host; lower them if memory is tight,
because a room is a goroutine and an engine held for up to its idle window. because a room is a goroutine and an engine held for up to its idle window.
Neither ceiling can be held by one client alone. A socket that has not sent
its `Hello` within ten seconds is closed, so an idle socket cannot sit on a
connection slot, and the room-creation budget is charged to the client
address as well as to the socket, so reconnecting does not refill it. The
per-address budget is deliberately wide (a burst of 30, refilling at one room
every two seconds), because without `NOITU_TRUSTED_PROXIES` every player
behind the proxy shares it.
A third ceiling, `NOITU_MAX_CONNECTIONS_PER_IP`, bounds how many of those A third ceiling, `NOITU_MAX_CONNECTIONS_PER_IP`, bounds how many of those
sockets one address may hold at once, and it is off by default. Turning it on sockets one address may hold at once, and it is off by default. Turning it on
is safe only once the client's own address (above) is the real one: behind a is safe only once the client's own address (above) is the real one: behind a
@@ -187,13 +233,19 @@ below, expvar always publishes the process's full command line and its
runtime memory statistics; that is the standard library's own doing, not runtime memory statistics; that is the standard library's own doing, not
something this server adds, and it is the whole reason `/debug/vars` lives on something this server adds, and it is the whole reason `/debug/vars` lives on
a separate address rather than a route on the public mux one config change a separate address rather than a route on the public mux one config change
could expose. The counters, all prefixed `noitu_`: connections open and could expose. In a container, binding it to `0.0.0.0` (`:6060`) makes it
reachable from every other container on the same Docker network, which on
Coolify can be the shared proxy network. Bind it to `127.0.0.1:6060`,
which only the container itself can reach, or leave it unset; do not put it on
a wildcard address on a shared network. The counters, all prefixed `noitu_`: connections open and
total; rooms live and total, each split `bot`/`pvp`; games started and total; rooms live and total, each split `bot`/`pvp`; games started and
finished the same way; words submitted, accepted, and rejected by reason; finished the same way; words submitted, accepted, and rejected by reason;
eliminations by reason; chat lines; join attempts refused, by whether it was eliminations by reason; chat lines; join attempts refused, by whether it was
the rate limit, an unknown code, or a full room; bot moves by difficulty; dead- the rate limit, an unknown code, or a full room; bot moves by difficulty; dead-
end claims by whether the position actually had no legal move; words reported end claims by whether the position actually had no legal move; words reported
as real by `ReportWord`; and resumes attempted versus succeeded. None of it is as real by `ReportWord`; resumes attempted versus succeeded; and
`word_rejected`/`word_reported` log lines dropped by the process-wide
rate limit on them (`noitu_corpus_log_suppressed`). None of it is
read by the game itself — it is a second write next to a decision already read by the game itself — it is a second write next to a decision already
made, not an input to one. made, not an input to one.
@@ -234,6 +286,26 @@ rooms, 503 once it has started draining (see below). Point a load balancer's
`/healthz`; pointing both at the same endpoint defeats the reason there are `/healthz`; pointing both at the same endpoint defeats the reason there are
two. two.
The image has no `curl` or `wget`, so a container health check cannot shell out
to one. The binary probes itself instead:
```sh
noitu-server -healthcheck
```
It sends `GET /healthz` to `NOITU_ADDR` (a bare `:8080`, or a wildcard host, is
dialled on `127.0.0.1`), and exits 0 on a 200 and 1 on anything else, printing
the reason to stderr. The Dockerfile's `HEALTHCHECK` runs it every 30 seconds
with a 10-second start period. It is a liveness check, so it deliberately does
not follow `/readyz` into a drain.
On Coolify, leave the dashboard's own HTTP health check switched off: it
executes `curl` or `wget` inside the container and cannot pass against this
image. Coolify picks the health check up from the Dockerfile after the next
deploy (the application then reports a custom health check found), and a
rolling update waits for the new container to be healthy before it removes the
old one. Confirm that in the deployment log after the first deploy.
A post-deploy check worth having beyond either is a real socket open, because A post-deploy check worth having beyond either is a real socket open, because
both health checks pass whether or not the proxy forwards upgrades. Opening both health checks pass whether or not the proxy forwards upgrades. Opening
the site and starting a game against the bot is the shortest version of that. the site and starting a game against the bot is the shortest version of that.
@@ -259,9 +331,31 @@ plain `kill` used to do that the instant the signal arrived, which is why
connected is told the server is restarting and the process shuts down as connected is told the server is restarting and the process shuts down as
it always did. it always did.
After the notice goes out the process waits two more seconds, so the writes
reach the sockets before it exits, then closes its listeners.
Each step logs the room and live-game count, so "did the deploy actually Each step logs the room and live-game count, so "did the deploy actually
wait, and for what" is answered from the log rather than guessed at. wait, and for what" is answered from the log rather than guessed at.
A second `SIGTERM` or `SIGINT` during a drain is not swallowed: the first
signal hands signal handling back to the runtime, so a second one ends the
process at once.
The container runtime, not this server, bounds the whole sequence. Docker
sends `SIGTERM`, waits its stop grace period (10 seconds by default), then
sends `SIGKILL`, which gives players no `server_restarting` notice and writes
no final log line. The rule is
NOITU_DRAIN_TIMEOUT + 2s < container stop grace
so with Docker's default grace `NOITU_DRAIN_TIMEOUT` must stay at or below
about `6s`. A drain shorter than one turn only helps games in their last
seconds; to let games ride out a full `NOITU_TURN_LIMIT` (30 seconds by
default), raise the stop grace first, on Coolify wherever the application's
container stop timeout is configured, and only then raise the drain timeout.
If the grace cannot be raised, keep the drain short instead of letting the
kill land mid-drain.
The default, `NOITU_DRAIN_TIMEOUT=0s`, is today's behaviour: nothing waits, The default, `NOITU_DRAIN_TIMEOUT=0s`, is today's behaviour: nothing waits,
every live game ends immediately. Setting it to something like `60s` turns every live game ends immediately. Setting it to something like `60s` turns
"deploy when the game is quiet" into "deploy whenever, and the games in "deploy when the game is quiet" into "deploy whenever, and the games in