mirror of
https://github.com/tiennm99/noitu.git
synced 2026-10-11 03:13:45 +00:00
build: container health check, licence in the image, CI hardening
HEALTHCHECK via noitu-server -healthcheck; the root LICENSE ships next to NOTICE and the CI image check requires it; .claude and .env files stay out of the build context. CI uses go-version stable, runs govulncheck and npm audit, checks out without persisted credentials, and proto.yml moves off the archived buf-setup-action. dependabot.yml is dropped; the audit steps are the dependency signal. deployment.md gains the Coolify/Traefik recipe, the stop-grace rule, the hello deadline and per-address room budget, and the suppressed-log counter.
This commit is contained in:
1 parent
1f2624c0e3
commit
00d3dadad4
6 files changed
+148
-15
No files matched your search
@@ -21,5 +21,9 @@ build-dictionary
|
||||
build-dictionary.exe
|
||||
plans
|
||||
docs
|
||||
.claude
|
||||
**/.claude
|
||||
.env*
|
||||
**/.env*
|
||||
*.md
|
||||
!data/ATTRIBUTION.md
|
||||
@@ -24,9 +24,11 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
persist-credentials: false
|
||||
- uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version-file: server/go.mod
|
||||
go-version: stable
|
||||
cache-dependency-path: server/go.sum
|
||||
|
||||
- name: Vet
|
||||
@@ -55,6 +57,14 @@ jobs:
|
||||
with:
|
||||
working-directory: server
|
||||
|
||||
# The module's go directive is only the minimum. CI runs the newest
|
||||
# toolchain, the way the image's golang:1 does, so vet and race results
|
||||
# come from the compiler that ships. govulncheck reads the call graph
|
||||
# against the current advisory database.
|
||||
- name: Vulnerability scan
|
||||
run: go run golang.org/x/vuln/cmd/govulncheck@latest ./...
|
||||
working-directory: server
|
||||
|
||||
# -race because the whole transport layer is goroutines and timers, and a
|
||||
# data race there is exactly the kind of defect that passes without it.
|
||||
- name: Test
|
||||
@@ -66,6 +76,8 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
persist-credentials: false
|
||||
- uses: actions/setup-node@v4
|
||||
with:
|
||||
node-version: 24
|
||||
@@ -83,6 +95,12 @@ jobs:
|
||||
run: npm run lint
|
||||
working-directory: web
|
||||
|
||||
# Runtime dependencies only: the frontend ships as static files, so a
|
||||
# dev-tool advisory cannot reach production.
|
||||
- name: Audit
|
||||
run: npm audit --omit=dev --audit-level=high
|
||||
working-directory: web
|
||||
|
||||
# npm test builds first, so this also proves the bundle compiles and
|
||||
# carries no wordlist.
|
||||
- name: Test
|
||||
@@ -94,9 +112,11 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
persist-credentials: false
|
||||
- uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version-file: server/go.mod
|
||||
go-version: stable
|
||||
cache-dependency-path: server/go.sum
|
||||
- uses: actions/setup-node@v4
|
||||
with:
|
||||
@@ -136,6 +156,8 @@ jobs:
|
||||
needs: [go, web, e2e]
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
# On a release the image is built from the real upstream release, which
|
||||
# is the only job in this file that downloads it. Every other run builds
|
||||
@@ -159,7 +181,7 @@ jobs:
|
||||
docker export check | tar -t > files.txt
|
||||
docker rm check
|
||||
|
||||
for required in app/data/LICENSE app/data/ATTRIBUTION.md app/NOTICE app/data/noitu.db; do
|
||||
for required in app/LICENSE app/data/LICENSE app/data/ATTRIBUTION.md app/NOTICE app/data/noitu.db; do
|
||||
grep -qx "$required" files.txt || { echo "missing from the image: $required"; exit 1; }
|
||||
done
|
||||
|
||||
@@ -179,6 +201,9 @@ jobs:
|
||||
sleep 1
|
||||
done
|
||||
curl -fsS http://localhost:8080/healthz
|
||||
# The image has no curl, so the HEALTHCHECK command is the binary
|
||||
# itself; run it the way the container runtime does.
|
||||
docker exec noitu /app/noitu-server -healthcheck
|
||||
# A deep link is a client route, so the binary has to answer it with
|
||||
# the app shell rather than a 404.
|
||||
curl -fsS -o /dev/null -w '%{http_code}\n' 'http://localhost:8080/play?difficulty=2' | grep -qx 200
|
||||
|
||||
@@ -19,19 +19,21 @@ jobs:
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
persist-credentials: false
|
||||
# buf breaking compares against main, which needs real history.
|
||||
fetch-depth: 0
|
||||
|
||||
# A moving version rather than an exact pin, per house rule: updates
|
||||
# arrive on the next run, and a breaking major would surface here before
|
||||
# it could surprise a contributor's own machine.
|
||||
- uses: bufbuild/buf-setup-action@v1
|
||||
# No version input, so the newest buf is used rather than an exact pin,
|
||||
# per house rule: updates arrive on the next run, and a breaking major
|
||||
# would surface here before it could surprise a contributor's own
|
||||
# machine.
|
||||
- uses: bufbuild/buf-action@v1
|
||||
with:
|
||||
version: latest
|
||||
setup_only: true
|
||||
|
||||
- uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version-file: server/go.mod
|
||||
go-version: stable
|
||||
cache-dependency-path: server/go.sum
|
||||
|
||||
- uses: actions/setup-node@v4
|
||||
|
||||
@@ -77,6 +77,7 @@ COPY --from=dict /out/noitu.db /app/data/noitu.db
|
||||
COPY data/LICENSE /app/data/LICENSE
|
||||
COPY data/ATTRIBUTION.md /app/data/ATTRIBUTION.md
|
||||
COPY NOTICE /app/NOTICE
|
||||
COPY LICENSE /app/LICENSE
|
||||
|
||||
ENV NOITU_ADDR=:8080 \
|
||||
NOITU_DB_PATH=/app/data/noitu.db \
|
||||
@@ -84,4 +85,10 @@ ENV NOITU_ADDR=:8080 \
|
||||
|
||||
EXPOSE 8080
|
||||
USER nonroot:nonroot
|
||||
|
||||
# The image has no curl or wget, so the health check is the binary itself:
|
||||
# -healthcheck GETs /healthz on NOITU_ADDR and exits 0 or 1. It is liveness
|
||||
# only; /readyz is the drain signal a load balancer polls.
|
||||
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
|
||||
CMD ["/app/noitu-server", "-healthcheck"]
|
||||
ENTRYPOINT ["/app/noitu-server"]
|
||||
@@ -218,7 +218,8 @@ Endpoints: `GET /ws` (Protobuf over binary WebSocket frames), `GET /healthz`
|
||||
(liveness), `GET /readyz` (readiness — 503 while draining), `GET /version`
|
||||
(plain text), and — when `NOITU_WEB_DIR` is set — the frontend on everything
|
||||
else, with unknown paths falling back to `index.html` because deep links are
|
||||
client routes.
|
||||
client routes. `noitu-server -healthcheck` probes `/healthz` on `NOITU_ADDR` and exits 0 or 1;
|
||||
it is what the image's `HEALTHCHECK` runs, since the image has no `curl`.
|
||||
|
||||
### Smoke-testing without a frontend
|
||||
|
||||
|
||||
+99
-5
@@ -86,9 +86,14 @@ one.
|
||||
### What travels with the data
|
||||
|
||||
The derived wordlist is CC BY-SA 4.0 while the code is Apache-2.0, so the image
|
||||
carries `data/LICENSE`, `data/ATTRIBUTION.md` and `NOTICE` alongside it. CI
|
||||
asserts all three are present, and that the upstream file is not. Removing them
|
||||
would put the image out of compliance.
|
||||
carries `data/LICENSE`, `data/ATTRIBUTION.md`, `NOTICE` and the Apache-2.0
|
||||
`LICENSE` that `NOTICE` refers to. CI asserts all four are present, and that
|
||||
the upstream file is not. Removing them would put the image out of compliance.
|
||||
|
||||
The image does not carry a notice file for the third-party Go and JavaScript
|
||||
dependencies. That is accepted while the image is only built and run by the
|
||||
operator; it must be generated before an image is ever published or handed to
|
||||
someone else.
|
||||
|
||||
## Behind a reverse proxy
|
||||
|
||||
@@ -135,6 +140,39 @@ noitu.example {
|
||||
}
|
||||
```
|
||||
|
||||
### Coolify and Traefik
|
||||
|
||||
Coolify fronts every app with Traefik. With its default forwarded-headers
|
||||
settings Traefik discards any `X-Forwarded-For` the client sent and appends the
|
||||
real peer, so it is safe to trust. Left at the defaults of this server it is
|
||||
not trusted, and the whole player base shares one client address: one join
|
||||
bucket, and no way to switch the per-address connection cap on. Set these on
|
||||
the app's environment variables:
|
||||
|
||||
| Variable | Value | What it changes |
|
||||
|---|---|---|
|
||||
| `NOITU_TRUSTED_PROXIES` | the subnet Traefik reaches the app on, for example `10.0.1.0/24` | `X-Forwarded-For` is believed when the socket peer is inside this range, so join limiting and the per-address cap key on the real player. Find the range with `docker network inspect coolify` (or the app's own network, if Coolify put it on one) and use the narrowest CIDR that contains the Traefik container. Never a range the public can connect from |
|
||||
| `NOITU_MAX_CONNECTIONS_PER_IP` | `32` as a starting point | One address may hold at most this many open sockets. Off (`0`) by default because it is only meaningful once the address above is the real client; 32 leaves room for a household or campus NAT. Without it one host can hold every socket up to `NOITU_MAX_CONNECTIONS` |
|
||||
| `NOITU_DRAIN_TIMEOUT` | below the container stop grace; see "Draining on deploy" | Live games get this long to finish on redeploy instead of ending at once |
|
||||
|
||||
Variables take effect on the next deploy. An entry in `NOITU_TRUSTED_PROXIES`
|
||||
that does not parse is logged as a warning at startup and skipped, so read the
|
||||
startup log after redeploying; a range that is too narrow to contain Traefik
|
||||
fails quietly, back to one shared address.
|
||||
|
||||
If Cloudflare or another CDN sits in front of Traefik, Traefik itself must
|
||||
trust that CDN's ranges (`forwardedHeaders.trustedIPs` on the entrypoint),
|
||||
otherwise it overwrites the header with the CDN edge address and every player
|
||||
behind the same edge shares one bucket again.
|
||||
|
||||
To make Traefik stop routing to an instance that has started draining, give the
|
||||
service a load-balancer health check on `/readyz` through Coolify's custom
|
||||
labels: `traefik.http.services.<service>.loadbalancer.healthcheck.path=/readyz`
|
||||
and `...healthcheck.interval=5s`, using the service name Coolify generated for
|
||||
the app (visible in the container's labels). Without it Traefik keeps sending
|
||||
new players to the old container until it is removed, and they are refused with
|
||||
`server_restarting`.
|
||||
|
||||
### The client's own address
|
||||
|
||||
Rate limiting keys on the client's address, and by default that is the
|
||||
@@ -169,6 +207,14 @@ client past it is disconnected rather than throttled. The defaults are
|
||||
generous for one binary on a small host; lower them if memory is tight,
|
||||
because a room is a goroutine and an engine held for up to its idle window.
|
||||
|
||||
Neither ceiling can be held by one client alone. A socket that has not sent
|
||||
its `Hello` within ten seconds is closed, so an idle socket cannot sit on a
|
||||
connection slot, and the room-creation budget is charged to the client
|
||||
address as well as to the socket, so reconnecting does not refill it. The
|
||||
per-address budget is deliberately wide (a burst of 30, refilling at one room
|
||||
every two seconds), because without `NOITU_TRUSTED_PROXIES` every player
|
||||
behind the proxy shares it.
|
||||
|
||||
A third ceiling, `NOITU_MAX_CONNECTIONS_PER_IP`, bounds how many of those
|
||||
sockets one address may hold at once, and it is off by default. Turning it on
|
||||
is safe only once the client's own address (above) is the real one: behind a
|
||||
@@ -187,13 +233,19 @@ below, expvar always publishes the process's full command line and its
|
||||
runtime memory statistics; that is the standard library's own doing, not
|
||||
something this server adds, and it is the whole reason `/debug/vars` lives on
|
||||
a separate address rather than a route on the public mux one config change
|
||||
could expose. The counters, all prefixed `noitu_`: connections open and
|
||||
could expose. In a container, binding it to `0.0.0.0` (`:6060`) makes it
|
||||
reachable from every other container on the same Docker network, which on
|
||||
Coolify can be the shared proxy network. Bind it to `127.0.0.1:6060`,
|
||||
which only the container itself can reach, or leave it unset; do not put it on
|
||||
a wildcard address on a shared network. The counters, all prefixed `noitu_`: connections open and
|
||||
total; rooms live and total, each split `bot`/`pvp`; games started and
|
||||
finished the same way; words submitted, accepted, and rejected by reason;
|
||||
eliminations by reason; chat lines; join attempts refused, by whether it was
|
||||
the rate limit, an unknown code, or a full room; bot moves by difficulty; dead-
|
||||
end claims by whether the position actually had no legal move; words reported
|
||||
as real by `ReportWord`; and resumes attempted versus succeeded. None of it is
|
||||
as real by `ReportWord`; resumes attempted versus succeeded; and
|
||||
`word_rejected`/`word_reported` log lines dropped by the process-wide
|
||||
rate limit on them (`noitu_corpus_log_suppressed`). None of it is
|
||||
read by the game itself — it is a second write next to a decision already
|
||||
made, not an input to one.
|
||||
|
||||
@@ -234,6 +286,26 @@ rooms, 503 once it has started draining (see below). Point a load balancer's
|
||||
`/healthz`; pointing both at the same endpoint defeats the reason there are
|
||||
two.
|
||||
|
||||
The image has no `curl` or `wget`, so a container health check cannot shell out
|
||||
to one. The binary probes itself instead:
|
||||
|
||||
```sh
|
||||
noitu-server -healthcheck
|
||||
```
|
||||
|
||||
It sends `GET /healthz` to `NOITU_ADDR` (a bare `:8080`, or a wildcard host, is
|
||||
dialled on `127.0.0.1`), and exits 0 on a 200 and 1 on anything else, printing
|
||||
the reason to stderr. The Dockerfile's `HEALTHCHECK` runs it every 30 seconds
|
||||
with a 10-second start period. It is a liveness check, so it deliberately does
|
||||
not follow `/readyz` into a drain.
|
||||
|
||||
On Coolify, leave the dashboard's own HTTP health check switched off: it
|
||||
executes `curl` or `wget` inside the container and cannot pass against this
|
||||
image. Coolify picks the health check up from the Dockerfile after the next
|
||||
deploy (the application then reports a custom health check found), and a
|
||||
rolling update waits for the new container to be healthy before it removes the
|
||||
old one. Confirm that in the deployment log after the first deploy.
|
||||
|
||||
A post-deploy check worth having beyond either is a real socket open, because
|
||||
both health checks pass whether or not the proxy forwards upgrades. Opening
|
||||
the site and starting a game against the bot is the shortest version of that.
|
||||
@@ -259,9 +331,31 @@ plain `kill` used to do that the instant the signal arrived, which is why
|
||||
connected is told the server is restarting and the process shuts down as
|
||||
it always did.
|
||||
|
||||
After the notice goes out the process waits two more seconds, so the writes
|
||||
reach the sockets before it exits, then closes its listeners.
|
||||
|
||||
Each step logs the room and live-game count, so "did the deploy actually
|
||||
wait, and for what" is answered from the log rather than guessed at.
|
||||
|
||||
A second `SIGTERM` or `SIGINT` during a drain is not swallowed: the first
|
||||
signal hands signal handling back to the runtime, so a second one ends the
|
||||
process at once.
|
||||
|
||||
The container runtime, not this server, bounds the whole sequence. Docker
|
||||
sends `SIGTERM`, waits its stop grace period (10 seconds by default), then
|
||||
sends `SIGKILL`, which gives players no `server_restarting` notice and writes
|
||||
no final log line. The rule is
|
||||
|
||||
NOITU_DRAIN_TIMEOUT + 2s < container stop grace
|
||||
|
||||
so with Docker's default grace `NOITU_DRAIN_TIMEOUT` must stay at or below
|
||||
about `6s`. A drain shorter than one turn only helps games in their last
|
||||
seconds; to let games ride out a full `NOITU_TURN_LIMIT` (30 seconds by
|
||||
default), raise the stop grace first, on Coolify wherever the application's
|
||||
container stop timeout is configured, and only then raise the drain timeout.
|
||||
If the grace cannot be raised, keep the drain short instead of letting the
|
||||
kill land mid-drain.
|
||||
|
||||
The default, `NOITU_DRAIN_TIMEOUT=0s`, is today's behaviour: nothing waits,
|
||||
every live game ends immediately. Setting it to something like `60s` turns
|
||||
"deploy when the game is quiet" into "deploy whenever, and the games in
|
||||
|
||||
Reference in new issue
Block a user