Add 5-phase plan to self-host on Coolify (docker-compose) with a MongoDB Atlas backend, an in-process cron scheduler, DynamoDB->Atlas migration, and full AWS teardown. Red-teamed and validated; decommission scope verified against live AWS/Cloudflare accounts. Include free-tier audit and S3-elimination research reports. Ignore wrangler local cache.
7.8 KiB
phase, title, status, priority, dependencies, effort
| phase | title | status | priority | dependencies | effort |
|---|---|---|---|---|---|
| 2 | In-Process Cron Scheduler | pending | P1 | S |
Phase 2: In-Process Cron Scheduler
Overview
Off AWS there is no EventBridge Scheduler to hit /cron/{name}. Add an opt-in in-process scheduler that reads each registered cron's existing Cron.Schedule field and fires Cron.Handler on time. Gated by CRON_MODE=internal so the Lambda/EventBridge path is unchanged by default.
Requirements
- Functional: when
CRON_MODE=internal, parse everyreg.Crons()[i].Scheduleand invoke its handler on schedule, in UTC (match EventBridge'sScheduleExpressionTimezone: UTC). - Functional: default (
CRON_MODEunset/external) starts NO scheduler — preserves current Lambda behavior where EventBridge owns timing. - Functional: the existing
/cron/{name}HTTP route stays as-is (still usable for manual/curl triggers and as the sidecar fallback). - Non-functional: each tick runs under the same 60s budget as the HTTP path (
defaultCronTimeout); a panicking handler is recovered and logged, scheduler keeps running. - Non-functional: scheduler stops cleanly on
rootCtxcancellation (graceful shutdown).
Architecture
The Cron struct already carries Schedule string (currently "documentation only", e.g. lolschedule "0 1 * * *"). Promote it to the real trigger source for self-host.
Use github.com/robfig/cron/v3 (standard, well-maintained 5-field cron parser) with cron.New(cron.WithLocation(time.UTC)). For each registered cron, c.AddFunc(schedule, fn) where fn dispatches the handler through the existing cron dispatcher path so logging/metrics/timeout/panic-recovery are identical to the HTTP route.
Dispatch: the existing exported helper is modules.DispatchScheduled(ctx, name, reg) (cron_dispatcher.go:19 — note arg order: name BEFORE reg). It does ONLY a registry lookup + cron.Handler(ctx, deps) — it has no timeout, no panic-recovery, no structured logging. That wrapping lives in the server-package cronHandler (internal/server/router.go), NOT in the dispatcher. So the scheduler must add its own:
- wrap each fire in
context.WithTimeout(ctx, 60s)(thedefaultCronTimeoutconstant is ininternal/server/timeouts.go; do not importinternal/serverinto the scheduler — define a localcronTimeout = 60*time.Secondor lift the constant to a shared package to avoid a layering inversion), recover()around the handler call, logging the cron name on panic so one bad cron doesn't kill the scheduler,- structured
log.Info("cron triggered", …)/log.Error("cron failed", …)mirroring the HTTP path. Callmodules.DispatchScheduled(ctx, name, reg)inside that wrapper.
New file internal/cron/scheduler.go (new small package, mirrors internal/metrics lifecycle style):
func Run(ctx context.Context, reg *modules.Registry) (stop func(), err error)
- Skips crons with empty
Schedule(logs a warning — a self-hosted cron with no schedule never fires). - Validates schedule strings at startup; a bad expression is fatal (fail fast, like other config errors).
Wire in cmd/server/main.go after modules.Install, gated:
if strings.EqualFold(cfg.CronMode, "internal") {
stop, err := cron.Run(rootCtx, reg)
if err != nil { log.Fatal("cron scheduler init failed", "err", err) }
defer stop()
log.Info("internal cron scheduler started", "crons", len(reg.Crons()))
}
Add CronMode to config from env CRON_MODE.
Related Code Files
- Create:
internal/cron/scheduler.go— scheduler lifecycle. - Create:
internal/cron/scheduler_test.go— fake registry with a fast schedule (@every 1sor injected clock) asserts handler fires; bad-schedule errors; ctx-cancel stops. - Modify:
cmd/server/main.go—CronModeconfig + gatedcron.Run. - Modify:
internal/modules/module.godoc comments — updateCronHandler/Cron.Scheduletext that currently says "real schedule lives in EventBridge" to note theCRON_MODE=internalpath. (The scheduler calls the existingmodules.DispatchScheduled; if the 60s-timeout/panic-recover wrapper is worth sharing with the HTTPcronHandler, lift it to a shared helper — but that is optional, not required.) - Modify:
go.mod/go.sum— addgithub.com/robfig/cron/v3. - Modify:
README.md— documentCRON_MODE(external default vs internal self-host). - Modify:
internal/modules/lolschedule/cron.go— add an idempotency guard: the daily-push handler reads/writes a KV "last push UTC date" key and no-ops if already pushed today (defends against all double-fire windows). This also makes the existing EventBridge path safe during cutover overlap.
Implementation Steps
go get github.com/robfig/cron/v3.- Call
modules.DispatchScheduled(ctx, name, reg)(cron_dispatcher.go:19) from inside a scheduler-local wrapper that adds the 60s timeout +recover()+ logging (the dispatcher provides none of these). - Write
internal/cron/scheduler.go: buildcron.New(cron.WithLocation(time.UTC)), register each non-empty schedule,c.Start(), return astopthat callsc.Stop()and waits for the context done. - Wire gated startup in
main.go; addCronModeconfig field + env read. - Update the now-stale "documentation only / EventBridge owns timing" comments on
Cron.ScheduleandCronHandler. - Tests +
make vet && make test.
Success Criteria
- With
CRON_MODE=internal, lolschedule daily push fires at 01:00 UTC; observable in logs (cron triggered). - With
CRON_MODEunset, no scheduler starts (Lambda/EventBridge path byte-for-byte unchanged). - Bad schedule string fails startup with a clear error.
- Handler panic is recovered (scheduler-local
recover()); scheduler survives and fires next tick. - Each fire runs under a 60s timeout (scheduler-local, not imported from
internal/server). - Daily push is idempotent per UTC date: invoking the handler twice on the same date sends subscribers exactly one digest.
- Scheduler stops within shutdown grace period on SIGTERM.
Risk Assessment
- Double-fire (Critical — the daily push is NOT idempotent):
lolschedulerunDailyPush(cron.go:128-191) fans out to every subscriber unconditionally — no "already sent today" marker. So ANY double-fire = every subscriber DM'd twice. Three concrete double-fire windows the opt-in flag does NOT cover:- Cutover overlap: EventBridge
AWS::Scheduler::Schedule(template.yaml:271-289) invokes the Lambda DIRECTLY, independent of the webhook URL — re-pointing the webhook does NOT stop it. It must be disabled/deleted before the Coolify container runsCRON_MODE=internal(ordered prerequisite in Phase 4, not "N days later" cleanup). - Rolling deploy: Coolify/compose may run old+new containers briefly; both run the in-process scheduler. A redeploy near 01:00 UTC double-fires.
- Operator misconfig (
internalset on a second instance). Primary mitigation (covers all three cheaply): add a KV "last push date" guard inlolschedule— the handler records the date it pushed and no-ops if already pushed for that UTC date. This makes the push idempotent regardless of trigger count. Secondary: prefer stop-first redeploy in Coolify; keepCRON_MODEopt-in and never setinternalon Lambda.
- Cutover overlap: EventBridge
- 5-field vs 6-field cron:
"0 1 * * *"is 5-field standard. Mitigation: use robfigcron/v3default 5-field parser (not the seconds-enabled one). - Single-instance assumption: if Coolify scales the service to >1 replica, crons fire per replica. Mitigation: document "run exactly 1 replica" (the bot is a single-instance webhook consumer anyway); revisit with a DB lock only if scaling is ever needed (YAGNI now).
- Missed fire while container restarts: a deploy at 01:00 UTC could skip that day's push. Mitigation: accept (same risk class as Lambda cold-start miss); not data-loss.