Files
tiennm99bot/plans/260627-1849-selfhost-coolify-mongodb/plan.md
T
tiennm99 b1171e7cbe docs(selfhost): plan Coolify + MongoDB Atlas self-host with AWS decommission
Add 5-phase plan to self-host on Coolify (docker-compose) with a MongoDB
Atlas backend, an in-process cron scheduler, DynamoDB->Atlas migration,
and full AWS teardown. Red-teamed and validated; decommission scope
verified against live AWS/Cloudflare accounts. Include free-tier audit
and S3-elimination research reports. Ignore wrangler local cache.
2026-06-27 20:46:43 +07:00

9.9 KiB

title, description, status, priority, branch, tags, blockedBy, blocks, created, createdBy, source
title description status priority branch tags blockedBy blocks created createdBy source
Self-host miti99bot on Coolify with MongoDB Atlas Add a MongoDB Atlas storage backend + in-process cron, containerize for Coolify docker-compose, migrate existing DynamoDB data. pending P2 feature/selfhosted
selfhost
coolify
mongodb
migration
2026-06-27T12:01:53.894Z ck:plan skill

Self-host miti99bot on Coolify with MongoDB Atlas

Overview

Run miti99bot as a long-lived container on Coolify (docker-compose) instead of AWS Lambda, using MongoDB Atlas (MONGO_URL + MONGO_DATABASE) instead of DynamoDB. Existing DynamoDB data is migrated into Atlas. At the code level this is additive — a 4th KV backend + a self-host run mode; the DynamoDB/Lambda code path is NOT ripped out (kept for portability and to run the migrator). At the infrastructure level, the deployed AWS stack is fully decommissioned after a verified cutover (Phase 5, validated decision).

Why this is low-risk: storage is already a pluggable KVProvider interface with 3 backends (memory, firestore, dynamodb); adding mongodb follows the exact firestore collection-per-module pattern. Secrets already fall back to plain env vars when *_PARAMETER_NAME is unset, so Coolify env vars need no code change. The only genuine gap is cron: EventBridge Scheduler triggers /cron/{name} today, which does not exist off-AWS.

Phases

Phase Name Status
1 MongoDB Storage Provider Pending
2 In-Process Cron Scheduler Pending
3 Containerize and Coolify Deploy Pending
4 Data Migration and Cutover Pending
5 AWS Full Decommission Pending

Dependencies

  • Phase 2 is independent of Phase 1 (cron touches no storage).
  • Phase 3 depends on 1 + 2 (container must boot with mongo + internal cron).
  • Phase 4 depends on Phase 1 (Mongo document schema must be final before copying data) and is the last step before flipping the Telegram webhook to the Coolify URL.
  • Phase 5 (AWS decommission) depends on Phase 4 --verify passing — it destroys DynamoDB, so it runs only after data is migrated and the bot is confirmed live on Coolify.

Suggested order: 1 → 2 → 3 → 4 → 5. Phases 1 and 2 can be done in parallel by separate developers (disjoint files).

Architecture Summary

                 BEFORE (AWS)                         AFTER (Coolify self-host)
   Telegram ──webhook──> Lambda Function URL    Telegram ──webhook──> Coolify domain
   EventBridge Scheduler ──> /cron/{name}        in-process scheduler ──> cron.Handler
   DynamoDB (pk=module, sk=key)                  MongoDB Atlas (db / collection-per-module)
   SSM Parameter Store secrets                   Coolify env vars (plain)

Same Go binary (cmd/server), same HTTP server on :8080, same module framework. Backend + cron + secret source are all selected by env vars at startup.

Acceptance Criteria

  • KV_PROVIDER=mongodb MONGO_URL=… MONGO_DATABASE=… boots and serves /webhook with persistent storage.
  • make test green; new mongo provider has parity tests with the firestore/dynamodb suites.
  • Self-hosted container fires the lolschedule daily push at 01:00 UTC (08:00 ICT) without EventBridge.
  • docker compose up (Coolify) brings the bot live behind a public HTTPS domain; / health check passes.
  • All existing prod DynamoDB items present in Atlas with identical values; migration is idempotent + verifiable by per-module counts.
  • After cutover, ALL miti99bot AWS resources are deleted — CloudFormation stack AND the manually-created SSM secrets, GitHub OIDC role, and SAM S3 bucket; app=miti99bot tag sweep is empty and Cost Explorer trends to $0.
  • .github/workflows/deploy.yml is removed/disabled so main pushes no longer recreate the AWS stack.

Red Team Review

Session — 2026-06-27

Findings: 15 (15 accepted, 0 rejected) — 3 reviewers (security/secrets, assumptions, failure-modes), all findings carried file:line evidence. Severity breakdown: 2 Critical, 6 High, 7 Medium.

# Finding Severity Disposition Applied To
1 nil-expected CAS is a live first-write path; $exists:false upsert is wrong primitive → use InsertOne + unique-_id + blocking concurrent test Critical Accept Phase 1
2 Rollback loses ALL post-cutover writes; no reverse path → state true RPO / reverse migrator Critical Accept Phase 4
3 buildProvider log line would leak MONGO_URL creds → log only database, never URL High Accept Phase 1
4 Cron dispatcher symbol misdescribed (DispatchScheduled(ctx,name,reg), no timeout/recover/log) → scheduler adds own High Accept Phase 2
5 Dockerfile omits gitSHA → deploynotify silently dead → inject build arg or document disabled High Accept Phase 3
6 Cron double-fire (EventBridge-live + rolling deploy; push non-idempotent) → disable schedule before internal cron + last-push-date guard High Accept Phase 2, Phase 4
7 Atlas 0.0.0.0/0 = public DB surface vs project boundary → egress-IP allowlist + least-priv user default High Accept Phase 3
8 Cutover write-loss gap; ambiguous webhook-unset → mandatory deleteWebhook→migrate→verify→setWebhook + pending_update_count High Accept Phase 4
9 Migrator raw UpdateOne diverges from Put encoding → write through Put; store updatedAt int64 Medium Accept Phase 1, Phase 4
10 Stray *_PARAMETER_NAME fatal off-AWS → .env.example lists all six "leave UNSET" + boot criterion Medium Accept Phase 3
11 / is plain text not JSON; -healthcheck flag doesn't exist → Coolify HTTP monitor, drop bogus CMD Medium Accept Phase 3
12 /cron/ redundant double-trigger when internal → leave CRON_SHARED_SECRET unset (route 404) Medium Accept Phase 3
13 Migrator IAM under-specified / admin profile → exact least-priv Scan-only + Scan-based verify, read-only profile Medium Accept Phase 4
14 No Mongo reconnect/health story → document auto-reconnect + pool opts; DB-aware healthcheck or accept trade-off Medium Accept Phase 3
15 validateKey rejects / but migrator bypasses it → migrator validateKey each sk, fail loud Medium Accept Phase 4

Verified non-issues (no change needed): List() $gte/$lt range avoids regex injection (sound, mirrors firestore); secrets fallback to plain env is verified correct; .env already gitignored.

Whole-Plan Consistency Sweep

Re-read all phase files after applying findings. Reconciled:

  • updatedAt storage type now consistently int64 nanos in Phase 1 doc-shape and Phase 4 migrator/risk (was "BSON datetime").
  • CAS absent-case is InsertOne everywhere (Phase 1 architecture + Phase 4 encoding note); the $exists:false upsert is removed.
  • Migrator writes through MongoKVStore.Put in Phase 4 architecture, risk, and success criteria (no raw UpdateOne).
  • Health endpoint described as plain text miti99bot ok in Phase 3 requirements, architecture, and risk (was "JSON"); bogus -healthcheck CMD removed from the compose snippet.
  • Cron dispatcher named DispatchScheduled in Phase 2 with scheduler-local timeout/recover.
  • Cutover ordering (disable EventBridge → migrate → CRON_MODE=internal → setWebhook) consistent across Phase 2 risk and Phase 4 runbook. No unresolved contradictions remain.

Validation Log

Session — 2026-06-27

Verification pass skipped: ## Red Team Review already carries full file:line evidence and no [UNVERIFIED] tags remain. Interview resolved all 8 open questions.

# Question Decision Affects
1 Rollback RPO Simple short-window cutover; NO reverse migrator. User coordinates users to pause around cutover, so no writes occur mid-migration. Phase 4
2 Decommission AWS Tear down the SAM stack after --verify passes. No long-term fallback. EventBridge schedule still disabled as an explicit pre-cutover step. Stop the GitHub Actions deploy workflow. Phase 4
3 Atlas network access 0.0.0.0/0 (no stable Coolify egress IP). Accepted trade-off: mandatory strong unique password + least-privilege DB user (readWrite on one DB, not admin). Documented as a knowing widening vs DynamoDB's IAM-gated posture. Phase 3
4 deploynotify Keep it — inject gitSHA via Dockerfile ARG GIT_SHA + Coolify build-arg. Phase 3
5 /cron/ HTTP route Disable in prod — leave CRON_SHARED_SECRET unset (route 404s); internal scheduler is the sole trigger. Phase 2, Phase 3
6 Mongo driver go.mongodb.org/mongo-driver/v2 (current stable); robfig/cron/v3 for the scheduler. Phase 1, Phase 2
7 Atlas tier Free M0 (512 MB) — sufficient for the tiny paper-trading KV. Phase 3
8 updatedAt reader Confirmed write-only today; store as int64 for cheap parity. No TTL/sort planned. Phase 1, Phase 4

Whole-Plan Consistency Sweep (post-validation)

  • Phase 4 rollback/teardown rewritten: AWS torn down after verify; reverse-migrator option removed; rollback framed as "coordinate users, short window" not "keep Lambda N days."
  • Phase 3 Atlas networking: 0.0.0.0/0 is now the chosen path (was "egress-IP default") with password + least-priv user as hard requirements.
  • deploynotify gitSHA injection and /cron/ disabled (CRON_SHARED_SECRET unset) were already the recommended defaults in Phases 2/3 — now confirmed, no contradiction.
  • No unresolved contradictions remain.

Open Questions

None — all resolved in the Validation Log above.