Add 5-phase plan to self-host on Coolify (docker-compose) with a MongoDB Atlas backend, an in-process cron scheduler, DynamoDB->Atlas migration, and full AWS teardown. Red-teamed and validated; decommission scope verified against live AWS/Cloudflare accounts. Include free-tier audit and S3-elimination research reports. Ignore wrangler local cache.
9.9 KiB
title, description, status, priority, branch, tags, blockedBy, blocks, created, createdBy, source
| title | description | status | priority | branch | tags | blockedBy | blocks | created | createdBy | source | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Self-host miti99bot on Coolify with MongoDB Atlas | Add a MongoDB Atlas storage backend + in-process cron, containerize for Coolify docker-compose, migrate existing DynamoDB data. | pending | P2 | feature/selfhosted |
|
2026-06-27T12:01:53.894Z | ck:plan | skill |
Self-host miti99bot on Coolify with MongoDB Atlas
Overview
Run miti99bot as a long-lived container on Coolify (docker-compose) instead of AWS Lambda, using MongoDB Atlas (MONGO_URL + MONGO_DATABASE) instead of DynamoDB. Existing DynamoDB data is migrated into Atlas. At the code level this is additive — a 4th KV backend + a self-host run mode; the DynamoDB/Lambda code path is NOT ripped out (kept for portability and to run the migrator). At the infrastructure level, the deployed AWS stack is fully decommissioned after a verified cutover (Phase 5, validated decision).
Why this is low-risk: storage is already a pluggable KVProvider interface with 3 backends (memory, firestore, dynamodb); adding mongodb follows the exact firestore collection-per-module pattern. Secrets already fall back to plain env vars when *_PARAMETER_NAME is unset, so Coolify env vars need no code change. The only genuine gap is cron: EventBridge Scheduler triggers /cron/{name} today, which does not exist off-AWS.
Phases
| Phase | Name | Status |
|---|---|---|
| 1 | MongoDB Storage Provider | Pending |
| 2 | In-Process Cron Scheduler | Pending |
| 3 | Containerize and Coolify Deploy | Pending |
| 4 | Data Migration and Cutover | Pending |
| 5 | AWS Full Decommission | Pending |
Dependencies
- Phase 2 is independent of Phase 1 (cron touches no storage).
- Phase 3 depends on 1 + 2 (container must boot with mongo + internal cron).
- Phase 4 depends on Phase 1 (Mongo document schema must be final before copying data) and is the last step before flipping the Telegram webhook to the Coolify URL.
- Phase 5 (AWS decommission) depends on Phase 4
--verifypassing — it destroys DynamoDB, so it runs only after data is migrated and the bot is confirmed live on Coolify.
Suggested order: 1 → 2 → 3 → 4 → 5. Phases 1 and 2 can be done in parallel by separate developers (disjoint files).
Architecture Summary
BEFORE (AWS) AFTER (Coolify self-host)
Telegram ──webhook──> Lambda Function URL Telegram ──webhook──> Coolify domain
EventBridge Scheduler ──> /cron/{name} in-process scheduler ──> cron.Handler
DynamoDB (pk=module, sk=key) MongoDB Atlas (db / collection-per-module)
SSM Parameter Store secrets Coolify env vars (plain)
Same Go binary (cmd/server), same HTTP server on :8080, same module framework. Backend + cron + secret source are all selected by env vars at startup.
Acceptance Criteria
KV_PROVIDER=mongodb MONGO_URL=… MONGO_DATABASE=…boots and serves/webhookwith persistent storage.make testgreen; new mongo provider has parity tests with the firestore/dynamodb suites.- Self-hosted container fires the lolschedule daily push at 01:00 UTC (08:00 ICT) without EventBridge.
docker compose up(Coolify) brings the bot live behind a public HTTPS domain;/health check passes.- All existing prod DynamoDB items present in Atlas with identical values; migration is idempotent + verifiable by per-module counts.
- After cutover, ALL miti99bot AWS resources are deleted — CloudFormation stack AND the manually-created SSM secrets, GitHub OIDC role, and SAM S3 bucket;
app=miti99bottag sweep is empty and Cost Explorer trends to $0. .github/workflows/deploy.ymlis removed/disabled somainpushes no longer recreate the AWS stack.
Red Team Review
Session — 2026-06-27
Findings: 15 (15 accepted, 0 rejected) — 3 reviewers (security/secrets, assumptions, failure-modes), all findings carried file:line evidence.
Severity breakdown: 2 Critical, 6 High, 7 Medium.
| # | Finding | Severity | Disposition | Applied To |
|---|---|---|---|---|
| 1 | nil-expected CAS is a live first-write path; $exists:false upsert is wrong primitive → use InsertOne + unique-_id + blocking concurrent test |
Critical | Accept | Phase 1 |
| 2 | Rollback loses ALL post-cutover writes; no reverse path → state true RPO / reverse migrator | Critical | Accept | Phase 4 |
| 3 | buildProvider log line would leak MONGO_URL creds → log only database, never URL |
High | Accept | Phase 1 |
| 4 | Cron dispatcher symbol misdescribed (DispatchScheduled(ctx,name,reg), no timeout/recover/log) → scheduler adds own |
High | Accept | Phase 2 |
| 5 | Dockerfile omits gitSHA → deploynotify silently dead → inject build arg or document disabled |
High | Accept | Phase 3 |
| 6 | Cron double-fire (EventBridge-live + rolling deploy; push non-idempotent) → disable schedule before internal cron + last-push-date guard | High | Accept | Phase 2, Phase 4 |
| 7 | Atlas 0.0.0.0/0 = public DB surface vs project boundary → egress-IP allowlist + least-priv user default |
High | Accept | Phase 3 |
| 8 | Cutover write-loss gap; ambiguous webhook-unset → mandatory deleteWebhook→migrate→verify→setWebhook + pending_update_count |
High | Accept | Phase 4 |
| 9 | Migrator raw UpdateOne diverges from Put encoding → write through Put; store updatedAt int64 |
Medium | Accept | Phase 1, Phase 4 |
| 10 | Stray *_PARAMETER_NAME fatal off-AWS → .env.example lists all six "leave UNSET" + boot criterion |
Medium | Accept | Phase 3 |
| 11 | / is plain text not JSON; -healthcheck flag doesn't exist → Coolify HTTP monitor, drop bogus CMD |
Medium | Accept | Phase 3 |
| 12 | /cron/ redundant double-trigger when internal → leave CRON_SHARED_SECRET unset (route 404) |
Medium | Accept | Phase 3 |
| 13 | Migrator IAM under-specified / admin profile → exact least-priv Scan-only + Scan-based verify, read-only profile |
Medium | Accept | Phase 4 |
| 14 | No Mongo reconnect/health story → document auto-reconnect + pool opts; DB-aware healthcheck or accept trade-off | Medium | Accept | Phase 3 |
| 15 | validateKey rejects / but migrator bypasses it → migrator validateKey each sk, fail loud |
Medium | Accept | Phase 4 |
Verified non-issues (no change needed): List() $gte/$lt range avoids regex injection (sound, mirrors firestore); secrets fallback to plain env is verified correct; .env already gitignored.
Whole-Plan Consistency Sweep
Re-read all phase files after applying findings. Reconciled:
updatedAtstorage type now consistently int64 nanos in Phase 1 doc-shape and Phase 4 migrator/risk (was "BSON datetime").- CAS absent-case is
InsertOneeverywhere (Phase 1 architecture + Phase 4 encoding note); the$exists:falseupsert is removed. - Migrator writes through
MongoKVStore.Putin Phase 4 architecture, risk, and success criteria (no rawUpdateOne). - Health endpoint described as plain text
miti99bot okin Phase 3 requirements, architecture, and risk (was "JSON"); bogus-healthcheckCMD removed from the compose snippet. - Cron dispatcher named
DispatchScheduledin Phase 2 with scheduler-local timeout/recover. - Cutover ordering (disable EventBridge → migrate →
CRON_MODE=internal→setWebhook) consistent across Phase 2 risk and Phase 4 runbook. No unresolved contradictions remain.
Validation Log
Session — 2026-06-27
Verification pass skipped: ## Red Team Review already carries full file:line evidence and no [UNVERIFIED] tags remain. Interview resolved all 8 open questions.
| # | Question | Decision | Affects |
|---|---|---|---|
| 1 | Rollback RPO | Simple short-window cutover; NO reverse migrator. User coordinates users to pause around cutover, so no writes occur mid-migration. | Phase 4 |
| 2 | Decommission AWS | Tear down the SAM stack after --verify passes. No long-term fallback. EventBridge schedule still disabled as an explicit pre-cutover step. Stop the GitHub Actions deploy workflow. |
Phase 4 |
| 3 | Atlas network access | 0.0.0.0/0 (no stable Coolify egress IP). Accepted trade-off: mandatory strong unique password + least-privilege DB user (readWrite on one DB, not admin). Documented as a knowing widening vs DynamoDB's IAM-gated posture. |
Phase 3 |
| 4 | deploynotify | Keep it — inject gitSHA via Dockerfile ARG GIT_SHA + Coolify build-arg. |
Phase 3 |
| 5 | /cron/ HTTP route |
Disable in prod — leave CRON_SHARED_SECRET unset (route 404s); internal scheduler is the sole trigger. |
Phase 2, Phase 3 |
| 6 | Mongo driver | go.mongodb.org/mongo-driver/v2 (current stable); robfig/cron/v3 for the scheduler. |
Phase 1, Phase 2 |
| 7 | Atlas tier | Free M0 (512 MB) — sufficient for the tiny paper-trading KV. | Phase 3 |
| 8 | updatedAt reader |
Confirmed write-only today; store as int64 for cheap parity. No TTL/sort planned. | Phase 1, Phase 4 |
Whole-Plan Consistency Sweep (post-validation)
- Phase 4 rollback/teardown rewritten: AWS torn down after verify; reverse-migrator option removed; rollback framed as "coordinate users, short window" not "keep Lambda N days."
- Phase 3 Atlas networking:
0.0.0.0/0is now the chosen path (was "egress-IP default") with password + least-priv user as hard requirements. - deploynotify
gitSHAinjection and/cron/disabled (CRON_SHARED_SECRETunset) were already the recommended defaults in Phases 2/3 — now confirmed, no contradiction. - No unresolved contradictions remain.
Open Questions
None — all resolved in the Validation Log above.