refactor(skills): fix dedup edge cases and modularize mt-* scripts

- fix checkDuplicate false-negative on Substack cover images (uuid with no size suffix)
- stop cleanUrl from corrupting the query string on unparseable input
- extract shared fetchWithTimeout into url-utils (consistent timer cleanup)
- split find-substack-post html/rss helpers into html-text-utils
- scan content/post for duplicates, matching the other scripts
- correct doc drift in mt-add-image/mt-add-url and drop plan-phase code comment
This commit is contained in:
tiennm99 committed 2026-05-30 20:31:50 +07:00
1 parent eada44d407
commit d262159742
6 files changed
+150 -127

No files matched your search

+3 -3
View File
@@ -21,7 +21,7 @@ A clean image URL (passed by `mt-add-url`, or given directly).
```bash
node .claude/skills/mt-add-image/scripts/detect-image-source.js "<url>"
```
→ `{ isSubstack, uuid?, innerUrl? }`.
→ `{ original_url, clean_url, isSubstack, uuid?, innerUrl? }`.
When invoked **directly** (not via `mt-add-url`), first run the router to get accessibility + duplicate status and skip accordingly:
```bash
@@ -34,11 +34,11 @@ Find the source post:
```bash
node .claude/skills/mt-add-image/scripts/find-substack-post.js --uuid <uuid>
```
- `found: false` → retry with the deeper sitemap crawl (slower — scans ~3 months; warn the user it may take ~15s):
- `found: false` → retry with the deeper sitemap crawl (slower — fetches posts ~3 months back, capped at 40 fetches total across all publications; warn the user it may take a while):
```bash
node .claude/skills/mt-add-image/scripts/find-substack-post.js --uuid <uuid> --deep
```
The result reports `scanned` + `cutoff` — mention how far back it looked.
On a miss the result reports `scanned` (posts fetched), `budget` (the 40-fetch cap), and `cutoff` (oldest date looked at) — mention how far back it looked.
- `found: false` after `--deep` → no source post; go to step 3 (ask) and/or step 4 (add publication).
**When `found: true` — pick the label (confirm-from-candidates):**