mirror of
https://github.com/tiennm99/blog.git
synced 2026-10-11 03:13:10 +00:00
refactor(skills): invoke the Node engine and centralise the command reference
Switches every call site to `node scripts/newsletter` and, in the same pass, fixes three structural problems in the skill layer. The seven subcommands were spelled out in five files; they are now enumerated once in docs/newsletter/engine-commands.md, and every other mention is a link except where a skill inlines the one or two commands it actually invokes. The shared post mechanics move out of mt-add-url's directory into docs/newsletter/, so no skill owns another skill's documentation and both runtimes reach it by a repository-relative path. The Codex adapters become symlinks to their canonical counterparts, which removes the parallel frontmatter that could drift. Two skills are renamed for a uniform mt-<verb>-<object> scheme: mt-add-post becomes mt-add-article, since it adds an article to a post, and mt-webfetch becomes mt-fetch-url. The other four already conformed. Note that $mt-add-post and mt-webfetch no longer resolve; no alias directories are left behind, because two names for one skill is the duplication this change removes. Trigger wording in the frontmatter descriptions is left alone apart from the renamed tokens and the corrected fetch-tier summary: those strings are how the runtimes decide whether to invoke a skill. Setup documentation now states the npm ci prerequisite. The engine has dependencies, so unlike its predecessor it does not run on a bare clone.
This commit is contained in:
1 parent
682d5239b8
commit
baab3d443a
19 files changed
+258
-134
No files matched your search
@@ -1,26 +1,26 @@
|
||||
---
|
||||
name: mt-add-post
|
||||
name: mt-add-article
|
||||
description: 'Article handler for the Hugo blog newsletter. Adds a single article/blog URL to the target newsletter post as a main-content entry with the original source title and a Vietnamese summary. Normally invoked by the mt-add-url meta skill after classification, but can be used directly for a known article URL.'
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
`mt-add-post` is the **article handler**: given a clean article/blog URL, it extracts the source title and content, writes a Vietnamese summary, and inserts it into the target newsletter post's main content (today's post unless the user pinned another one). It does **not** classify or route URLs — that is `mt-add-url`'s job. For YouTube/images/other types, use `mt-add-url`.
|
||||
`mt-add-article` is the **article handler**: given a clean article/blog URL, it extracts the source title and content, writes a Vietnamese summary, and inserts it into the target newsletter post's main content (today's post unless the user pinned another one). It does **not** classify or route URLs — that is `mt-add-url`'s job. For YouTube/images/other types, use `mt-add-url`.
|
||||
|
||||
Shared scripts: `scripts/newsletter/`. Shared procedure: `../mt-add-url/references/newsletter-post-mechanics.md` (resolve/create the target post, newsletter numbering, section insertion, language rules) — **follow it** for all post mechanics.
|
||||
Shared scripts: `scripts/newsletter/`. Shared procedure: `docs/newsletter/post-mechanics.md` (resolve/create the target post, newsletter numbering, section insertion, language rules) — **follow it** for all post mechanics.
|
||||
|
||||
## Input
|
||||
|
||||
A clean article URL (passed by `mt-add-url`, or given directly). If a raw URL is provided directly, you may run the classifier to clean/dedup it first:
|
||||
```bash
|
||||
go run ./scripts/newsletter add-url "<url>"
|
||||
node scripts/newsletter add-url "<url>"
|
||||
```
|
||||
Trust `route: article`; skip if `duplicate` or not `accessible`.
|
||||
|
||||
## Workflow
|
||||
|
||||
1. **Post mechanics** — follow `../mt-add-url/references/newsletter-post-mechanics.md` to resolve/create the target post and (if new) get the newsletter number + apply the template.
|
||||
2. **Extract** the original article title and main content. Try `WebFetch` first; if the host blocks it (403, Cloudflare challenge, empty body), work through the `mt-webfetch` fallback chain before giving up. A `accessible: false` from the router is not by itself a reason to skip — many public tech blogs block plain fetches but serve the fallback fetchers fine. Preserve the source title exactly enough to remain recognizable: do not translate/localize it; keep source-language wording, capitalization, punctuation, and proper nouns from metadata (`og:title`, page title, or fetcher frontmatter).
|
||||
1. **Post mechanics** — follow `docs/newsletter/post-mechanics.md` to resolve/create the target post and (if new) get the newsletter number + apply the template.
|
||||
2. **Extract** the original article title and main content. Try `WebFetch` first; if the host blocks it (403, Cloudflare challenge, empty body), work through the `mt-fetch-url` fallback chain before giving up. A `accessible: false` from the router is not by itself a reason to skip — many public tech blogs block plain fetches but serve the fallback fetchers fine. Preserve the source title exactly enough to remain recognizable: do not translate/localize it; keep source-language wording, capitalization, punctuation, and proper nouns from metadata (`og:title`, page title, or fetcher frontmatter).
|
||||
3. **Summarize in Vietnamese** — 1-2 paragraphs, max 300 words, professional tone for junior developers. Summary paragraphs only; do not add a key-points bullet list.
|
||||
4. **Write** the main-content block and insert it **before** the `### Bonus` section (or append at end of file if there is no Bonus yet — do not create an empty Bonus):
|
||||
|
||||
@@ -7,7 +7,7 @@ description: 'Image handler for the Hugo blog newsletter. Adds an image URL to t
|
||||
|
||||
`mt-add-image` is the **image handler**: given an image URL, it resolves a human-readable **label** and inserts `` into the target newsletter post's **Bonus → Images** (today's post unless the user pinned another one). It does not classify/route — that is `mt-add-url`'s job.
|
||||
|
||||
Shared scripts: `scripts/newsletter/`. Shared procedure: `../mt-add-url/references/newsletter-post-mechanics.md` — **follow it** for post find/create, Bonus insertion, and language rules.
|
||||
Shared scripts: `scripts/newsletter/`. Shared procedure: `docs/newsletter/post-mechanics.md` — **follow it** for post find/create, Bonus insertion, and language rules.
|
||||
|
||||
**Label priority:** figure caption → source post title → your typed input.
|
||||
|
||||
@@ -21,13 +21,13 @@ A clean image URL (passed by `mt-add-url`, or given directly).
|
||||
|
||||
### 1. Detect source
|
||||
```bash
|
||||
go run ./scripts/newsletter detect-image-source "<url>"
|
||||
node scripts/newsletter detect-image-source "<url>"
|
||||
```
|
||||
→ `{ original_url, clean_url, isSubstack, uuid?, innerUrl? }`.
|
||||
|
||||
When invoked **directly** (not via `mt-add-url`), first run the router to get accessibility + duplicate status and skip accordingly:
|
||||
```bash
|
||||
go run ./scripts/newsletter add-url "<url>" # expect route:image; skip if duplicate/!accessible
|
||||
node scripts/newsletter add-url "<url>" # expect route:image; skip if duplicate/!accessible
|
||||
```
|
||||
(When dispatched by `mt-add-url`, that check already ran — don't repeat it.)
|
||||
|
||||
@@ -37,11 +37,11 @@ go run ./scripts/newsletter add-url "<url>" # expect route:image; skip if dupl
|
||||
|
||||
Otherwise, find the source post:
|
||||
```bash
|
||||
go run ./scripts/newsletter find-substack-post --uuid <uuid>
|
||||
node scripts/newsletter find-substack-post --uuid <uuid>
|
||||
```
|
||||
- `found: false` → retry with the deeper sitemap crawl. The quick RSS pass only covers the newest posts, so a miss is the normal result for anything older than the current feed window — go straight to `--deep` rather than treating the miss as a dead end. It is slower (fetches posts ~3 months back, capped at 40 fetches total across all publications), so warn the user it may take a while:
|
||||
```bash
|
||||
go run ./scripts/newsletter find-substack-post --uuid <uuid> --deep
|
||||
node scripts/newsletter find-substack-post --uuid <uuid> --deep
|
||||
```
|
||||
On a miss the result reports `scanned` (posts fetched), `budget` (the 40-fetch cap), and `cutoff` (oldest date looked at) — mention how far back it looked.
|
||||
- `found: false` after `--deep` → no source post; go to step 3 (ask) and/or step 4 (add publication).
|
||||
@@ -71,7 +71,7 @@ If no label was detected, use `AskUserQuestion`:
|
||||
If a Substack image wasn't found and the user tells you which publication it's from, offer to append that host to `scripts/newsletter/config/substack-publications.json`, then retry step 2a. This grows coverage for next time.
|
||||
|
||||
### 5. Insert into Bonus → Images
|
||||
Follow `../mt-add-url/references/newsletter-post-mechanics.md` to resolve/create the target post, then add under **Images**:
|
||||
Follow `docs/newsletter/post-mechanics.md` to resolve/create the target post, then add under **Images**:
|
||||
```markdown
|
||||
**Images:**
|
||||

|
||||
|
||||
@@ -50,7 +50,7 @@ Read the full post body. Identify:
|
||||
|
||||
Before generating, run:
|
||||
```bash
|
||||
go run ./scripts/newsletter list-existing-tags
|
||||
node scripts/newsletter list-existing-tags
|
||||
```
|
||||
When a proposed tag matches an existing one case-insensitively, use the existing casing.
|
||||
-->
|
||||
|
||||
@@ -1,14 +1,14 @@
|
||||
---
|
||||
name: mt-add-url
|
||||
description: 'Meta entry for adding URLs to the Hugo blog newsletter. Use whenever the user provides one or more URLs to add to their newsletter (articles, YouTube videos, images, etc.). Classifies each URL and auto-dispatches to the right handler skill (mt-add-post for articles, mt-add-video for YouTube, mt-add-image for images). For unsupported types it asks the user how to proceed. This is the default entry point for newsletter URL processing.'
|
||||
description: 'Meta entry for adding URLs to the Hugo blog newsletter. Use whenever the user provides one or more URLs to add to their newsletter (articles, YouTube videos, images, etc.). Classifies each URL and auto-dispatches to the right handler skill (mt-add-article for articles, mt-add-video for YouTube, mt-add-image for images). For unsupported types it asks the user how to proceed. This is the default entry point for newsletter URL processing.'
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
`mt-add-url` is the **meta dispatcher**: it classifies each URL and auto-invokes the matching handler skill. Handlers (`mt-add-post`, `mt-add-video`, `mt-add-image`) own the actual content writing. Shared scripts live in `scripts/newsletter/`; shared post mechanics in `references/newsletter-post-mechanics.md`.
|
||||
`mt-add-url` is the **meta dispatcher**: it classifies each URL and auto-invokes the matching handler skill. Handlers (`mt-add-article`, `mt-add-video`, `mt-add-image`) own the actual content writing. The shared engine lives in `scripts/newsletter/` — see [docs/newsletter/engine-commands.md](../../../docs/newsletter/engine-commands.md) for every command it offers, and `docs/newsletter/post-mechanics.md` for shared post mechanics.
|
||||
|
||||
**Supported routes (this version):**
|
||||
- `article` → `mt-add-post`
|
||||
- `article` → `mt-add-article`
|
||||
- `youtube` → `mt-add-video`
|
||||
- `image` → `mt-add-image`
|
||||
|
||||
@@ -20,7 +20,7 @@ Everything else (direct `video` file, `document`, or anything unrecognized) is *
|
||||
|
||||
For every URL the user provides:
|
||||
```bash
|
||||
go run ./scripts/newsletter add-url "<url>"
|
||||
node scripts/newsletter add-url "<url>"
|
||||
```
|
||||
Output (JSON): `{ original_url, clean_url, http_status, accessible, duplicate, route, title?, author? }`.
|
||||
|
||||
@@ -30,13 +30,13 @@ Output (JSON): `{ original_url, clean_url, http_status, accessible, duplicate, r
|
||||
### 2. Skip non-actionable URLs
|
||||
|
||||
- `duplicate: true` → skip, note in report (already in a newsletter).
|
||||
- `accessible: false` → **not an automatic skip.** The classifier does a plain fetch, so a bot-blocked host (403, Cloudflare challenge) reports `accessible: false` even when the page is public and the fallback fetchers can read it. Dispatch on `route` as normal and let the handler's fetch chain decide; only report the URL as skipped when every fetcher in `mt-webfetch` has failed. A `404`/dead URL is a genuine skip.
|
||||
- `accessible: false` → **not an automatic skip.** The classifier does a plain fetch, so a bot-blocked host (403, Cloudflare challenge) reports `accessible: false` even when the page is public and the fallback fetchers can read it. Dispatch on `route` as normal and let the handler's fetch chain decide; only report the URL as skipped when every fetcher in `mt-fetch-url` has failed. A `404`/dead URL is a genuine skip.
|
||||
|
||||
### 3. Dispatch on route
|
||||
|
||||
| route | Action |
|
||||
|-------|--------|
|
||||
| `article` | Invoke the **`mt-add-post`** skill, passing `clean_url` |
|
||||
| `article` | Invoke the **`mt-add-article`** skill, passing `clean_url` |
|
||||
| `youtube` | Invoke the **`mt-add-video`** skill, passing `clean_url` |
|
||||
| `image` | Invoke the **`mt-add-image`** skill, passing `clean_url` |
|
||||
| `video` (direct file) / `document` / anything else | **Fallback** — see step 4 |
|
||||
@@ -62,7 +62,7 @@ Act on the user's choice. If they choose add/update, proceed to design that skil
|
||||
Close with the target post's TL;DR tally, then the per-URL detail. Read the tally from the post itself so it reflects everything the post now holds, not just this batch:
|
||||
|
||||
```bash
|
||||
go run ./scripts/newsletter post-stats content/post/YYYY/MM/DD/index.md
|
||||
node scripts/newsletter post-stats content/post/YYYY/MM/DD/index.md
|
||||
```
|
||||
|
||||
Aggregate across all URLs:
|
||||
@@ -73,7 +73,7 @@ Aggregate across all URLs:
|
||||
(omit zero counts; documents too when present)
|
||||
|
||||
✅ Dispatched: [count]
|
||||
- [count] → mt-add-post (articles)
|
||||
- [count] → mt-add-article (articles)
|
||||
- [count] → mt-add-video (YouTube)
|
||||
- [count] → mt-add-image (images)
|
||||
|
||||
@@ -84,9 +84,9 @@ Aggregate across all URLs:
|
||||
- [url] (route: [route]): [user decision]
|
||||
```
|
||||
|
||||
The tally is report-only — never write it into `index.md`. See *Post tally* in `references/newsletter-post-mechanics.md`.
|
||||
The tally is report-only — never write it into `index.md`. See *Post tally* in `docs/newsletter/post-mechanics.md`.
|
||||
|
||||
## Notes
|
||||
|
||||
- Handlers (`mt-add-post`, `mt-add-video`, `mt-add-image`) remain directly invocable for single-purpose use, but `mt-add-url` is the normal entry point when a user pastes a URL.
|
||||
- Shared mechanics (numbering, post find/create, Bonus insertion, language rules) are defined once in `references/newsletter-post-mechanics.md`; handlers reference it.
|
||||
- Handlers (`mt-add-article`, `mt-add-video`, `mt-add-image`) remain directly invocable for single-purpose use, but `mt-add-url` is the normal entry point when a user pastes a URL.
|
||||
- Shared mechanics (numbering, post find/create, Bonus insertion, language rules) are defined once in `docs/newsletter/post-mechanics.md`; handlers reference it.
|
||||
@@ -1,134 +0,0 @@
|
||||
# Newsletter Post Mechanics (shared)
|
||||
|
||||
Shared procedure used by the newsletter handler skills (`mt-add-post`, `mt-add-video`, `mt-add-image`).
|
||||
All shared scripts live in `scripts/newsletter/`.
|
||||
|
||||
**Project Context:**
|
||||
- Hugo static site, theme `hugo-theme-stack`
|
||||
- Language: Vietnamese (99%) + minimal English tech terms (1%)
|
||||
- Timezone: Asia/Ho_Chi_Minh (UTC+7)
|
||||
- Content path: `content/post/YYYY/MM/DD/index.md`
|
||||
|
||||
## 1. Find / create the target post
|
||||
|
||||
The target defaults to today's date but the user can pin a different one.
|
||||
|
||||
**Target date resolution:**
|
||||
- Default: current date in `YYYY-MM-DD` (UTC+7).
|
||||
- **Pinned target:** when the user names a post or date to keep working on — typically an unpublished draft started on an earlier day ("keep adding to the 7/9 post") — that post is the target for the rest of the session, including across a date rollover mid-session. Do not silently start a new post for the new day; if a pin might have lapsed, ask before creating one.
|
||||
|
||||
Check `content/post/YYYY/MM/DD/index.md` for the resolved target date:
|
||||
- **Exists** → update this file.
|
||||
- **Missing** → create it (new newsletter number; template below). Create directories as needed.
|
||||
|
||||
The newsletter number always comes from the target post, not from today's date.
|
||||
|
||||
## 2. Newsletter number
|
||||
|
||||
```bash
|
||||
go run ./scripts/newsletter find-newsletter-number
|
||||
```
|
||||
Searches backwards from today for the most recent newsletter and returns the next number. Only needed when **creating** a new post — when the target post already exists, read its number from its own `title`.
|
||||
|
||||
## 2a. Terminology
|
||||
|
||||
Refer to each numbered publication as a **newsletter**, not an issue. Use wording like `Newsletter #120`, `add newsletter 120`, or `add Newsletter #120` in reports and commit messages.
|
||||
|
||||
## 3. New post template
|
||||
|
||||
```markdown
|
||||
---
|
||||
title: "Newsletter #[number]"
|
||||
date: YYYY-MM-DD
|
||||
tags: ["AI-Assisted"]
|
||||
categories: ["Newsletter"]
|
||||
---
|
||||
|
||||
*Mời bạn thưởng thức Newsletter #[number].*
|
||||
```
|
||||
(Handlers then append their own content block — article body, or Bonus entry.)
|
||||
|
||||
## 4. Section insertion (no clobbering)
|
||||
|
||||
**Never rewrite the whole `index.md`.** Always insert by anchoring an Edit on an existing string (e.g. `### Bonus`, `**Videos:**`) and prepending/appending around it. Multiple URLs targeting the same day's post must be applied **sequentially** so one edit doesn't clobber another.
|
||||
|
||||
**Articles go before the `### Bonus` section; Bonus assets go inside it.**
|
||||
Article headings must use the original source title. Do not translate/localize article titles; keep the source-language wording, capitalization, punctuation, and proper nouns from metadata (`og:title`, page title, or fetcher frontmatter). Summaries stay Vietnamese and are prose paragraphs only — no key-points bullet list.
|
||||
|
||||
To insert an article safely, anchor the Edit on `### Bonus` and prepend:
|
||||
```
|
||||
old_string: "### Bonus"
|
||||
new_string: "## [Original Article Title](clean_url)\n\n[Summary]\n\n### Bonus"
|
||||
```
|
||||
|
||||
If the post has **no `### Bonus`** yet:
|
||||
- Article → append at end of file. Do NOT create an empty Bonus section.
|
||||
- Bonus asset (video/image/doc) → create the `### Bonus` section now (only because there's an asset to put in it).
|
||||
|
||||
**Bonus section format:**
|
||||
```markdown
|
||||
### Bonus
|
||||
|
||||
**Images:**
|
||||

|
||||
|
||||
**Videos:**
|
||||
[Tiêu đề video tiếng Việt](https://www.youtube.com/watch?v=ID)
|
||||
> Tóm tắt 1-2 câu về nội dung video.
|
||||
|
||||
[Video: title](url) <!-- direct video file -->
|
||||
|
||||
**Documents:**
|
||||
[PDF: title](url)
|
||||
```
|
||||
When a subsection (e.g. `**Videos:**`) already exists, append under it; otherwise create it. Keep subsections in this order: **Images** → **Videos** → **Documents**.
|
||||
|
||||
## 4a. Post tally (report line)
|
||||
|
||||
After every successful insertion, report the target post's running totals so the user can see what the post now holds:
|
||||
|
||||
```bash
|
||||
go run ./scripts/newsletter post-stats content/post/YYYY/MM/DD/index.md
|
||||
```
|
||||
Output (JSON): `{ post, newsletter, articles, images, videos, documents, total }`.
|
||||
|
||||
Render it as a single TL;DR line in the handler's report, omitting zero counts:
|
||||
|
||||
```
|
||||
📊 4 articles · 2 videos · 1 image
|
||||
```
|
||||
|
||||
Use the counts the command returns — do not tally by hand. This is a **report-only** line: never write it into `index.md`.
|
||||
|
||||
## 5. Post content language guidelines
|
||||
|
||||
These rules apply only to text written into `content/post/**/index.md`. Keep user-facing questions, status updates, reports, and final responses in English unless the user explicitly requests another language.
|
||||
|
||||
**Primary**: Vietnamese (99%) | **Secondary**: English (1% — unavoidable technical terms only)
|
||||
|
||||
**Allow English when no Vietnamese equivalent exists:**
|
||||
- Technology: AI, API, GitHub, Rust, Go, Docker, Kubernetes
|
||||
- Databases: PostgreSQL, MongoDB, Redis
|
||||
- Companies: Google, Microsoft, Databricks
|
||||
- Acronyms: MCP, CI/CD, JSON, SQL, HTTP
|
||||
|
||||
**Always translate to Vietnamese:**
|
||||
- Common terms: code → mã nguồn, software → phần mềm, deployment → triển khai, performance → hiệu năng
|
||||
- Action verbs: use → sử dụng, check → kiểm tra, build → xây dựng, test → kiểm thử
|
||||
- Any other term: translate if a common Vietnamese equivalent exists
|
||||
|
||||
**Preserve original source language:**
|
||||
- Article headings and image labels must retain the source wording, capitalization, and punctuation.
|
||||
- Do not translate or localize article headings or image labels unless the user explicitly asks.
|
||||
- Video titles follow the `mt-add-video` handler's localization rules.
|
||||
|
||||
**Content requirements:** audience = junior developers; tone = professional, clear, accessible; verify content is ≥99% Vietnamese before saving.
|
||||
|
||||
## 6. Error handling
|
||||
|
||||
| Problem | Action |
|
||||
|-------|--------|
|
||||
| URL inaccessible | Skip, note in report |
|
||||
| Duplicate URL | Skip, note in report |
|
||||
| Extraction fails | Skip, treat as error |
|
||||
| Directory missing | Create directories |
|
||||
@@ -9,7 +9,7 @@ description: 'YouTube video handler for the Hugo blog newsletter. Adds a YouTube
|
||||
|
||||
Scope: **YouTube links only** (`watch`, `youtu.be`, `shorts`). Direct video files (`.mp4` etc.) are not handled here — they go through `mt-add-url`'s fallback.
|
||||
|
||||
Shared scripts: `scripts/newsletter/`. Shared procedure: `../mt-add-url/references/newsletter-post-mechanics.md` — **follow it** for post find/create, numbering, Bonus insertion, and language rules.
|
||||
Shared scripts: `scripts/newsletter/`. Shared procedure: `docs/newsletter/post-mechanics.md` — **follow it** for post find/create, numbering, Bonus insertion, and language rules.
|
||||
|
||||
## Input
|
||||
|
||||
@@ -19,7 +19,7 @@ A clean YouTube URL (passed by `mt-add-url`, or given directly).
|
||||
|
||||
1. **Classify / fetch title** — run the router to get the canonical URL + title:
|
||||
```bash
|
||||
go run ./scripts/newsletter add-url "<url>"
|
||||
node scripts/newsletter add-url "<url>"
|
||||
```
|
||||
Confirm `route: youtube`; skip if `duplicate` or not `accessible`. Use the returned `clean_url` (canonical `watch?v=ID`) and `title`.
|
||||
- If `title` is missing (oEmbed failed), fetch the title via WebFetch on the watch URL.
|
||||
@@ -32,7 +32,7 @@ A clean YouTube URL (passed by `mt-add-url`, or given directly).
|
||||
- Worst case, derive a single sentence from the title.
|
||||
Keep it to 1-2 sentences (KISS).
|
||||
|
||||
4. **Post mechanics** — follow `../mt-add-url/references/newsletter-post-mechanics.md` to resolve/create the target post and locate the Bonus section.
|
||||
4. **Post mechanics** — follow `docs/newsletter/post-mechanics.md` to resolve/create the target post and locate the Bonus section.
|
||||
|
||||
5. **Insert under Bonus → Videos:**
|
||||
```markdown
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
name: mt-webfetch
|
||||
description: "Fallback web content fetchers for pages that blocked the built-in fetch. Use ONLY when the built-in WebFetch tool has already failed with 403 Forbidden, bot detection, Cloudflare challenge, empty content, or similarly blocked response. Tries defuddle.md first, then a reader proxy — both fetch the page server-side from a different IP and return clean markdown. Do NOT use as a first-choice fetcher — try WebFetch first. Does NOT bypass paywalls, login walls, or pages that require JavaScript execution."
|
||||
name: mt-fetch-url
|
||||
description: "Fallback web content fetchers for pages that blocked the built-in fetch. Use ONLY when the built-in WebFetch tool has already failed with 403 Forbidden, bot detection, Cloudflare challenge, empty content, or similarly blocked response. Tries local defuddle extraction, then the defuddle.md proxy, then a reader proxy — the proxies fetch the page server-side from a different IP and return clean markdown. Do NOT use as a first-choice fetcher — try WebFetch first. Does NOT bypass paywalls, login walls, or pages that require JavaScript execution."
|
||||
---
|
||||
|
||||
## Scope
|
||||
@@ -26,29 +26,40 @@ Use this skill only after a WebFetch attempt returned one of:
|
||||
## Workflow
|
||||
|
||||
1. Confirm WebFetch already failed on the target URL
|
||||
2. **Tier 1 — defuddle.** Run the fetch script:
|
||||
2. **Tier 1 — defuddle.** Run the fetch command:
|
||||
```bash
|
||||
go run ./scripts/newsletter fetch-via-defuddle "<target_url>"
|
||||
node scripts/newsletter fetch-via-defuddle "<target_url>"
|
||||
```
|
||||
Alternatively, use WebFetch with the defuddle-prefixed URL:
|
||||
```
|
||||
WebFetch(url: "https://defuddle.md/<target_url>", prompt: "<extraction prompt>")
|
||||
```
|
||||
3. **Tier 2 — reader proxy.** If tier 1 returns an error (commonly `502 / empty body`) or a body with no usable content, try a reader proxy through WebFetch:
|
||||
It is itself two-stage: it extracts locally first, and falls back to the
|
||||
`https://defuddle.md/<target_url>` proxy when local extraction fails or
|
||||
returns nothing. Stderr names which stage failed. A single invocation covers
|
||||
both — do not run it twice.
|
||||
3. **Tier 2 — reader proxy.** If tier 1 exits nonzero or returns a body with no usable content, try a reader proxy through WebFetch:
|
||||
```
|
||||
WebFetch(url: "https://r.jina.ai/<target_url>", prompt: "<extraction prompt>")
|
||||
```
|
||||
Tier 1 and tier 2 fail independently — a site blocking one often still serves the other, so always attempt tier 2 before giving up.
|
||||
4. Parse the returned markdown (tier 1 has YAML frontmatter with title/description/etc.)
|
||||
4. Parse the returned markdown — both tier-1 stages emit YAML frontmatter with title/description/etc. ahead of the body
|
||||
5. If every tier fails, stop and report which tiers were tried and what each returned — do not keep retrying.
|
||||
|
||||
Attempt each tier at most once. The whole chain is: built-in WebFetch → defuddle → reader proxy → report failure.
|
||||
Attempt each tier at most once. The whole chain is: built-in WebFetch → local defuddle → defuddle.md → reader proxy → report failure.
|
||||
|
||||
## How defuddle works
|
||||
|
||||
- URL pattern: `https://defuddle.md/<target_url>` (target URL appended as path, works with or without scheme)
|
||||
- Returns: Markdown body with YAML frontmatter containing metadata (title, author, description, site name)
|
||||
- Server-side HTTP fetch from defuddle's IP + extraction via Defuddle library (clean main-content extraction)
|
||||
- **Local stage** — the engine fetches the page itself and runs the Defuddle
|
||||
library in-process, so the chain does not depend on a third-party service
|
||||
being up. Returns YAML frontmatter (title, author, description, site,
|
||||
published, source) followed by the markdown body. An extraction that looks
|
||||
like a bot challenge rather than the page counts as a failure here, so the
|
||||
proxy stage still runs.
|
||||
- **Proxy stage** — `https://defuddle.md/<target_url>` (target URL appended as
|
||||
path, works with or without scheme). Returns the same shape: YAML frontmatter
|
||||
(title, author, description, site name) plus a markdown body. This stage fetches from
|
||||
defuddle's IP, which is what matters when *this* machine's IP is the blocked
|
||||
one — that is why it is kept behind the local stage rather than replaced by it.
|
||||
|
||||
Exit codes: `0` content returned, `1` both stages failed, `2` bad arguments. See
|
||||
[docs/newsletter/engine-commands.md](../../../docs/newsletter/engine-commands.md).
|
||||
|
||||
## Output handling
|
||||
|
||||
@@ -61,7 +72,7 @@ Give up once every tier has been tried once. Treat these as a tier failure and m
|
||||
- Empty markdown body
|
||||
- Only frontmatter with no body
|
||||
|
||||
When the last tier fails, report it plainly — name each tier and its result (e.g. "WebFetch 403, defuddle 502, reader proxy empty") so the user can decide whether to paste the text or supply another source. If a fetcher fails repeatedly across sessions for the same host family, say so: that is a signal to reorder or extend the chain, not to keep retrying.
|
||||
When the last tier fails, report it plainly — name each tier and its result (e.g. "WebFetch 403, defuddle exit 1, reader proxy empty") so the user can decide whether to paste the text or supply another source. If a fetcher fails repeatedly across sessions for the same host family, say so: that is a signal to reorder or extend the chain, not to keep retrying.
|
||||
|
||||
Never loop. Never retry a tier more than once.
|
||||
|
||||
@@ -78,10 +89,10 @@ Never loop. Never retry a tier more than once.
|
||||
```
|
||||
User wanted to extract content from https://example.com/article
|
||||
WebFetch returned: "Request failed with status code 403"
|
||||
→ Trigger mt-webfetch
|
||||
→ Tier 1: go run ./scripts/newsletter fetch-via-defuddle "https://example.com/article"
|
||||
→ success: parse markdown output, summarize as usual
|
||||
→ "502 / empty body": continue
|
||||
→ Trigger mt-fetch-url
|
||||
→ Tier 1: node scripts/newsletter fetch-via-defuddle "https://example.com/article"
|
||||
→ success (exit 0): parse markdown output, summarize as usual
|
||||
→ exit 1 after both stages: continue
|
||||
→ Tier 2: WebFetch(url: "https://r.jina.ai/https://example.com/article", prompt: ...)
|
||||
→ success: parse markdown output, summarize as usual
|
||||
→ failure: report "WebFetch 403, defuddle 502, reader proxy failed" and stop
|
||||
Reference in new issue
Block a user