mirror of
https://github.com/tiennm99/blog.git
synced 2026-10-11 03:13:10 +00:00
The duplicate-check boundary set had '<' but not '>', so a URL stored as <https://…> (autolink, used by older posts) was never detected as a duplicate. Inherited from the JS engine; found while validating the Go port against 2020-era posts.
3.7 KiB
3.7 KiB
Parity Report — Newsletter Engine JS → Go (2026-08-18)
Method: identical args to node scripts/newsletter/<script>.js and go run ./scripts/newsletter <cmd>, diff on stdout, exit codes compared. Live network cases against real upstreams.
Results
| Case | Result |
|---|---|
| find-newsletter-number | IDENTICAL |
| list-existing-tags (full 40-line ranking) | IDENTICAL |
| detect-image-source: substackcdn wrapper (real URL from posts) | IDENTICAL |
| detect-image-source: raw S3 URL | IDENTICAL |
| detect-image-source: non-Substack image w/ utm param | IDENTICAL |
| detect-image-source: garbage non-URL | IDENTICAL |
| add-url: YouTube watch / youtu.be / shorts (same video id) | IDENTICAL ×3; all canonicalize to same watch URL, oEmbed title+author present |
add-url: stored article + utm_source+fbclid+x=1 |
IDENTICAL; trackers dropped, x=1 kept, duplicate: true |
add-url: prefix probe /p/goclaw-30-mot-buoc vs stored …-ngoat-lon |
IDENTICAL; duplicate: false — lookahead→index-scan rewrite verified |
| add-url: stored S3 image (uuid dedup) | IDENTICAL; route: image, duplicate: true |
add-url: .pdf |
IDENTICAL; route: document |
| find-substack-post: live ByteByteGo RSS uuid | IDENTICAL (hit shape incl. candidates list) |
| find-substack-post: fabricated uuid | IDENTICAL ({found:false}) |
find-substack-post: fabricated uuid --deep |
IDENTICAL (miss shape: scanned/budget/cutoff) |
| fetch-via-defuddle: example.com | exit 0=0, body byte-identical |
| fetch-via-defuddle: unresolvable host | exit 1=1 |
| fetch-via-defuddle: no args | node 2, go-run 1 — see drift |
Accepted Drift
go runcollapses nonzero exits to 1. In-program codes (1 fail / 2 bad-args) are correct in a compiled binary, butgo runreports any child failure as 1. No skill distinguishes 1 vs 2 (verified mt-webfetch SKILL.md acts on stderr/body, not codes). Noted infetch_via_defuddle.goheader.- Entity decoding: Go uses stdlib
html.UnescapeString(superset of the JS hand-rolled map). No divergence observed on live data; Go can only be more correct on exotic entities. clean_urlquery strategy: Go keeps surviving query pairs verbatim (no re-encode) vs JS re-serialization. No divergence observed in the matrix; could differ on already-percent-encoded params. Dedup unaffected (path/identity-based).- list-existing-tags tie order: Go walks lexically (deterministic); Node readdir order is FS-dependent. Full output was identical on current content; a future tie could order differently between engines — moot once JS is deleted.
Verdict
All matrix cases pass. Go engine is behavior-equivalent for skill consumption. Proceed to cutover.
Post-cutover validation addendum (old-post URLs, 2026-08-18)
Ran the Go engine against URLs stored in old posts (2020 and 2025/02):
| Case | Result |
|---|---|
| 2025/02 article, plain + tracker-variant | duplicate:true both, trackers stripped |
| 2025/02 YouTube stored watch URL + youtu.be variant | both canonicalize, duplicate:true, oEmbed title |
| 2025/02 S3 image (add-url + detect-image-source) | route:image, duplicate:true, uuid extracted |
| 18-month-old uuid via find-substack-post | {found:false} — correct, outside RSS window |
| 404 negative control | duplicate:false, http_status 404 |
2020 URL stored as markdown autolink <https://…> |
duplicate:false — defect found |
Defect: the URL boundary set ()]"'?#<_&,, inherited verbatim from the JS) lacked >, so URLs stored in autolink form never dedup. Pre-existing JS behavior, not a port regression. Fixed post-cutover by adding > to the set in url_utils.go; regression-checked prefix-probe (still false), stored full URL (still true), deterministic outputs (unchanged).