Files
blog/plans/reports/parity-260818-newsletter-go-migration-report.md
tiennm99 c35a14f9bb fix(scripts): dedup URLs stored in markdown autolink form
The duplicate-check boundary set had '<' but not '>', so a URL stored
as <https://…> (autolink, used by older posts) was never detected as a
duplicate. Inherited from the JS engine; found while validating the Go
port against 2020-era posts.
2026-08-18 21:43:57 +07:00

3.7 KiB
Raw Permalink Blame History

Parity Report — Newsletter Engine JS → Go (2026-08-18)

Method: identical args to node scripts/newsletter/<script>.js and go run ./scripts/newsletter <cmd>, diff on stdout, exit codes compared. Live network cases against real upstreams.

Results

Case Result
find-newsletter-number IDENTICAL
list-existing-tags (full 40-line ranking) IDENTICAL
detect-image-source: substackcdn wrapper (real URL from posts) IDENTICAL
detect-image-source: raw S3 URL IDENTICAL
detect-image-source: non-Substack image w/ utm param IDENTICAL
detect-image-source: garbage non-URL IDENTICAL
add-url: YouTube watch / youtu.be / shorts (same video id) IDENTICAL ×3; all canonicalize to same watch URL, oEmbed title+author present
add-url: stored article + utm_source+fbclid+x=1 IDENTICAL; trackers dropped, x=1 kept, duplicate: true
add-url: prefix probe /p/goclaw-30-mot-buoc vs stored …-ngoat-lon IDENTICAL; duplicate: false — lookahead→index-scan rewrite verified
add-url: stored S3 image (uuid dedup) IDENTICAL; route: image, duplicate: true
add-url: .pdf IDENTICAL; route: document
find-substack-post: live ByteByteGo RSS uuid IDENTICAL (hit shape incl. candidates list)
find-substack-post: fabricated uuid IDENTICAL ({found:false})
find-substack-post: fabricated uuid --deep IDENTICAL (miss shape: scanned/budget/cutoff)
fetch-via-defuddle: example.com exit 0=0, body byte-identical
fetch-via-defuddle: unresolvable host exit 1=1
fetch-via-defuddle: no args node 2, go-run 1 — see drift

Accepted Drift

  1. go run collapses nonzero exits to 1. In-program codes (1 fail / 2 bad-args) are correct in a compiled binary, but go run reports any child failure as 1. No skill distinguishes 1 vs 2 (verified mt-webfetch SKILL.md acts on stderr/body, not codes). Noted in fetch_via_defuddle.go header.
  2. Entity decoding: Go uses stdlib html.UnescapeString (superset of the JS hand-rolled map). No divergence observed on live data; Go can only be more correct on exotic entities.
  3. clean_url query strategy: Go keeps surviving query pairs verbatim (no re-encode) vs JS re-serialization. No divergence observed in the matrix; could differ on already-percent-encoded params. Dedup unaffected (path/identity-based).
  4. list-existing-tags tie order: Go walks lexically (deterministic); Node readdir order is FS-dependent. Full output was identical on current content; a future tie could order differently between engines — moot once JS is deleted.

Verdict

All matrix cases pass. Go engine is behavior-equivalent for skill consumption. Proceed to cutover.

Post-cutover validation addendum (old-post URLs, 2026-08-18)

Ran the Go engine against URLs stored in old posts (2020 and 2025/02):

Case Result
2025/02 article, plain + tracker-variant duplicate:true both, trackers stripped
2025/02 YouTube stored watch URL + youtu.be variant both canonicalize, duplicate:true, oEmbed title
2025/02 S3 image (add-url + detect-image-source) route:image, duplicate:true, uuid extracted
18-month-old uuid via find-substack-post {found:false} — correct, outside RSS window
404 negative control duplicate:false, http_status 404
2020 URL stored as markdown autolink <https://…> duplicate:false — defect found

Defect: the URL boundary set ()]"'?#<_&,, inherited verbatim from the JS) lacked >, so URLs stored in autolink form never dedup. Pre-existing JS behavior, not a port regression. Fixed post-cutover by adding > to the set in url_utils.go; regression-checked prefix-probe (still false), stored full URL (still true), deterministic outputs (unchanged).