feat: split data update from site build; publish via Cloudflare Pages

The updater fetched GitHub metadata and rendered the dashboard payload in
a single step, so anything that published the site needed a GITHUB_TOKEN.
Cloudflare Pages builds from a Git webhook and has no business holding
one.

Split the tool into three modes. The default update step still fetches
GitHub and now records every API-sourced field in data/metadata.json,
which is committed alongside README.md and data/history.jsonl. The new
-build mode joins that snapshot with data/agents.yml and
data/history.jsonl to render dist/ with no network access and no token;
-check is unchanged.

Supporting changes:

- computeDeltaAt anchors the delta windows to when the data was fetched
  rather than the wall clock, so a redeploy days later reproduces the
  same Δ7d instead of sliding the window past its slack allowance.
- writeSiteData takes updatedAt explicitly for the same reason: the
  timestamp labels data freshness, not build time.
- sortStats is extracted from fetchStats so the build step re-ranks
  identically from committed metadata.
- An agents.yml entry with no metadata yet is omitted with a warning
  instead of failing the build, which would otherwise block every deploy
  between merging a new entry and the next nightly run.
- update.yml drops the GitHub Pages deploy steps and commits
  data/metadata.json; ci.yml runs `go run . -build` so a build that would
  break on deploy breaks in CI first.
- site/_headers stops the edge serving a stale data.json after a refresh.
- docs/DEPLOY.md covers the Cloudflare setup, including the one-time
  bootstrap of data/metadata.json that the build depends on.

Also carries the in-progress curation work already in the tree: the
module rename to awesome-ai-dev-tools, retagged entries, and removal of
the archived Roo-Code, void, continue and suna entries.
This commit is contained in:
tiennm99 committed 2026-09-16 22:24:48 +07:00
1 parent 3310ed36d4
commit 00c90c5be9
22 files changed
+728 -116

No files matched your search

+4
View File
@@ -11,6 +11,7 @@ on:
- '.github/workflows/ci.yml'
- 'data/agents.yml'
- 'templates/**'
- 'site/**'
permissions:
contents: read
@@ -30,6 +31,9 @@ jobs:
- run: go test ./...
- run: go build ./...
- run: go run . -check
# Runs exactly what Cloudflare Pages runs, so a build that would fail on
# deploy fails here instead. Needs no token — that is the point of the split.
- run: go run . -build
lint:
runs-on: ubuntu-latest
+12 -29
View File
@@ -8,7 +8,6 @@ on:
branches: [main]
paths:
- 'data/agents.yml'
- 'site/**'
- 'templates/**'
- '**.go'
- 'go.mod'
@@ -18,11 +17,11 @@ permissions:
# Uses GITHUB_TOKEN intentionally — a PAT would cause infinite trigger loops
# on the auto-commit ("chore: daily ranking refresh") because GITHUB_TOKEN-authored
# pushes do NOT re-trigger workflows, whereas a PAT would.
#
# This job only refreshes data. Publishing is Cloudflare Pages' job: it builds
# from the commit this job pushes (webhook-driven, so the GITHUB_TOKEN caveat
# above does not apply to it) by running `go run . -build`, which needs no token.
contents: write
# Pages deploy happens in THIS workflow (a separate push-triggered Pages
# workflow would never fire — see GITHUB_TOKEN note above).
pages: write
id-token: write
concurrency:
group: update-rankings
@@ -31,9 +30,6 @@ concurrency:
jobs:
update:
runs-on: ubuntu-latest
environment:
name: github-pages
url: ${{ steps.deploy.outputs.page_url }}
steps:
- uses: actions/checkout@v7
@@ -42,6 +38,9 @@ jobs:
go-version: 'stable'
cache: true
# The only step that touches the network or needs a token. It refreshes
# README.md, data/history.jsonl and data/metadata.json; rendering the
# site from those files happens later, on Cloudflare.
- name: Run updater
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
@@ -51,7 +50,7 @@ jobs:
run: |
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git add README.md data/history.jsonl
git add README.md data/history.jsonl data/metadata.json
if git diff --staged --quiet; then
echo "no changes"
exit 0
@@ -59,7 +58,7 @@ jobs:
git commit -m "chore: daily ranking refresh"
# Rebase + retry guards against the race where two runs commit near-simultaneously.
# Strategy: plain rebase first; only force-resolve bot-generated files (README.md,
# data/history.jsonl) when they are the sole conflicts. If any human-maintained file
# data/history.jsonl, data/metadata.json) when they are the sole conflicts. If any human-maintained file
# (e.g. data/agents.yml) conflicts, abort and fail loudly so a human can investigate.
for attempt in 1 2 3; do
if git push; then exit 0; fi
@@ -72,31 +71,15 @@ jobs:
conflicted=$(git diff --name-only --diff-filter=U)
echo "conflicted files: $conflicted"
# Only auto-resolve if ALL conflicts are in bot-generated files.
non_bot=$(echo "$conflicted" | grep -v -E '^(README\.md|data/history\.jsonl)$' || true)
non_bot=$(echo "$conflicted" | grep -v -E '^(README\.md|data/history\.jsonl|data/metadata\.json)$' || true)
if [ -n "$non_bot" ]; then
echo "ERROR: conflict in human-maintained file(s): $non_bot — aborting rebase"
git rebase --abort
exit 1
fi
# Safe to force-resolve: keep our (freshest) bot-generated content.
git checkout --theirs README.md data/history.jsonl 2>/dev/null || true
git add README.md data/history.jsonl
git checkout --theirs README.md data/history.jsonl data/metadata.json 2>/dev/null || true
git add README.md data/history.jsonl data/metadata.json
GIT_EDITOR=true git rebase --continue
done
exit 1
# site/data.json was generated by the updater run above; deploy the
# static dashboard (site/) to GitHub Pages every run.
- name: Configure Pages
uses: actions/configure-pages@v6
with:
enablement: true
- name: Upload Pages artifact
uses: actions/upload-pages-artifact@v5
with:
path: site
- name: Deploy to GitHub Pages
id: deploy
uses: actions/deploy-pages@v5