The updater fetched GitHub metadata and rendered the dashboard payload in
a single step, so anything that published the site needed a GITHUB_TOKEN.
Cloudflare Pages builds from a Git webhook and has no business holding
one.
Split the tool into three modes. The default update step still fetches
GitHub and now records every API-sourced field in data/metadata.json,
which is committed alongside README.md and data/history.jsonl. The new
-build mode joins that snapshot with data/agents.yml and
data/history.jsonl to render dist/ with no network access and no token;
-check is unchanged.
Supporting changes:
- computeDeltaAt anchors the delta windows to when the data was fetched
rather than the wall clock, so a redeploy days later reproduces the
same Δ7d instead of sliding the window past its slack allowance.
- writeSiteData takes updatedAt explicitly for the same reason: the
timestamp labels data freshness, not build time.
- sortStats is extracted from fetchStats so the build step re-ranks
identically from committed metadata.
- An agents.yml entry with no metadata yet is omitted with a warning
instead of failing the build, which would otherwise block every deploy
between merging a new entry and the next nightly run.
- update.yml drops the GitHub Pages deploy steps and commits
data/metadata.json; ci.yml runs `go run . -build` so a build that would
break on deploy breaks in CI first.
- site/_headers stops the edge serving a stale data.json after a refresh.
- docs/DEPLOY.md covers the Cloudflare setup, including the one-time
bootstrap of data/metadata.json that the build depends on.
Also carries the in-progress curation work already in the tree: the
module rename to awesome-ai-dev-tools, retagged entries, and removal of
the archived Roo-Code, void, continue and suna entries.
The list was already drifting past "coding agents" — ADEs were admitted in
the previous commit. Rather than keep widening criterion 1 one category at a
time, it now describes the actual subject: developer tools built around AI.
The dividing line becomes tools you use vs. building blocks you import, which
keeps libraries, SDKs, model weights and skill collections out.
The star floor drops from a soft "roughly 10,000+" to a hard 1,000:
- enforceStarFloor drops any below-floor entry from the ranking and emits an
::error:: annotation, so the published list can never violate the rule.
Dropping rather than failing keeps one bad entry from blocking the refresh
of every other repo.
- -check cannot catch this (star counts need the API, -check runs offline);
docs/CONTRIBUTING.md says so explicitly.
Side effect: Orkas (1,998 stars) now clears the floor it previously missed.
go run . -check: 43 agents valid. go test ./...: ok.
README.md and site/data.json were written with os.Create/os.WriteFile,
so a write that failed midway left the repo front page truncated, while
history.jsonl already used a temp-file-plus-rename. Extract that pattern
into atomicWriteFile and route all three writers through it.
Also sort snapshots by date when reading history.jsonl: delta windows
pick the newest snapshot inside the window by scan order, which silently
produces wrong deltas if a hand edit or a merge of two concurrent runs
interleaves lines.
Cover the README renderer, which had no test beyond sanitizeCell, and
name the generated paths as constants instead of repeating literals.
Coverage 65.1% -> 75.7%.
- Add pi, OpenHands, warp, gpt-pilot, qwen-code, kilocode, onlook,
dyad, trae-agent, copilot-cli (19 -> 29 tracked repos)
- Update renamed slugs to canonical owners (anomalyco/opencode,
aaif-goose/goose, AntonOsika/gpt-engineer) with history key
migrations so star deltas survive the rename
- Generate site/data.json each run and deploy site/ (interactive
table + star-history chart) to GitHub Pages from the daily
workflow, since GITHUB_TOKEN bot pushes cannot trigger a separate
Pages workflow
GitHub fetcher (github.go):
- add 30s HTTP client timeout (was http.DefaultClient with no bound)
- chunk GraphQL alias requests at 50 repos to stay clear of abuse detection
- abort the run on any partial GraphQL error or missing repo rather than
silently shrinking the README and poisoning the next delta
- retry transient failures (network, 5xx, 429) with 2s/4s/8s backoff
History layer (history.go):
- key snapshots by canonical owner/repo from agents.yml instead of the
rename-resolved NameWithOwner returned by the API; carry a lazy
migration map so existing aaif-goose/goose entries fold into block/goose
on next read with no manual data edit
- tighten the 7d delta window to (cutoff-3d, cutoff] so a missed cron week
no longer mislabels a 90d-old comparison as Delta7d
- replace the snapshots[:0] aliased filter loop with slices.DeleteFunc
- log malformed JSONL lines to stderr with line numbers instead of
silently skipping them
- write history.jsonl atomically via tmp file + rename so a crash
mid-write can no longer truncate accumulated history
Plus collapse a few redundant fmt.Errorf wraps, drop a named Config type
that was used once, inline the single-call sortByStars helper with a
deterministic tiebreaker on canonical key, and use filepath.Base instead
of hand-rolling a basename.
Includes unit tests covering the 7d window edges, canonical-key
migration, atomic write path, malformed-line tolerance, YAML validation,
and markdown cell escaping.
Go updater that fetches AI agent coding tool repo stats via GitHub GraphQL
(batched, one query), sorts by star count, appends a daily snapshot to
data/history.jsonl, and regenerates README.md from templates/readme.tmpl.
Daily workflow at .github/workflows/update.yml refreshes rankings and
commits changes. Seed list in data/agents.yml covers 19 tracked repos.