fix(github): split a saturated quarter by month to recover dropped repos

Quarterly windows were not enough for the busiest quarter: a live run
warned that 2026-04-01..2026-06-30 still came back at the 100-repo
ceiling, so its seed list was truncated and those repos never got
probed for commit history.

Re-ask a saturated quarter month by month and merge only the repo
lists. The quarter has already contributed its calendar days and commit
totals, so the month passes skip that fold; verified by forcing the
ceiling down to 20 and confirming all six calendar-derived cards stay
byte-identical. Warn only when a month itself saturates, since there is
no finer window left to ask for.
This commit is contained in:
tiennm99 committed 2026-08-13 13:34:08 +07:00
1 parent 7832842c5d
commit 2c3667d28c
4 files changed
+104 -28

No files matched your search

+2 -2
View File
@@ -169,7 +169,7 @@ ghstats -user tiennm99 -themes dracula,github_dark -tz Asia/Saigon -out output
## How attribution works
**Repo sampling** uses a seed list built from `contributionsCollection.commitContributionsByRepository`, unioned across every active contribution year. This catches every repo you've committed in — not just your top-starred ones. Each year is queried a quarter at a time: the API caps that field at 100 repos per query and drops the rest without saying so, which a prolific year hits easily.
**Repo sampling** uses a seed list built from `contributionsCollection.commitContributionsByRepository`, unioned across every active contribution year. This catches every repo you've committed in — not just your top-starred ones. Each year is queried a quarter at a time: the API caps that field at 100 repos per query and drops the rest without saying so, which a prolific year hits easily. A quarter that still comes back at the cap is re-asked month by month to recover the tail.
**Which repos count where.** The commit-driven cards (most-commit-language, productive time, productive weekday, and everything derived from the contribution calendar) cover repos in *any* namespace you committed to — your own, your orgs', and upstream repos you sent PRs to. The repo-driven cards (stars, repo count, repos-per-language, top-starred) look only at repos you own. Set `include_org_repos` / `-include-org-repos` to also count org-owned repos where your permission is `ADMIN`; org repos you merely have read or write access to are never counted.
@@ -179,7 +179,7 @@ ghstats -user tiennm99 -themes dracula,github_dark -tz Asia/Saigon -out output
- For per-file accuracy, a future `-accurate-languages` mode is planned (per-commit REST + go-enry).
**Cost per run** (current defaults, typical user):
- ~1 profile query + ~4 queries per active year + ~50 commit-history pages ≈ **80-100 GraphQL calls**.
- ~1 profile query + ~4 queries per active year (+3 for any quarter that saturates) + ~50 commit-history pages ≈ **80-100 GraphQL calls**.
- Zero REST calls. Well under the 5000 points/hr budget.
## Themes