Two README statements were false rather than merely incomplete. "--dry-run takes no lock at all, so it can always be run against an export that is currently in progress" contradicted the paragraph four lines above it warning that two clients sharing one session is a corruption hazard: dry-run opens the session file like any other run. Both places now say what is true - no *export* lock, but a session lock, so an alongside dry-run needs its own --session. The exit table promised 1 for a bad argument. argparse exits 2 on its own, before main()'s try block can map anything, so a bad --limit collided with the dry-run SHORT verdict. Documented as the shared code it is instead of a code the tool never returns. Also notes that both locks are flock-based, and that flock is advisory or per-client on NFS, so single-instance enforcement is not guaranteed there. Setup gains the venv creation step it assumed. Phase 7 gains step 12b: measure the sweep cost under a date filter before deciding whether to optimize it. With reverse=True and no offset_date, telethon starts the sweep at message id 1 - verified in the pinned wheel - so --since walks the whole history discarding messages in keep(), and --until never terminates early. Neither fix is free: offset_date can start the sweep mid-album, which would take a post_id that is not the album's lowest id and break filter-invariant identity, and an --until break assumes dates rise with ids, false for imported history. The swept-to-kept ratio decides whether either risk is worth taking. Reports from the review pass are recorded under plans/reports/, including the findings left unfixed: resolve_entity catching only ValueError, _find_in_dialogs being the one network loop outside a flood primitive, SystemExit in session.py routing around the Abort contract, and completed_at being cleared before any work begins.
8.9 KiB
phase, title, status, priority, dependencies, effort
| phase | title | status | priority | dependencies | effort | |
|---|---|---|---|---|---|---|
| 7 | End-to-End Validation | pending | P1 |
|
~2h + export runtime |
Phase 7: End-to-End Validation
Overview
Prove the tool against a real group, with deliberate interruption. The failure modes that matter — cursor ahead of reality, split albums, flood-wait escalation — are invisible in unit tests and surface only against live data.
Every step below has a checkable outcome. No step may pass by escape clause.
Requirements
- Functional: full export completes, resumes correctly after kill, produces a consistent sidecar
- Non-functional: README documents setup and semantics; no secret or third-party data is committable; the spike's server-side session is revoked
Validation Script
-
Dry run —
tg-export --group <g> --out ./exports --dry-run. Record posts, files, total bytes, verdict. Confirm no export root was created. -
Limited real run —
--limit 20. Inspect by hand: albums collapsed into one folder, standalone messages in their own, filenames{message_id}_{name},title.txtwritten, directory namedg<chat_id>. -
Sidecar cross-check — scripted, not by eye:
python3 - <<'EOF' # 1. every files[].path in messages.jsonl exists on disk # 2. every file on disk appears in the UNION of records for its message_id # 3. no two post folders share a grouped_id <- catches album splitting EOF -
Idempotency —
--reset-state, then re-run--limit 20. Expect zero downloads, all skips. (Re-running--limit 20without resetting would correctly continue to posts 21–40 — the cursor advanced because those 20 posts genuinely completed. Resetting is what isolates the dedupe path, which is the thing under test.) -
Interruption — start a full run,
kill -9after ~50 files. Re-run. Confirm: no surviving.part, no re-downloads of completed files, no gaps, the in-flight album complete. -
Network loss — drop the interface or block Telegram with iptables mid-download for ~30s, then restore. Confirm the transient path retries with backoff and the run continues.
kill -9alone does not exercise exception propagation or reconnect, so it cannot stand in for this. -
Filter guard — re-run with
--types videoagainst the existing state. Confirm it refuses, names the changed filters, and exits 7. Then--reset-stateand confirm it proceeds and rotatesmessages.jsonl. -
Concurrency guard — start a run; while it holds the lock, start a second against the same
--out. Confirm exit 3 with pid/host/start-time, and that no file was touched. Confirm a run against a different group proceeds. -
Path safety — covered by
tests/test_paths.pyandtests/test_download_decisions.py, which include the hostile-filename case offline. (The previous "otherwise rely on unit tests and note no live sample existed" escape clause is deleted — it made passing the default.) Optionally self-send a message named../../../etc/passwdto a scratch group for live confirmation. -
Full export — run to completion. Confirm
completed_atis set. Reconcile against step 1: post and file counts exactly equal; bytes within 2% (photo approximation only). Any file-count delta is a bug to investigate, never accepted — expected differences are only deleted/expired media, which appear inerrors[]and the exit summary. -
Flood-wait behavior — grep logs: each wait slept through once, no retry storm, no restart during an active wait, and every entry carries a computed wake time. Separately confirm
--max-flood-wait 1exits 6 cleanly with a resumable cursor, since that path is unreachable by default.
11b. --include-text — run with the flag on a scratch range. Confirm text-only messages appear with "files":[], create no folders, and that toggling the flag against the existing cursor triggers the exit-7 refusal.
- Memory —
/usr/bin/time -v tg-export --dry-run→ peak RSS under 200 MB. The album tripwire set is O(albums), not O(1); the criterion states the real bound.
12b. Sweep cost under a date filter — measure before deciding whether to optimize it. With reverse=True and no offset_date, telethon 1.44's _MessagesIter._init starts the sweep at message id 1, so --since walks the entire history discarding messages in keep(). Time --dry-run --since <recent date> against a large group and record the wall clock, the getHistory round-trip count, and the ratio of messages swept to messages kept. Do the same for --until, which currently never terminates the sweep early.
Both fixes carry an invariant risk, which is why this is a measurement rather than a change:
- server-side
offset_datecan start the sweep mid-album, so_closewould take apost_idthat is not the album's lowest id — breaking filter-invariant post identity. Needs a mid-album guard before it is safe. - an early break on
--untilassumes message dates rise monotonically with ids, which is false for imported history.
Decide with the numbers: if the ratio is small the current sweep is fine, and neither risk is worth taking.
- Revoke the spike session — Telegram → Settings → Devices → terminate the phase 1 probe session. Deleting
spike/probe.sessiondoes not revoke server-side authorization; a live auth key otherwise survives for an account you believe is clean.
Related Code Files
- Modify:
README.md - Delete:
spike/(or leave gitignored)
README Must Document
- Obtaining
api_id/api_hash; the credential policy table (env vs prompt vs getpass; what.envmay contain) - First-run login incl. 2FA; session path and why it defaults outside the repo
- Why the Bot API cannot do this — the first question any reader will have
- Resume semantics: the cursor means "handled through here under these filters", and why changing filters requires
--reset-state - Sidecar semantics: records are append-only events; union over records per
message_id, not last-write-wins - Export directory is
g<chat_id>with the name intitle.txt, and why (path safety + surviving group renames) - Why downloads are sequential — flood-wait escalation, and that
--workersdeliberately does not exist --include-text: off by default; when set, text-only messages appear in the sidecar with"files":[]and no folder. It is part of the stored filter set, so toggling it requires--reset-state- Flood waits have no ceiling by default — the tool sleeps as long as Telegram demands, logging the computed wake time each time. A long sleep is a sleep, not a hang; check the log before assuming otherwise. Pass
--max-flood-wait SECONDSto exit resumably instead
- That the dry-run verdict is advisory; the per-file disk check enforces. A
SHORTverdict means the run stops partway, resumably — not that it is unsafe cryptgis optional and not installed by default: native code in Telethon's AES path, trusted (if adopted) because the Telethon author maintains it — not because it is "CPU-side only"
Success Criteria
- Every step 1–13, including 11b, passes with its stated outcome — no step passes by escape clause
- Full export completes with
completed_atset; counts reconcile within the stated tolerance - Kill-and-resume produces a complete export with no gaps and no duplicates
- Network-loss interruption recovers without operator intervention
- Filter guard fires with a named diff and clears with
--reset-state - Concurrency guard exits 3; a different group is unaffected
pytestgreen;pip install -r requirements.txt -e .from a clean venv yields a workingtg-exportgit statusclean of session files, exports, state, lock, and.env- Spike session terminated server-side
- README covers every item listed above
Risk Assessment
- Group larger than free disk → step 1 reports it; the per-file guard stops cleanly and resumably. Scope down with
--max-size/--typesor export to external storage. - Long export interrupted by network loss → step 6 is the proof; step 5 alone would not be.
- Estimate/actual discrepancy → expected sources are deleted/expired media (visible in
errors[]) and photo approximation. Anything else is a filter or traversal bug — investigate. - Album split → step 3's grouped_id check plus phase 3's runtime tripwire. A split that reaches disk is silent otherwise.
- Account flood-limited during validation → keep
--limitsmall in steps 2–8; run step 10 once, unattended. - Third-party data committed → the sidecar contains message text and sender names. Export root is gitignored and
assert_not_committablerefuses an un-ignored path inside the repo; verify before any commit. - Spike credential left live → step 13. The most-missed cleanup item, because the local file looks like the whole credential.