Commit Graph
11 Commits
Author SHA1 Message Date
tiennm99 ec0584e34e chore: churn project - strip implementation, keep tdl + webdav instructions
The Python/Telethon exporter is retired in favor of iyear/tdl, which
covers the same purpose and is actively maintained. README now carries
only migration instructions, including the rclone rolling-pipeline
recipe for exporting to WebDAV. Full implementation remains at d170986.
2026-08-25 15:03:32 +07:00
tiennm99 d1709866e3 docs: correct the dry-run lock and exit-code claims, record the sweep measurement
Two README statements were false rather than merely incomplete. "--dry-run
takes no lock at all, so it can always be run against an export that is
currently in progress" contradicted the paragraph four lines above it
warning that two clients sharing one session is a corruption hazard:
dry-run opens the session file like any other run. Both places now say
what is true - no *export* lock, but a session lock, so an alongside
dry-run needs its own --session.

The exit table promised 1 for a bad argument. argparse exits 2 on its
own, before main()'s try block can map anything, so a bad --limit
collided with the dry-run SHORT verdict. Documented as the shared code
it is instead of a code the tool never returns.

Also notes that both locks are flock-based, and that flock is advisory or
per-client on NFS, so single-instance enforcement is not guaranteed
there. Setup gains the venv creation step it assumed.

Phase 7 gains step 12b: measure the sweep cost under a date filter before
deciding whether to optimize it. With reverse=True and no offset_date,
telethon starts the sweep at message id 1 - verified in the pinned wheel -
so --since walks the whole history discarding messages in keep(), and
--until never terminates early. Neither fix is free: offset_date can
start the sweep mid-album, which would take a post_id that is not the
album's lowest id and break filter-invariant identity, and an --until
break assumes dates rise with ids, false for imported history. The
swept-to-kept ratio decides whether either risk is worth taking.

Reports from the review pass are recorded under plans/reports/, including
the findings left unfixed: resolve_entity catching only ValueError,
_find_in_dialogs being the one network loop outside a flood primitive,
SystemExit in session.py routing around the Abort contract, and
completed_at being cleared before any work begins.
2026-08-22 23:55:00 +07:00
tiennm99 c37598e008 fix: lock the session before opening it, and stop a failed export reading as finished
The session lock was taken after the thing it guards. It lived inside
connected_client, so connect() and _login() had already written the
shared SQLite session by the time the lock existed. Two runs on the
default session each passed their own per-root lock, both opened the same
database, and the loser exited 3 *after* causing the corruption its
message described. The lock now wraps the whole client lifetime, and the
committable-path check runs first so a refused run leaves no lock file
next to a session it was never allowed to create.

--dry-run is inside that lock too. It writes nothing to the export tree,
but it opens the same session file, which is the resource the lock is
about - so running it alongside an export now needs its own --session.

Refusing to mark an export complete required that *nothing* had
succeeded: `failed and not (downloaded or skipped)`. A single
already-present file made skipped non-zero and disabled the guard
outright, so any resume across a partly-complete export could fail every
remaining file and still stamp completed_at with a cursor at
end-of-history. A broken export then answered "did my export finish?"
with a confident yes. The comparison is now against downloaded + skipped;
a legitimate tail of present files still outnumbers its own stray
failures and completes normally.

Also:

- Renewing an expired file reference counted as a failed attempt, so
  expiry on the final attempt burned the last slot and the fresh
  reference was never fetched - reported as "exhausted 3 attempts" after
  two. The refreshed latch already bounds that arm.
- Sidecar repair scanned a fixed window back from EOF for the last
  newline. A partial record larger than the window contains none, so the
  file was truncated to end-window: still unreadable, one megabyte
  shorter. The window now grows until a newline is found.
- --reset-state made the zeroed cursor durable before rotating the old
  sidecar, leaving exactly the mixed-generation log that rotating exists
  to prevent. Rotation now precedes State.open, which is sound only on
  this path because there is no compatibility check to fail.
- Two resets inside one second silently clobbered the first archive
  through os.replace, and the rename was the one here not fsynced.
- title.txt was the only untrusted string written raw. The export root is
  safe because the title never becomes a path component, but cat title.txt
  handed ANSI escapes and a right-to-left override to the operator. It is
  stripped of the same Unicode categories filenames are, from one shared
  set so the two rules cannot drift, and written atomically.

Cost, on the most common path of every resume:

- The post directory was fsynced every post, including posts where every
  file was already present and nothing had been renamed - one fsync per
  post to re-record a directory entry an earlier run had already made
  durable. Gated on an actual download.
- An already-present file was stat'd three times: exists(), stat(), and
  again inside _result. One stat now, reused as the recorded size.

Moving fsync off the loop is not available: the AST scan forbids
to_thread and run_in_executor, deliberately, because parallelism here
buys nothing and escalates flood waits.

Each fix has a test that fails without it, confirmed by reverting the
fix and re-running. 204 tests to 214, coverage unchanged at 93%.
2026-08-22 23:20:00 +07:00
tiennm99 3a51f33068 fix: repair the real export path and close taxonomy and sanitizer gaps
The real run was dead. Removing Config.max_flood_wait left cli._real_run
still passing it, so every non-dry-run export raised an uncaught
TypeError after taking both locks and writing title.txt and the state
file, but before downloading anything. Dry-run was unaffected, which is
why it looked healthy.

That escaped 180 passing tests because every one of them constructed
Config by hand and none invoked main(). test_cli_dispatch.py now drives
the whole dispatch against a fake client - real run, resume, dry run,
--limit 0, exit 7, exit 3, and an unexpected exception still honoring
the exit contract. That test file, not the one-line fix, is the remedy.

Also:

- AuthKeyNotFound subclasses plain Exception and matched no clause, yet
  it is what actually arrives when the connection drops mid-download:
  MTProtoSender sets it on every in-flight request. It now aborts with
  exit 4 like its rarer RPC-level sibling.
- CdnFileTamperedError was swallowed as a network error and retried
  three times. A hash mismatch on CDN-served bytes is a trust event, not
  transient noise, so it gets its own clause and its own message.
- Unicode Cf/Zl/Zp characters survived filename sanitization. U+202E
  RIGHT-TO-LEFT OVERRIDE renders 'a<RLO>gpj.exe' as 'a.exe.jpg' in any
  terminal or file manager - the ANSI vector was closed while the older
  extension-spoofing one stayed open. Stripping by Unicode category
  subsumes the previous control-character table.
- A run where every file failed and none succeeded no longer sets
  completed_at. A dead session looks exactly like that from inside the
  loop, and the field is supposed to answer 'did my export finish?'
  without guessing.
- Network errors reaching main() are no longer labelled 'filesystem',
  and a catch-all keeps any unexpected exception inside the exit-code
  contract.
- The retry ladder no longer sleeps after its final attempt.

Verified unchanged: filenames for ordinary extensions, so no previously
downloaded file is orphaned on resume.
2026-08-12 23:55:00 +07:00
tiennm99 a12ddcb9c0 docs: sync the plan with implementation and review outcomes
Phases 2-6 complete, phase 1 partial, phase 7 pending.

Records what was settled without live credentials: offset_id is an
exclusive lower bound under reverse=True, read from Telethon's iterator
rather than inferred from one sample, so the cursor stores the last
handled id. Also logs the defects found along the way and the two
critical findings from the post-implementation review, plus the one
finding left unresolved rather than decided unilaterally - whether a
file that exhausted its retries should stay abandoned behind a set
completed_at.
2026-08-12 23:40:00 +07:00
tiennm99 30f68b43ff docs: document credentials, resume semantics and flood-wait behavior
Covers why the Bot API structurally cannot do this, what .env may and
may not hold, why the export directory is a chat id, and the sidecar's
union-over-records rule - reading it as last-write-wins reports a
narrower file list than what is on disk after a filter change.

Also states plainly that flood waits have no ceiling by default: the
tool sleeps as long as Telegram demands and logs the computed wake time,
so a long sleep reads as a sleep rather than a hang.
2026-08-12 22:15:00 +07:00
tiennm99 5bf0992fe6 test: cover path safety, resume, flood waits and the concurrency rule
180 tests, all offline against synthetic message stubs: no credentials,
no network, no live group.

The cases that matter are the ones where 'looks like success' and 'is
success' diverge - a cursor ahead of what is on disk, a split album, a
zero-byte file renamed to its target, an empty filter set serialized as
no filter at all.

Two guards are tested for their ability to fail rather than assumed:
the AST scan that forbids parallelism primitives is checked against a
sample that must trip it, and lock contention is proven with a real
second process, since two acquisitions in one process would succeed and
prove nothing.
2026-08-12 21:30:00 +07:00
tiennm99 275ec47576 feat: add tg-export, a resumable Telegram group media exporter
Exports all media from a group to disk, one folder per logical post with
albums collapsed, alongside a messages.jsonl sidecar linking every file
to its message.

Three properties the design is built around:

- Identity is filter-invariant. A post's folder is the lowest message id
  in the full grouped_id group and filenames key on message id, so
  --types photo and --types video address the same files. A positional
  index would shift whenever a member message was deleted.
- The cursor commits only after a complete post and is unreachable from
  every error path, so it can never point past in-flight work.
- Filters are stored verbatim; a mismatch refuses the run and names what
  changed, because the cursor means 'handled through here under these
  filters'.

Exactly one untrusted string becomes a path component: the filename
leaf, sanitized stem and extension both. Every other component is an int
enforced by type. The export directory is the chat id rather than the
group title, which makes traversal via a renamed group structurally
impossible and stops a mid-export rename from orphaning the download.

Downloads are sequential. Telethon does not parallelize transfers, and
concurrency mainly accelerates flood-wait escalation from seconds to
hours.
2026-08-12 20:10:00 +07:00
tiennm99 65f8abdf18 build: add packaging and pin the verified Telethon version
Telethon 2.0 is alpha with a renamed API surface, hence the <2 bound.
requirements.txt records 1.44.0, the version the API probes were run
against, so a later resolver bump cannot silently change the semantics
the traversal depends on.
2026-08-12 15:45:00 +07:00
tiennm99 de9a08796e chore: add gitignore covering sessions, exports, and credentials 2026-08-12 15:01:53 +07:00
tiennm99 4ce369d953 Initial commit 2026-08-11 14:07:26 +07:00