The verifier reconstructs {DialogID}_{MessageID}_{FileName} from the export
and looks it up in the remote listing. A file stored under any other name is
not the file the export asked for, so it counts as absent however close the
name looks.
That leaves a stale copy behind when the re-download lands beside it, so
absent files whose message id is already present on the remote are now listed
separately with both names, ready to be removed.
Remote paths are reduced to their basename first, so a grouped subdirectory
layout matches the same way a flat one does.
Backends that finish an upload as a server-side task, pikpak among them,
leave it pending long enough that rclone abandons the transfer once
--low-level-retries polls run out. Default RCLONE_TRANSFERS to 2 to keep
that task queue short and RCLONE_LOW_LEVEL_RETRIES to 20 to wait a slow
commit out. Both use the :=default form, so an exported value still wins.
Shorten the --min-age default to 45s: excluding '*.tmp' is what actually
keeps an unfinished download from being uploaded, so the age guard only
needs to cover a rename racing a sweep, and two minutes of dwell made
staging larger than it had to be.
Sweep every 120s rather than 300s in the driver, which shrinks the backlog
each sweep has to move, and raise the pass limit to 30 to suit chats that
need more rounds than 20 to converge.
Add -m SIZE to run.sh, passed through by export-until-complete.sh, so the
staging directory cannot outgrow the local disk when tdl downloads faster
than rclone uploads. Staging is measured every 10 seconds, independently of
the sweep interval; at the cap tdl is suspended with SIGSTOP and rclone
sweeps until staging is back under it, then tdl resumes and --continue picks
its .tmp files back up. The drain loops while it frees space and gives up
once only in-flight .tmp files remain, which no sweep can move, so a cap set
below what the concurrent downloads hold warns instead of stalling. The
failure streak still counts only sweeps that actually ran, so a quiet cap
check cannot clear it.
The sweeps that run once at the end now report progress rather than going
quiet for minutes: rclone's redrawn bar on a terminal, one-line stats every
30s when output is redirected. Stats are logged at INFO, so only the stats
level is raised to NOTICE, avoiding the line-per-file output -v would add.
Also send SIGCONT before SIGTERM when stopping tdl, since a suspended
process never sees the term.
run.sh finishes when tdl finishes, which does not mean every file arrived: a
dropped session, a stalled remote or an interrupted pass leaves gaps that
nothing reported.
verify-export.sh rebuilds the filename tdl produces for each media message in
the export JSON and checks the remote and staging for it. Messages carrying no
media are skipped rather than counted as gaps. A zero-byte file counts as
missing because rclone overwrites a size-mismatched destination, so a retry
repairs it; files under 1 KiB are reported but not retried, since real media is
sometimes that small and retrying them would never terminate.
export-until-complete.sh loops that check with run.sh, narrowing the export to
the outstanding ids each pass. It stops when the verifier reports complete, when
a pass fetches nothing new, when the remote runs low on space, or on Ctrl-C.
tdl's progress bar is shown on a terminal and suppressed when output is
redirected, so logs stay readable.
Exports, staging and logs are now ignored: they hold real message ids, file
names and text from an account.
-c already passed through to tdl, which resolves a numeric argument as an
MTProto id and anything else through gotd's resolver (@name, name, t.me and
tg:// links). Two forms needed handling: a Bot API '-100' prefixed id, which
MTProto does not use and which tdl never strips, is now converted; and a
message link, which names a message rather than a chat, is rejected with a
clear error instead of a resolver failure.
The export JSON now defaults to export-<chat>.json when -c is given, so
exporting a second chat from the same directory cannot silently download the
first chat again from a stale export.json. An existing file is still reused,
but the script now says which file it is using.
The pipeline never depended on WebDAV; only the naming, docs and preflight
did. Replace the WebDAV-root listing with a portable check: a named remote
must appear in 'rclone listremotes', and creating the destination proves
reachability and credentials. On-the-fly connection strings have no config
entry, so the name check is skipped for them.
Read the remote list into a variable rather than piping it into 'grep -q':
grep exits on the first match and kills rclone with SIGPIPE, which pipefail
reports as failure, intermittently rejecting a configured remote.
Document that rclone's flags are all settable through their RCLONE_*
environment variables, so the upload side is tunable without new options.
The README now covers only what the tool does and how to run it: setup,
options, tdl flag pass-through, subset exports, resume semantics, the
failure modes it guards against, and exit codes.
Flags are documented from tdl's own source rather than assumed: dl -f takes
the exported JSON while dl -i/-e are file-extension filters, and the
expression filter is chat export -f. Chats are addressed by id or domain.
Also warn that the inline rclone loops need --exclude '*.tmp' themselves.
Without it a download stalled by a flood wait is uploaded half-written and
its local copy deleted, which breaks tdl --continue for that file.
tdl has no remote destination, so exporting a group larger than local disk
needs a staging dir plus a concurrent uploader. The script runs tdl into a
small staging dir while rclone moves finished files out to WebDAV, so local
disk only holds the files in flight plus one sync interval of throughput.
Partial downloads are excluded by name rather than by age: tdl writes
<name>.tmp and renames on completion, so a download stalled by a flood wait
stops touching its .tmp and would otherwise age past --min-age and be
uploaded half-written, destroying resume for that file.
--delete-empty-src-dirs runs only in the final sweep, because removing a
directory under a running tdl makes it fail to create its next file. tdl is
stopped on exit, interrupt or termination so no download is orphaned, the
unrestricted sweep runs only after tdl exits 0, and five consecutive rclone
failures abort the run before staging fills the disk.
The Python/Telethon exporter is retired in favor of iyear/tdl, which
covers the same purpose and is actively maintained. README now carries
only migration instructions, including the rclone rolling-pipeline
recipe for exporting to WebDAV. Full implementation remains at d170986.
Two README statements were false rather than merely incomplete. "--dry-run
takes no lock at all, so it can always be run against an export that is
currently in progress" contradicted the paragraph four lines above it
warning that two clients sharing one session is a corruption hazard:
dry-run opens the session file like any other run. Both places now say
what is true - no *export* lock, but a session lock, so an alongside
dry-run needs its own --session.
The exit table promised 1 for a bad argument. argparse exits 2 on its
own, before main()'s try block can map anything, so a bad --limit
collided with the dry-run SHORT verdict. Documented as the shared code
it is instead of a code the tool never returns.
Also notes that both locks are flock-based, and that flock is advisory or
per-client on NFS, so single-instance enforcement is not guaranteed
there. Setup gains the venv creation step it assumed.
Phase 7 gains step 12b: measure the sweep cost under a date filter before
deciding whether to optimize it. With reverse=True and no offset_date,
telethon starts the sweep at message id 1 - verified in the pinned wheel -
so --since walks the whole history discarding messages in keep(), and
--until never terminates early. Neither fix is free: offset_date can
start the sweep mid-album, which would take a post_id that is not the
album's lowest id and break filter-invariant identity, and an --until
break assumes dates rise with ids, false for imported history. The
swept-to-kept ratio decides whether either risk is worth taking.
Reports from the review pass are recorded under plans/reports/, including
the findings left unfixed: resolve_entity catching only ValueError,
_find_in_dialogs being the one network loop outside a flood primitive,
SystemExit in session.py routing around the Abort contract, and
completed_at being cleared before any work begins.
The session lock was taken after the thing it guards. It lived inside
connected_client, so connect() and _login() had already written the
shared SQLite session by the time the lock existed. Two runs on the
default session each passed their own per-root lock, both opened the same
database, and the loser exited 3 *after* causing the corruption its
message described. The lock now wraps the whole client lifetime, and the
committable-path check runs first so a refused run leaves no lock file
next to a session it was never allowed to create.
--dry-run is inside that lock too. It writes nothing to the export tree,
but it opens the same session file, which is the resource the lock is
about - so running it alongside an export now needs its own --session.
Refusing to mark an export complete required that *nothing* had
succeeded: `failed and not (downloaded or skipped)`. A single
already-present file made skipped non-zero and disabled the guard
outright, so any resume across a partly-complete export could fail every
remaining file and still stamp completed_at with a cursor at
end-of-history. A broken export then answered "did my export finish?"
with a confident yes. The comparison is now against downloaded + skipped;
a legitimate tail of present files still outnumbers its own stray
failures and completes normally.
Also:
- Renewing an expired file reference counted as a failed attempt, so
expiry on the final attempt burned the last slot and the fresh
reference was never fetched - reported as "exhausted 3 attempts" after
two. The refreshed latch already bounds that arm.
- Sidecar repair scanned a fixed window back from EOF for the last
newline. A partial record larger than the window contains none, so the
file was truncated to end-window: still unreadable, one megabyte
shorter. The window now grows until a newline is found.
- --reset-state made the zeroed cursor durable before rotating the old
sidecar, leaving exactly the mixed-generation log that rotating exists
to prevent. Rotation now precedes State.open, which is sound only on
this path because there is no compatibility check to fail.
- Two resets inside one second silently clobbered the first archive
through os.replace, and the rename was the one here not fsynced.
- title.txt was the only untrusted string written raw. The export root is
safe because the title never becomes a path component, but cat title.txt
handed ANSI escapes and a right-to-left override to the operator. It is
stripped of the same Unicode categories filenames are, from one shared
set so the two rules cannot drift, and written atomically.
Cost, on the most common path of every resume:
- The post directory was fsynced every post, including posts where every
file was already present and nothing had been renamed - one fsync per
post to re-record a directory entry an earlier run had already made
durable. Gated on an actual download.
- An already-present file was stat'd three times: exists(), stat(), and
again inside _result. One stat now, reused as the recorded size.
Moving fsync off the loop is not available: the AST scan forbids
to_thread and run_in_executor, deliberately, because parallelism here
buys nothing and escalates flood waits.
Each fix has a test that fails without it, confirmed by reverting the
fix and re-running. 204 tests to 214, coverage unchanged at 93%.
The real run was dead. Removing Config.max_flood_wait left cli._real_run
still passing it, so every non-dry-run export raised an uncaught
TypeError after taking both locks and writing title.txt and the state
file, but before downloading anything. Dry-run was unaffected, which is
why it looked healthy.
That escaped 180 passing tests because every one of them constructed
Config by hand and none invoked main(). test_cli_dispatch.py now drives
the whole dispatch against a fake client - real run, resume, dry run,
--limit 0, exit 7, exit 3, and an unexpected exception still honoring
the exit contract. That test file, not the one-line fix, is the remedy.
Also:
- AuthKeyNotFound subclasses plain Exception and matched no clause, yet
it is what actually arrives when the connection drops mid-download:
MTProtoSender sets it on every in-flight request. It now aborts with
exit 4 like its rarer RPC-level sibling.
- CdnFileTamperedError was swallowed as a network error and retried
three times. A hash mismatch on CDN-served bytes is a trust event, not
transient noise, so it gets its own clause and its own message.
- Unicode Cf/Zl/Zp characters survived filename sanitization. U+202E
RIGHT-TO-LEFT OVERRIDE renders 'a<RLO>gpj.exe' as 'a.exe.jpg' in any
terminal or file manager - the ANSI vector was closed while the older
extension-spoofing one stayed open. Stripping by Unicode category
subsumes the previous control-character table.
- A run where every file failed and none succeeded no longer sets
completed_at. A dead session looks exactly like that from inside the
loop, and the field is supposed to answer 'did my export finish?'
without guessing.
- Network errors reaching main() are no longer labelled 'filesystem',
and a catch-all keeps any unexpected exception inside the exit-code
contract.
- The retry ladder no longer sleeps after its final attempt.
Verified unchanged: filenames for ordinary extensions, so no previously
downloaded file is orphaned on resume.
Phases 2-6 complete, phase 1 partial, phase 7 pending.
Records what was settled without live credentials: offset_id is an
exclusive lower bound under reverse=True, read from Telethon's iterator
rather than inferred from one sample, so the cursor stores the last
handled id. Also logs the defects found along the way and the two
critical findings from the post-implementation review, plus the one
finding left unresolved rather than decided unilaterally - whether a
file that exhausted its retries should stay abandoned behind a set
completed_at.
Covers why the Bot API structurally cannot do this, what .env may and
may not hold, why the export directory is a chat id, and the sidecar's
union-over-records rule - reading it as last-write-wins reports a
narrower file list than what is on disk after a filter change.
Also states plainly that flood waits have no ceiling by default: the
tool sleeps as long as Telegram demands and logs the computed wake time,
so a long sleep reads as a sleep rather than a hang.
180 tests, all offline against synthetic message stubs: no credentials,
no network, no live group.
The cases that matter are the ones where 'looks like success' and 'is
success' diverge - a cursor ahead of what is on disk, a split album, a
zero-byte file renamed to its target, an empty filter set serialized as
no filter at all.
Two guards are tested for their ability to fail rather than assumed:
the AST scan that forbids parallelism primitives is checked against a
sample that must trip it, and lock contention is proven with a real
second process, since two acquisitions in one process would succeed and
prove nothing.
Exports all media from a group to disk, one folder per logical post with
albums collapsed, alongside a messages.jsonl sidecar linking every file
to its message.
Three properties the design is built around:
- Identity is filter-invariant. A post's folder is the lowest message id
in the full grouped_id group and filenames key on message id, so
--types photo and --types video address the same files. A positional
index would shift whenever a member message was deleted.
- The cursor commits only after a complete post and is unreachable from
every error path, so it can never point past in-flight work.
- Filters are stored verbatim; a mismatch refuses the run and names what
changed, because the cursor means 'handled through here under these
filters'.
Exactly one untrusted string becomes a path component: the filename
leaf, sanitized stem and extension both. Every other component is an int
enforced by type. The export directory is the chat id rather than the
group title, which makes traversal via a renamed group structurally
impossible and stops a mid-export rename from orphaning the download.
Downloads are sequential. Telethon does not parallelize transfers, and
concurrency mainly accelerates flood-wait escalation from seconds to
hours.
Telethon 2.0 is alpha with a renamed API surface, hence the <2 bound.
requirements.txt records 1.44.0, the version the API probes were run
against, so a later resolver bump cannot silently change the semantics
the traversal depends on.