pikpak commits an upload as a server-side async task, and rclone polls that
task only as long as its low-level retries last. A task still in
PHASE_TYPE_PENDING when that budget runs out is reported as "can't verify the
task is completed" — and checking the live archive ten minutes later, those
tasks had not committed: the objects were simply absent.
The transfer itself was fine; only the confirmation timed out. Abandoning the
file meant re-downloading it from Telegram on the next pass, which for this
archive can be two gigabytes, so a run against a slow remote paid for the same
bytes repeatedly. The staged copy is still on disk when a move fails — rclone
says so in the same breath — so another attempt costs seconds instead.
Every attempt after the first checks the remote before re-uploading. A pending
task may have committed during the backoff, and pikpak allows two files under
one name, so re-uploading blind is how one file becomes two — which
verification then reports as an ambiguous basename on every future run.
rclone's log now goes through the renderer. It writes to stderr on its own
schedule, so its error lines were landing mid-redraw and shredding the display
exactly when there was most to read.
Also: "capped at uncapped" now reads "uncapped", and the ETA column fits the
three-digit hour counts a slow remote produces.
A run reported one aggregate line, which could say how much was done but never
what was happening: which files were moving, whether a stall was a slow
download or a slow upload, or how long the rest would take. On a transfer
measured in hours those are the only questions worth answering.
The pipeline now emits a per-item lifecycle rather than only Stats, and a
terminal renders it as an overall bar plus one bar per file in flight. Uploads
get a spinner rather than a bar because rclone's MoveFile is a single blocking
call with no byte callbacks; naming the file is still the point, since a run
that looks stalled is usually waiting on one large object.
Before the transfer, the run states what it found and what it will do. The
survey breaks the outstanding set down by reason — never fetched, zero-byte,
wrong size, unarchivable — because that is the difference between a run that
will converge and one that cannot, and the old "N to fetch" hid it. The plan
states the cap, staging path and concurrency, so a wrong setting is visible
before hours of transfer rather than after.
Redirected output keeps plain lines and gains one per archived file. Bars are
continuous cursor movement, and a captured log of them is what tdl's progress
bar did to the shell pipeline's logs.
cmd/uidemo renders the whole thing against fake data. It is how the layout was
checked without a session, and it earned its place immediately: the overall bar
sat at zero because Stats only tracked the file count and never advanced it.
Both phases ran for minutes printing nothing between their opening line and
their result, so a working run looked exactly like a hung one — which is how it
was reported.
The message counter lives in Walk rather than in the caller's loop because most
of a chat is not media: text-only and service messages are filtered out inside
the walk, so a caller counting yielded items still sees nothing while crossing a
long stretch of conversation.
Cadence follows the existing reporter: a terminal redraws one line, a redirected
run gets a periodic one, since ANSI redraws are what turned the shell pipeline's
captured logs into megabytes of control characters.
run.sh, export-until-complete.sh and verify-export.sh are replaced by the
tgexport binary. Everything expensive in them existed because tdl and rclone
could not see each other's state — a staging directory polled with du, an age
guard to guess when a download had finished, SIGSTOP/SIGCONT to enforce the
disk cap, and an outer loop re-narrowing a JSON export between passes. One
process needs none of it.
The README keeps the comparison, including the pikpak-specific tunables and
the size judgements worth carrying over.
Verification matched on name and non-zero size, so an upload that died partway
was counted archived permanently. Comparing against the size Telegram reports
found six such objects in the live archive, one of them 221 MiB standing in for
a 2006 MiB video. They are folded into the outstanding set; rclone overwrites a
size mismatch, so another pass repairs them.
A zero Report claimed the archive was complete — nothing expected, nothing
missing — which is the value both commands hold before their Telegram callback
populates it, so any early return printed COMPLETE and exited 0 on an untouched
chat. A Report now knows whether it ran.
A name that can never be written kept the run outstanding forever while the
download set deliberately excluded it, so a driver looping on "incomplete"
walked the whole history and re-indexed the whole remote every pass for work
that could not be done. Such a run now reports STALLED and exits 4.
A destination that stopped accepting uploads was reported and then discarded,
exiting 1. The same driver would retry against a full or unreachable remote
indefinitely, downloading gigabytes each pass to upload none. It exits 3.
Free space is re-checked during the run, not only before it. An archive this
size runs for hours, and the remote can fill in the middle; discovering it
through five failed multi-gigabyte uploads wastes the download for all of them.
Also: parseSize silently wrapped to a negative or zero on a large input, which
reads downstream as "no cap"; list printed attacker-chosen filenames raw, so a
tab shifted the columns and an escape sequence reached the terminal; a missing
backend blamed credentials rather than the build; humanBytes indexed past its
unit table above 1 PiB; a second Init reported success against a config that
never loaded; and sync did not surface the basename collisions verify warned
about, though sync is the command that acts on the verdict.
The env-override test could not observe what it claimed: rclone reads RCLONE_*
at package init, so t.Setenv came too late and the assertion held with the
guard removed. It runs in a subprocess now, as does the new one covering the
index against inherited filters.
rclone does not address a file by the bytes os.OpenFile wrote. Names handed to
an Fs go through the backend encoder and names listed back are re-encoded to
the standard set, neither of which the write path performs. A filename holding
one of the rewritten characters was therefore stored under one string and
looked up under another: confirmed against the local backend, where a written
"a‛b.jpg" reports object not found and a written "a\nb.jpg" lists back as
"a␊b.jpg". That is the same two-derivation divergence this program was written
to remove, with rclone's encoder standing where filenamify used to. Such names
are rejected, not encoded, for the same reason every other name is.
The length limit ignored the ".part" suffix that is opened first, so a name
just inside NAME_MAX passed the check and then failed to open on every pass,
stalling the walk on that message forever. The suffix now lives beside the
limit that has to account for it.
filter.NewFilter(nil) does not build a neutral filter; it copies the package
global, which rclone has already filled from RCLONE_*. Indexing inherited the
operator's environment, so a stray RCLONE_MIN_SIZE emptied the index and
re-downloaded the archive. Every narrowing field is now set explicitly and the
result is asserted inactive. A subprocess test covers it, since the env is read
at package init and t.Setenv is too late to observe anything.
An ErrorDirNotFound from a subdirectory was also treated as an empty
destination, returning a partial index as authoritative.
t.me/c/<id> and t.me/s/<name> were passed through whole, and gotd reads the
first path component as the username — resolving "c" or "s", which is a
confusing failure at best and someone else's chat at worst, since
one-character usernames exist. Both now yield the chat, and t.me/s/<name>/<id>
is refused like any other message link. Two tests asserted the old behaviour.
core's dcpool.Takeout deadlocks when takeout init fails: it holds the pool
mutex and recovers by calling Client, which takes the same non-reentrant
mutex. Telegram returns TAKEOUT_INIT_DELAY for a takeout started recently and
takeout is on by default, so two runs in succession hang the process with no
output and no response to cancellation. The session is established once here
instead, falling back to a plain client, and the pool's own Takeout is never
called.
rclone writes straight to the final remote name on any backend that does not
advertise PartialUploads, and cleans up after a failed Put only when it did
not. Pikpak advertises neither, so a transfer that died halfway left a
fragment under exactly the name verification matches on — counted archived by
that run and every run after it, with the local copy already deleted. A failed
move now looks for that object and removes it, leaving a complete one alone
since pikpak's async commit can still land it correctly.
An unwritable filename ended the whole walk, so one hostile name could strand
every message behind it. It is skipped and reported instead. The comment there
had described that behaviour all along.
Also: remove the part file when promoting it fails, since the caller hands the
reservation back and the cap would stay over-committed; report upload failures
alongside a download error rather than instead of it, which on Ctrl-C hid that
finished files had been discarded.
Run had no test of its own because it called Download directly. That step is
now indirected, covering the properties only the composition has: uploads
closed after the last send, the budget balanced across failures, and a tripped
breaker halting downloads rather than walking the whole chat.
core's Download returns without waiting on its worker group when the
iterator reports an error, so surfacing one through Iter.Err left workers
sending into a channel the caller had already closed. The iterator now
always reports a nil error and stashes the real one, read after Download
returns.
The circuit breaker cancelled only the upload context, which left
downloads running full speed against a remote refusing them: every file
stayed in staging and every reservation came back, so a broken remote
filled local disk faster than a working one. It now stops the download
iterator instead.
A failed move leaves the local copy in place, so the byte reservation
cannot be handed back until the file is removed. A confirmed-short object
is deleted rather than left under a name verification would count as
archived forever.
Also: release the reservation when opening the staging file fails, report
results even when an upload errored, bound the recorded errors, and drop
the reporter's lock before writing so terminal latency cannot throttle
downloads.
Phases 4 through 6: the two legs and the command that joins them.
Downloads go to <name>.part and are renamed only once complete, so a file
without the suffix is always whole. That is what lets the upload leg treat
"exists" as "finished" — the property run.sh could only approximate with a
filename convention plus an age guard, because it could not see inside tdl.
Every finished file is checked against the size Telegram reported, and that
check rather than the error is the authoritative signal. core's downloader logs
a failed transfer and returns nil, and its completion callback is deferred on
that named return, so a failure arrives indistinguishable from a success.
Trusting it would promote a truncated file and archive it as complete.
The disk cap is a semaphore over bytes. A download reserves its own size before
starting and releases it only after the upload confirms, so a slow remote
stalls downloads by itself. Blocking the iterator is safe because the
downloader calls it from its dispatch loop while workers run in a group, so a
blocked iterator never stops the uploads that free the space. Gone with it: the
du polling, the SIGSTOP and SIGCONT suspension, the min-age guard, the
temp-file filter and the sweep-failure counter.
A cap smaller than the largest file is refused up front. The semaphore could
never admit it, and a run blocked on a file it can never start looks exactly
like a stalled remote.
Uploads re-state each object to prove its size before the local copy is gone,
closing a gap where a truncated upload was only noticed by a later verify.
The destination is created before the chat is read. It is also the credentials
check, and doing it first means a bad destination fails in seconds rather than
after a full history walk.
One invocation converges: each item is checked against the index immediately
before download, so there are no passes and re-running is the resume path.
Options that no longer exist say what replaced them instead of failing as
unknown flags.
Verified end to end against the live chat and a scratch remote path: two files
downloaded, uploaded, confirmed present at the right size, staging left empty.
Third slice: verify-export.sh, without the subprocess or the python.
One rclone listing builds an in-memory index, and the report is computed from
it. missing-ids.txt and gap.json are gone; so is every python3 heredoc.
Presence is answered from a whole name and never from a message id. The
id-keyed map exists only to tell "absent" apart from "absent, but a stale copy
under an older name is sitting there", and it stays unexported so nothing can
reach for it as an answer. That distinction is the bug this rewrite exists to
remove, so it is enforced by structure rather than by comment.
Names are checked for path containment before use. Storing them verbatim means
a filename chosen by whoever uploaded the file can contain a separator or a
parent reference, and tdl never had to care because its template rewrote those
away. Over-long names are refused for the same reason: the filesystem would
reject them at create time, and a file that can never be written would be
reported absent on every pass forever.
Objects are addressed by the path rclone knows them by, not by the basename
used for matching. The two differ once a remote has directory structure, and
deleting by basename would miss the object or remove a same-named one from the
root. Basenames appearing at more than one path make the snapshot ambiguous, so
they are reported rather than silently resolved.
Indexing runs at full depth with filters cleared. Inheriting RCLONE_MAX_DEPTH
or RCLONE_EXCLUDE would not fail, it would quietly report archived files as
absent and fetch them all again.
Deleting stale copies stays opt-in and confirmed; a non-interactive stdin
declines rather than proceeding. Filenames are quoted wherever they are
printed, so an embedded escape cannot redraw the list an operator approves.
Verified against the live remote: identical to verify-export.sh on the same
state — 18155 expected, 15548 present, 2607 absent, ids 9857-18013, exit 1.
Second slice: the read path, and the fix for the bug that motivated the
rewrite.
The shell pipeline derived a filename twice. `tdl chat export` wrote the raw
Telegram name into a JSON, while `tdl dl` rendered it through a template whose
default pipes it through filenamify, which rewrites reserved characters and
collapses runs of '!'. A message whose name contained '!!' was therefore looked
up under one name and stored under another; the verifier never found it and
re-fetched it on every pass. Here a single function produces the name, and the
string it returns is used both to test for presence and to write the file, so
the two cannot disagree.
Names are stored exactly as Telegram reports them rather than reproducing
filenamify. That is a deliberate break from what the old pipeline wrote: a file
it stored under a rewritten name is not recognised and will be fetched again.
For the one chat archived so far that is a single file out of 18155, already
removed. Because names are verbatim they are not path-safe, so the code that
turns one into a path must enforce containment.
Walk yields a sequence rather than taking a callback, since the downloader
consumes a pull iterator and range-over-func converts either way without anyone
owning a goroutine. It pages newest first, where the old pipeline went oldest
first, which changes what an interrupted run leaves behind.
Message links are refused rather than guessed at, across every host Telegram
uses and the tg:// forms that carry the message id in a query parameter. A
private channel link and a public message link have the same shape, so the two
are told apart by parsing rather than by pattern.
Verified against the live chat: 18155 media messages, matching the shell
verifier, and every one of the 15548 objects already on the remote is found
under a derived name.
First slice of replacing the three-script pipeline with one process. The
scripts coordinate tdl and rclone as separate programs, so everything
expensive in them exists to work around the fact that neither can see the
other's state. A single process does not need that machinery.
This slice covers only the foundations: open the session tdl login already
wrote, resolve an rclone destination, and report on both via a doctor
command. Downloading, uploading and verification follow.
The session store is shared with the tdl CLI rather than copied, so the two
cannot run against one namespace at the same time; -n selects another.
AppID and AppHash are read from the store rather than hardcoded, because a
session is bound to the application that created it and tdl records which
one it used.
No middlewares are passed to tclient.New, which already prepends its own
defaults; the DC pool gets them instead, since gotd applies a client's
middlewares only to direct invocations and not to pooled connections.
The rclone config is loaded up front because the lazy path calls os.Exit on
a config it cannot read, which would bypass every defer and exit with the
code this tool reserves for an incomplete run.
Exit codes follow the shell pipeline: 0 ok, 1 incomplete, 2 usage, 3 remote
failure, 130 SIGINT, 143 SIGTERM. Cancellation is checked explicitly after
the Telegram client returns, because gotd reports an interrupted run as
success and a driver would read that as a finished archive.
The verifier reconstructs {DialogID}_{MessageID}_{FileName} from the export
and looks it up in the remote listing. A file stored under any other name is
not the file the export asked for, so it counts as absent however close the
name looks.
That leaves a stale copy behind when the re-download lands beside it, so
absent files whose message id is already present on the remote are now listed
separately with both names, ready to be removed.
Remote paths are reduced to their basename first, so a grouped subdirectory
layout matches the same way a flat one does.
Backends that finish an upload as a server-side task, pikpak among them,
leave it pending long enough that rclone abandons the transfer once
--low-level-retries polls run out. Default RCLONE_TRANSFERS to 2 to keep
that task queue short and RCLONE_LOW_LEVEL_RETRIES to 20 to wait a slow
commit out. Both use the :=default form, so an exported value still wins.
Shorten the --min-age default to 45s: excluding '*.tmp' is what actually
keeps an unfinished download from being uploaded, so the age guard only
needs to cover a rename racing a sweep, and two minutes of dwell made
staging larger than it had to be.
Sweep every 120s rather than 300s in the driver, which shrinks the backlog
each sweep has to move, and raise the pass limit to 30 to suit chats that
need more rounds than 20 to converge.
Add -m SIZE to run.sh, passed through by export-until-complete.sh, so the
staging directory cannot outgrow the local disk when tdl downloads faster
than rclone uploads. Staging is measured every 10 seconds, independently of
the sweep interval; at the cap tdl is suspended with SIGSTOP and rclone
sweeps until staging is back under it, then tdl resumes and --continue picks
its .tmp files back up. The drain loops while it frees space and gives up
once only in-flight .tmp files remain, which no sweep can move, so a cap set
below what the concurrent downloads hold warns instead of stalling. The
failure streak still counts only sweeps that actually ran, so a quiet cap
check cannot clear it.
The sweeps that run once at the end now report progress rather than going
quiet for minutes: rclone's redrawn bar on a terminal, one-line stats every
30s when output is redirected. Stats are logged at INFO, so only the stats
level is raised to NOTICE, avoiding the line-per-file output -v would add.
Also send SIGCONT before SIGTERM when stopping tdl, since a suspended
process never sees the term.
run.sh finishes when tdl finishes, which does not mean every file arrived: a
dropped session, a stalled remote or an interrupted pass leaves gaps that
nothing reported.
verify-export.sh rebuilds the filename tdl produces for each media message in
the export JSON and checks the remote and staging for it. Messages carrying no
media are skipped rather than counted as gaps. A zero-byte file counts as
missing because rclone overwrites a size-mismatched destination, so a retry
repairs it; files under 1 KiB are reported but not retried, since real media is
sometimes that small and retrying them would never terminate.
export-until-complete.sh loops that check with run.sh, narrowing the export to
the outstanding ids each pass. It stops when the verifier reports complete, when
a pass fetches nothing new, when the remote runs low on space, or on Ctrl-C.
tdl's progress bar is shown on a terminal and suppressed when output is
redirected, so logs stay readable.
Exports, staging and logs are now ignored: they hold real message ids, file
names and text from an account.
-c already passed through to tdl, which resolves a numeric argument as an
MTProto id and anything else through gotd's resolver (@name, name, t.me and
tg:// links). Two forms needed handling: a Bot API '-100' prefixed id, which
MTProto does not use and which tdl never strips, is now converted; and a
message link, which names a message rather than a chat, is rejected with a
clear error instead of a resolver failure.
The export JSON now defaults to export-<chat>.json when -c is given, so
exporting a second chat from the same directory cannot silently download the
first chat again from a stale export.json. An existing file is still reused,
but the script now says which file it is using.
The pipeline never depended on WebDAV; only the naming, docs and preflight
did. Replace the WebDAV-root listing with a portable check: a named remote
must appear in 'rclone listremotes', and creating the destination proves
reachability and credentials. On-the-fly connection strings have no config
entry, so the name check is skipped for them.
Read the remote list into a variable rather than piping it into 'grep -q':
grep exits on the first match and kills rclone with SIGPIPE, which pipefail
reports as failure, intermittently rejecting a configured remote.
Document that rclone's flags are all settable through their RCLONE_*
environment variables, so the upload side is tunable without new options.
The README now covers only what the tool does and how to run it: setup,
options, tdl flag pass-through, subset exports, resume semantics, the
failure modes it guards against, and exit codes.
Flags are documented from tdl's own source rather than assumed: dl -f takes
the exported JSON while dl -i/-e are file-extension filters, and the
expression filter is chat export -f. Chats are addressed by id or domain.
Also warn that the inline rclone loops need --exclude '*.tmp' themselves.
Without it a download stalled by a flood wait is uploaded half-written and
its local copy deleted, which breaks tdl --continue for that file.
tdl has no remote destination, so exporting a group larger than local disk
needs a staging dir plus a concurrent uploader. The script runs tdl into a
small staging dir while rclone moves finished files out to WebDAV, so local
disk only holds the files in flight plus one sync interval of throughput.
Partial downloads are excluded by name rather than by age: tdl writes
<name>.tmp and renames on completion, so a download stalled by a flood wait
stops touching its .tmp and would otherwise age past --min-age and be
uploaded half-written, destroying resume for that file.
--delete-empty-src-dirs runs only in the final sweep, because removing a
directory under a running tdl makes it fail to create its next file. tdl is
stopped on exit, interrupt or termination so no download is orphaned, the
unrestricted sweep runs only after tdl exits 0, and five consecutive rclone
failures abort the run before staging fills the disk.
The Python/Telethon exporter is retired in favor of iyear/tdl, which
covers the same purpose and is actively maintained. README now carries
only migration instructions, including the rclone rolling-pipeline
recipe for exporting to WebDAV. Full implementation remains at d170986.
Two README statements were false rather than merely incomplete. "--dry-run
takes no lock at all, so it can always be run against an export that is
currently in progress" contradicted the paragraph four lines above it
warning that two clients sharing one session is a corruption hazard:
dry-run opens the session file like any other run. Both places now say
what is true - no *export* lock, but a session lock, so an alongside
dry-run needs its own --session.
The exit table promised 1 for a bad argument. argparse exits 2 on its
own, before main()'s try block can map anything, so a bad --limit
collided with the dry-run SHORT verdict. Documented as the shared code
it is instead of a code the tool never returns.
Also notes that both locks are flock-based, and that flock is advisory or
per-client on NFS, so single-instance enforcement is not guaranteed
there. Setup gains the venv creation step it assumed.
Phase 7 gains step 12b: measure the sweep cost under a date filter before
deciding whether to optimize it. With reverse=True and no offset_date,
telethon starts the sweep at message id 1 - verified in the pinned wheel -
so --since walks the whole history discarding messages in keep(), and
--until never terminates early. Neither fix is free: offset_date can
start the sweep mid-album, which would take a post_id that is not the
album's lowest id and break filter-invariant identity, and an --until
break assumes dates rise with ids, false for imported history. The
swept-to-kept ratio decides whether either risk is worth taking.
Reports from the review pass are recorded under plans/reports/, including
the findings left unfixed: resolve_entity catching only ValueError,
_find_in_dialogs being the one network loop outside a flood primitive,
SystemExit in session.py routing around the Abort contract, and
completed_at being cleared before any work begins.
The session lock was taken after the thing it guards. It lived inside
connected_client, so connect() and _login() had already written the
shared SQLite session by the time the lock existed. Two runs on the
default session each passed their own per-root lock, both opened the same
database, and the loser exited 3 *after* causing the corruption its
message described. The lock now wraps the whole client lifetime, and the
committable-path check runs first so a refused run leaves no lock file
next to a session it was never allowed to create.
--dry-run is inside that lock too. It writes nothing to the export tree,
but it opens the same session file, which is the resource the lock is
about - so running it alongside an export now needs its own --session.
Refusing to mark an export complete required that *nothing* had
succeeded: `failed and not (downloaded or skipped)`. A single
already-present file made skipped non-zero and disabled the guard
outright, so any resume across a partly-complete export could fail every
remaining file and still stamp completed_at with a cursor at
end-of-history. A broken export then answered "did my export finish?"
with a confident yes. The comparison is now against downloaded + skipped;
a legitimate tail of present files still outnumbers its own stray
failures and completes normally.
Also:
- Renewing an expired file reference counted as a failed attempt, so
expiry on the final attempt burned the last slot and the fresh
reference was never fetched - reported as "exhausted 3 attempts" after
two. The refreshed latch already bounds that arm.
- Sidecar repair scanned a fixed window back from EOF for the last
newline. A partial record larger than the window contains none, so the
file was truncated to end-window: still unreadable, one megabyte
shorter. The window now grows until a newline is found.
- --reset-state made the zeroed cursor durable before rotating the old
sidecar, leaving exactly the mixed-generation log that rotating exists
to prevent. Rotation now precedes State.open, which is sound only on
this path because there is no compatibility check to fail.
- Two resets inside one second silently clobbered the first archive
through os.replace, and the rename was the one here not fsynced.
- title.txt was the only untrusted string written raw. The export root is
safe because the title never becomes a path component, but cat title.txt
handed ANSI escapes and a right-to-left override to the operator. It is
stripped of the same Unicode categories filenames are, from one shared
set so the two rules cannot drift, and written atomically.
Cost, on the most common path of every resume:
- The post directory was fsynced every post, including posts where every
file was already present and nothing had been renamed - one fsync per
post to re-record a directory entry an earlier run had already made
durable. Gated on an actual download.
- An already-present file was stat'd three times: exists(), stat(), and
again inside _result. One stat now, reused as the recorded size.
Moving fsync off the loop is not available: the AST scan forbids
to_thread and run_in_executor, deliberately, because parallelism here
buys nothing and escalates flood waits.
Each fix has a test that fails without it, confirmed by reverting the
fix and re-running. 204 tests to 214, coverage unchanged at 93%.
The real run was dead. Removing Config.max_flood_wait left cli._real_run
still passing it, so every non-dry-run export raised an uncaught
TypeError after taking both locks and writing title.txt and the state
file, but before downloading anything. Dry-run was unaffected, which is
why it looked healthy.
That escaped 180 passing tests because every one of them constructed
Config by hand and none invoked main(). test_cli_dispatch.py now drives
the whole dispatch against a fake client - real run, resume, dry run,
--limit 0, exit 7, exit 3, and an unexpected exception still honoring
the exit contract. That test file, not the one-line fix, is the remedy.
Also:
- AuthKeyNotFound subclasses plain Exception and matched no clause, yet
it is what actually arrives when the connection drops mid-download:
MTProtoSender sets it on every in-flight request. It now aborts with
exit 4 like its rarer RPC-level sibling.
- CdnFileTamperedError was swallowed as a network error and retried
three times. A hash mismatch on CDN-served bytes is a trust event, not
transient noise, so it gets its own clause and its own message.
- Unicode Cf/Zl/Zp characters survived filename sanitization. U+202E
RIGHT-TO-LEFT OVERRIDE renders 'a<RLO>gpj.exe' as 'a.exe.jpg' in any
terminal or file manager - the ANSI vector was closed while the older
extension-spoofing one stayed open. Stripping by Unicode category
subsumes the previous control-character table.
- A run where every file failed and none succeeded no longer sets
completed_at. A dead session looks exactly like that from inside the
loop, and the field is supposed to answer 'did my export finish?'
without guessing.
- Network errors reaching main() are no longer labelled 'filesystem',
and a catch-all keeps any unexpected exception inside the exit-code
contract.
- The retry ladder no longer sleeps after its final attempt.
Verified unchanged: filenames for ordinary extensions, so no previously
downloaded file is orphaned on resume.
Phases 2-6 complete, phase 1 partial, phase 7 pending.
Records what was settled without live credentials: offset_id is an
exclusive lower bound under reverse=True, read from Telethon's iterator
rather than inferred from one sample, so the cursor stores the last
handled id. Also logs the defects found along the way and the two
critical findings from the post-implementation review, plus the one
finding left unresolved rather than decided unilaterally - whether a
file that exhausted its retries should stay abandoned behind a set
completed_at.
Covers why the Bot API structurally cannot do this, what .env may and
may not hold, why the export directory is a chat id, and the sidecar's
union-over-records rule - reading it as last-write-wins reports a
narrower file list than what is on disk after a filter change.
Also states plainly that flood waits have no ceiling by default: the
tool sleeps as long as Telegram demands and logs the computed wake time,
so a long sleep reads as a sleep rather than a hang.
180 tests, all offline against synthetic message stubs: no credentials,
no network, no live group.
The cases that matter are the ones where 'looks like success' and 'is
success' diverge - a cursor ahead of what is on disk, a split album, a
zero-byte file renamed to its target, an empty filter set serialized as
no filter at all.
Two guards are tested for their ability to fail rather than assumed:
the AST scan that forbids parallelism primitives is checked against a
sample that must trip it, and lock contention is proven with a real
second process, since two acquisitions in one process would succeed and
prove nothing.
Exports all media from a group to disk, one folder per logical post with
albums collapsed, alongside a messages.jsonl sidecar linking every file
to its message.
Three properties the design is built around:
- Identity is filter-invariant. A post's folder is the lowest message id
in the full grouped_id group and filenames key on message id, so
--types photo and --types video address the same files. A positional
index would shift whenever a member message was deleted.
- The cursor commits only after a complete post and is unreachable from
every error path, so it can never point past in-flight work.
- Filters are stored verbatim; a mismatch refuses the run and names what
changed, because the cursor means 'handled through here under these
filters'.
Exactly one untrusted string becomes a path component: the filename
leaf, sanitized stem and extension both. Every other component is an int
enforced by type. The export directory is the chat id rather than the
group title, which makes traversal via a renamed group structurally
impossible and stops a mid-export rename from orphaning the download.
Downloads are sequential. Telethon does not parallelize transfers, and
concurrency mainly accelerates flood-wait escalation from seconds to
hours.
Telethon 2.0 is alpha with a renamed API surface, hence the <2 bound.
requirements.txt records 1.44.0, the version the API probes were run
against, so a later resolver bump cannot silently change the semantics
the traversal depends on.