15 Commits
Author SHA1 Message Date
tiennm99 b459b1f5e9 fix: re-read each message for a live file reference before downloading it
The walk collects every item's file reference up front, and Telegram expires
those. The lifetime is undocumented; a run was seen to outlast an hour and fail
before two, which is less than a large chat spends downloading. So every fetch
past that point returned FILE_REFERENCE_EXPIRED, five in a row tripped the
breaker, and the run abandoned the rest of its todo list.

Telegram's contract is to cache a reference together with the source it came
from and re-read that source when it expires, so each message is re-read for a
live reference immediately before its own download. That is one round trip per
transfer and needs no guess at how long a reference lives. It runs after the
staging acquire rather than before: that acquire is the run's brake and blocks
for as long as the destination is slow, which on a stalled remote is long enough
to expire a fresh reference all over again.

A re-read that describes a different file — different name or different length —
is refused, because archiving it under the name this run recorded would store
the wrong bytes. That case, a deleted message, and one that no longer holds a
file are all skipped: nothing can be archived for them, and one must not strand
the rest of the chat.

Any other re-read failure is recorded as a failed download instead, because that
is what it is. The connection that serves the re-read is the one that serves the
file, so these arrive exactly when transfers are failing too. An item that never
reaches a worker never reaches OnAdd or OnDone, so on the skip path it would
move no counter, feed no streak and produce no outcome to retry — a source
outage would walk the whole list one dead round trip at a time and still exit as
merely incomplete.

tutil.GetSingleMessage is not reused for the re-read: it reports a deleted
message three different ways, two of them wrapping a nil error, and only one is
distinguishable — so deletions would be misclassified as transient and spend the
breaker's streak.
2026-09-08 22:47:52 +07:00
tiennm99 efb29c9c39 fix: report why a download failed, retry it, and stop a failing source
core's downloader logs a failed transfer and returns nil, and nothing here
installed a logger, so logctx handed out a nop and the reason was destroyed.
The size check in finish was all that survived, which reported a failed
two-gigabyte fetch as "short download: got 0 bytes" — a symptom with no cause
an operator can act on. The logger it reaches for now feeds a core that keeps
error entries and stores them on the element they belong to, so finish reports
the reason ahead of the byte count.

A failed download is also retried inside the run, up to three passes over
whatever is still missing and spaced minutes apart, because the failure this
repairs is transient: a dead connection takes every transfer in flight with it
and all of them are fetchable again afterwards, while leaving them to the next
run costs a full re-walk of the chat and a re-index of the remote first. An
item is counted once however many attempts it took, and a pass that recovers
reports nothing.

The download leg gains the breaker the upload leg already had. A source that
refuses everything refuses the rest of the list in milliseconds, so the pass
gives up after five failures in a row and the run exits 3 rather than spending
the entire outstanding list finding that out.
2026-09-08 18:20:24 +07:00
tiennm99 846c2bafe9 feat: count download and upload separately
One combined figure could say how much had moved but not which half was
moving it. That is the question a slowing run actually raises: is Telegram
the bottleneck, or the remote? The two legs now have their own totals, files
and bytes each.

The gap between them is the useful part. They normally track a file or two
apart; a widening gap is the remote falling behind, which is also staging
filling up — visible now before the byte cap starts throttling downloads.

Uploads advance a whole file at a time because rclone's MoveFile is one
blocking call with no byte callbacks, so a partial upload counts for nothing
until it lands. That is the honest reading anyway: what the upload total
reports is what is actually on the remote.

Both renderers read the same counters, so the terminal and a captured log
cannot disagree about the numbers.
2026-09-07 00:05:39 +07:00
tiennm99 a0aa0c98fa test: prove an upload recovers on retry after a real move failure 2026-09-06 22:17:47 +07:00
tiennm99 b05247abed fix: retry a failed upload instead of abandoning the file
pikpak commits an upload as a server-side async task, and rclone polls that
task only as long as its low-level retries last. A task still in
PHASE_TYPE_PENDING when that budget runs out is reported as "can't verify the
task is completed" — and checking the live archive ten minutes later, those
tasks had not committed: the objects were simply absent.

The transfer itself was fine; only the confirmation timed out. Abandoning the
file meant re-downloading it from Telegram on the next pass, which for this
archive can be two gigabytes, so a run against a slow remote paid for the same
bytes repeatedly. The staged copy is still on disk when a move fails — rclone
says so in the same breath — so another attempt costs seconds instead.

Every attempt after the first checks the remote before re-uploading. A pending
task may have committed during the backoff, and pikpak allows two files under
one name, so re-uploading blind is how one file becomes two — which
verification then reports as an ambiguous basename on every future run.

rclone's log now goes through the renderer. It writes to stderr on its own
schedule, so its error lines were landing mid-redraw and shredding the display
exactly when there was most to read.

Also: "capped at uncapped" now reads "uncapped", and the ETA column fits the
three-digit hour counts a slow remote produces.
2026-09-06 21:54:38 +07:00
tiennm99 4624c506e5 feat: show per-file progress and state the plan before a run starts
A run reported one aggregate line, which could say how much was done but never
what was happening: which files were moving, whether a stall was a slow
download or a slow upload, or how long the rest would take. On a transfer
measured in hours those are the only questions worth answering.

The pipeline now emits a per-item lifecycle rather than only Stats, and a
terminal renders it as an overall bar plus one bar per file in flight. Uploads
get a spinner rather than a bar because rclone's MoveFile is a single blocking
call with no byte callbacks; naming the file is still the point, since a run
that looks stalled is usually waiting on one large object.

Before the transfer, the run states what it found and what it will do. The
survey breaks the outstanding set down by reason — never fetched, zero-byte,
wrong size, unarchivable — because that is the difference between a run that
will converge and one that cannot, and the old "N to fetch" hid it. The plan
states the cap, staging path and concurrency, so a wrong setting is visible
before hours of transfer rather than after.

Redirected output keeps plain lines and gains one per archived file. Bars are
continuous cursor movement, and a captured log of them is what tdl's progress
bar did to the shell pipeline's logs.

cmd/uidemo renders the whole thing against fake data. It is how the layout was
checked without a session, and it earned its place immediately: the overall bar
sat at zero because Stats only tracked the file count and never advanced it.
2026-09-06 21:21:07 +07:00
tiennm99 fafeab4750 feat: report progress while reading a chat and listing a remote
Both phases ran for minutes printing nothing between their opening line and
their result, so a working run looked exactly like a hung one — which is how it
was reported.

The message counter lives in Walk rather than in the caller's loop because most
of a chat is not media: text-only and service messages are filtered out inside
the walk, so a caller counting yielded items still sees nothing while crossing a
long stretch of conversation.

Cadence follows the existing reporter: a terminal redraws one line, a redirected
run gets a periodic one, since ANSI redraws are what turned the shell pipeline's
captured logs into megabytes of control characters.
2026-09-06 21:05:59 +07:00
tiennm99 aea29826c2 fix: detect truncated uploads, and stop reporting unreachable work as retryable
Verification matched on name and non-zero size, so an upload that died partway
was counted archived permanently. Comparing against the size Telegram reports
found six such objects in the live archive, one of them 221 MiB standing in for
a 2006 MiB video. They are folded into the outstanding set; rclone overwrites a
size mismatch, so another pass repairs them.

A zero Report claimed the archive was complete — nothing expected, nothing
missing — which is the value both commands hold before their Telegram callback
populates it, so any early return printed COMPLETE and exited 0 on an untouched
chat. A Report now knows whether it ran.

A name that can never be written kept the run outstanding forever while the
download set deliberately excluded it, so a driver looping on "incomplete"
walked the whole history and re-indexed the whole remote every pass for work
that could not be done. Such a run now reports STALLED and exits 4.

A destination that stopped accepting uploads was reported and then discarded,
exiting 1. The same driver would retry against a full or unreachable remote
indefinitely, downloading gigabytes each pass to upload none. It exits 3.

Free space is re-checked during the run, not only before it. An archive this
size runs for hours, and the remote can fill in the middle; discovering it
through five failed multi-gigabyte uploads wastes the download for all of them.

Also: parseSize silently wrapped to a negative or zero on a large input, which
reads downstream as "no cap"; list printed attacker-chosen filenames raw, so a
tab shifted the columns and an escape sequence reached the terminal; a missing
backend blamed credentials rather than the build; humanBytes indexed past its
unit table above 1 PiB; a second Init reported success against a config that
never loaded; and sync did not surface the basename collisions verify warned
about, though sync is the command that acts on the verdict.

The env-override test could not observe what it claimed: rclone reads RCLONE_*
at package init, so t.Setenv came too late and the assertion held with the
guard removed. It runs in a subprocess now, as does the new one covering the
index against inherited filters.
2026-09-06 20:44:08 +07:00
tiennm99 19837ceb85 fix: reject names rclone rewrites, and stop resolving link markers as chats
rclone does not address a file by the bytes os.OpenFile wrote. Names handed to
an Fs go through the backend encoder and names listed back are re-encoded to
the standard set, neither of which the write path performs. A filename holding
one of the rewritten characters was therefore stored under one string and
looked up under another: confirmed against the local backend, where a written
"a‛b.jpg" reports object not found and a written "a\nb.jpg" lists back as
"a␊b.jpg". That is the same two-derivation divergence this program was written
to remove, with rclone's encoder standing where filenamify used to. Such names
are rejected, not encoded, for the same reason every other name is.

The length limit ignored the ".part" suffix that is opened first, so a name
just inside NAME_MAX passed the check and then failed to open on every pass,
stalling the walk on that message forever. The suffix now lives beside the
limit that has to account for it.

filter.NewFilter(nil) does not build a neutral filter; it copies the package
global, which rclone has already filled from RCLONE_*. Indexing inherited the
operator's environment, so a stray RCLONE_MIN_SIZE emptied the index and
re-downloaded the archive. Every narrowing field is now set explicitly and the
result is asserted inactive. A subprocess test covers it, since the env is read
at package init and t.Setenv is too late to observe anything.

An ErrorDirNotFound from a subdirectory was also treated as an empty
destination, returning a partial index as authoritative.

t.me/c/<id> and t.me/s/<name> were passed through whole, and gotd reads the
first path component as the username — resolving "c" or "s", which is a
confusing failure at best and someone else's chat at worst, since
one-character usernames exist. Both now yield the chat, and t.me/s/<name>/<id>
is refused like any other message link. Two tests asserted the old behaviour.

core's dcpool.Takeout deadlocks when takeout init fails: it holds the pool
mutex and recovers by calling Client, which takes the same non-reentrant
mutex. Telegram returns TAKEOUT_INIT_DELAY for a takeout started recently and
takeout is on by default, so two runs in succession hang the process with no
output and no response to cancellation. The session is established once here
instead, falling back to a plain client, and the pool's own Takeout is never
called.
2026-09-06 20:28:09 +07:00
tiennm99 3e75138c59 fix: clean up a fragment left by a failed upload, and skip unwritable names
rclone writes straight to the final remote name on any backend that does not
advertise PartialUploads, and cleans up after a failed Put only when it did
not. Pikpak advertises neither, so a transfer that died halfway left a
fragment under exactly the name verification matches on — counted archived by
that run and every run after it, with the local copy already deleted. A failed
move now looks for that object and removes it, leaving a complete one alone
since pikpak's async commit can still land it correctly.

An unwritable filename ended the whole walk, so one hostile name could strand
every message behind it. It is skipped and reported instead. The comment there
had described that behaviour all along.

Also: remove the part file when promoting it fails, since the caller hands the
reservation back and the cap would stay over-committed; report upload failures
alongside a download error rather than instead of it, which on Ctrl-C hid that
finished files had been discarded.

Run had no test of its own because it called Download directly. That step is
now indirected, covering the properties only the composition has: uploads
closed after the last send, the budget balanced across failures, and a tripped
breaker halting downloads rather than walking the whole chat.
2026-09-06 19:51:24 +07:00
tiennm99 bd816156eb fix: join download workers before closing the upload channel
core's Download returns without waiting on its worker group when the
iterator reports an error, so surfacing one through Iter.Err left workers
sending into a channel the caller had already closed. The iterator now
always reports a nil error and stashes the real one, read after Download
returns.

The circuit breaker cancelled only the upload context, which left
downloads running full speed against a remote refusing them: every file
stayed in staging and every reservation came back, so a broken remote
filled local disk faster than a working one. It now stops the download
iterator instead.

A failed move leaves the local copy in place, so the byte reservation
cannot be handed back until the file is removed. A confirmed-short object
is deleted rather than left under a name verification would count as
archived forever.

Also: release the reservation when opening the staging file fails, report
results even when an upload errored, bound the recorded errors, and drop
the reporter's lock before writing so terminal latency cannot throttle
downloads.
2026-09-06 19:37:30 +07:00
tiennm99 a81e7eaacb feat: download, upload and drive a chat to completion in one process
Phases 4 through 6: the two legs and the command that joins them.

Downloads go to <name>.part and are renamed only once complete, so a file
without the suffix is always whole. That is what lets the upload leg treat
"exists" as "finished" — the property run.sh could only approximate with a
filename convention plus an age guard, because it could not see inside tdl.

Every finished file is checked against the size Telegram reported, and that
check rather than the error is the authoritative signal. core's downloader logs
a failed transfer and returns nil, and its completion callback is deferred on
that named return, so a failure arrives indistinguishable from a success.
Trusting it would promote a truncated file and archive it as complete.

The disk cap is a semaphore over bytes. A download reserves its own size before
starting and releases it only after the upload confirms, so a slow remote
stalls downloads by itself. Blocking the iterator is safe because the
downloader calls it from its dispatch loop while workers run in a group, so a
blocked iterator never stops the uploads that free the space. Gone with it: the
du polling, the SIGSTOP and SIGCONT suspension, the min-age guard, the
temp-file filter and the sweep-failure counter.

A cap smaller than the largest file is refused up front. The semaphore could
never admit it, and a run blocked on a file it can never start looks exactly
like a stalled remote.

Uploads re-state each object to prove its size before the local copy is gone,
closing a gap where a truncated upload was only noticed by a later verify.

The destination is created before the chat is read. It is also the credentials
check, and doing it first means a bad destination fails in seconds rather than
after a full history walk.

One invocation converges: each item is checked against the index immediately
before download, so there are no passes and re-running is the resume path.
Options that no longer exist say what replaced them instead of failing as
unknown flags.

Verified end to end against the live chat and a scratch remote path: two files
downloaded, uploaded, confirmed present at the right size, staging left empty.
2026-09-06 19:21:46 +07:00
tiennm99 1393abdae2 feat: index a remote and verify a chat against it in-process
Third slice: verify-export.sh, without the subprocess or the python.

One rclone listing builds an in-memory index, and the report is computed from
it. missing-ids.txt and gap.json are gone; so is every python3 heredoc.

Presence is answered from a whole name and never from a message id. The
id-keyed map exists only to tell "absent" apart from "absent, but a stale copy
under an older name is sitting there", and it stays unexported so nothing can
reach for it as an answer. That distinction is the bug this rewrite exists to
remove, so it is enforced by structure rather than by comment.

Names are checked for path containment before use. Storing them verbatim means
a filename chosen by whoever uploaded the file can contain a separator or a
parent reference, and tdl never had to care because its template rewrote those
away. Over-long names are refused for the same reason: the filesystem would
reject them at create time, and a file that can never be written would be
reported absent on every pass forever.

Objects are addressed by the path rclone knows them by, not by the basename
used for matching. The two differ once a remote has directory structure, and
deleting by basename would miss the object or remove a same-named one from the
root. Basenames appearing at more than one path make the snapshot ambiguous, so
they are reported rather than silently resolved.

Indexing runs at full depth with filters cleared. Inheriting RCLONE_MAX_DEPTH
or RCLONE_EXCLUDE would not fail, it would quietly report archived files as
absent and fetch them all again.

Deleting stale copies stays opt-in and confirmed; a non-interactive stdin
declines rather than proceeding. Filenames are quoted wherever they are
printed, so an embedded escape cannot redraw the list an operator approves.

Verified against the live remote: identical to verify-export.sh on the same
state — 18155 expected, 15548 present, 2607 absent, ids 9857-18013, exit 1.
2026-09-06 18:50:53 +07:00
tiennm99 b0c163ed87 feat: resolve chats, walk history, and derive one canonical filename
Second slice: the read path, and the fix for the bug that motivated the
rewrite.

The shell pipeline derived a filename twice. `tdl chat export` wrote the raw
Telegram name into a JSON, while `tdl dl` rendered it through a template whose
default pipes it through filenamify, which rewrites reserved characters and
collapses runs of '!'. A message whose name contained '!!' was therefore looked
up under one name and stored under another; the verifier never found it and
re-fetched it on every pass. Here a single function produces the name, and the
string it returns is used both to test for presence and to write the file, so
the two cannot disagree.

Names are stored exactly as Telegram reports them rather than reproducing
filenamify. That is a deliberate break from what the old pipeline wrote: a file
it stored under a rewritten name is not recognised and will be fetched again.
For the one chat archived so far that is a single file out of 18155, already
removed. Because names are verbatim they are not path-safe, so the code that
turns one into a path must enforce containment.

Walk yields a sequence rather than taking a callback, since the downloader
consumes a pull iterator and range-over-func converts either way without anyone
owning a goroutine. It pages newest first, where the old pipeline went oldest
first, which changes what an interrupted run leaves behind.

Message links are refused rather than guessed at, across every host Telegram
uses and the tg:// forms that carry the message id in a query parameter. A
private channel link and a public message link have the same shape, so the two
are told apart by parsing rather than by pattern.

Verified against the live chat: 18155 media messages, matching the shell
verifier, and every one of the 15548 objects already on the remote is found
under a derived name.
2026-09-06 17:37:18 +07:00
tiennm99 b73b74efd5 feat: add Go binary embedding tdl and rclone as libraries
First slice of replacing the three-script pipeline with one process. The
scripts coordinate tdl and rclone as separate programs, so everything
expensive in them exists to work around the fact that neither can see the
other's state. A single process does not need that machinery.

This slice covers only the foundations: open the session tdl login already
wrote, resolve an rclone destination, and report on both via a doctor
command. Downloading, uploading and verification follow.

The session store is shared with the tdl CLI rather than copied, so the two
cannot run against one namespace at the same time; -n selects another.

AppID and AppHash are read from the store rather than hardcoded, because a
session is bound to the application that created it and tdl records which
one it used.

No middlewares are passed to tclient.New, which already prepends its own
defaults; the DC pool gets them instead, since gotd applies a client's
middlewares only to direct invocations and not to pooled connections.

The rclone config is loaded up front because the lazy path calls os.Exit on
a config it cannot read, which would bypass every defer and exit with the
code this tool reserves for an incomplete run.

Exit codes follow the shell pipeline: 0 ok, 1 incomplete, 2 usage, 3 remote
failure, 130 SIGINT, 143 SIGTERM. Cancellation is checked explicitly after
the Telegram client returns, because gotd reports an interrupted run as
success and a driver would read that as a finished archive.
2026-09-06 17:00:45 +07:00