Commit Graph
24 Commits
Author SHA1 Message Date
tiennm99 a81e7eaacb feat: download, upload and drive a chat to completion in one process
Phases 4 through 6: the two legs and the command that joins them.

Downloads go to <name>.part and are renamed only once complete, so a file
without the suffix is always whole. That is what lets the upload leg treat
"exists" as "finished" — the property run.sh could only approximate with a
filename convention plus an age guard, because it could not see inside tdl.

Every finished file is checked against the size Telegram reported, and that
check rather than the error is the authoritative signal. core's downloader logs
a failed transfer and returns nil, and its completion callback is deferred on
that named return, so a failure arrives indistinguishable from a success.
Trusting it would promote a truncated file and archive it as complete.

The disk cap is a semaphore over bytes. A download reserves its own size before
starting and releases it only after the upload confirms, so a slow remote
stalls downloads by itself. Blocking the iterator is safe because the
downloader calls it from its dispatch loop while workers run in a group, so a
blocked iterator never stops the uploads that free the space. Gone with it: the
du polling, the SIGSTOP and SIGCONT suspension, the min-age guard, the
temp-file filter and the sweep-failure counter.

A cap smaller than the largest file is refused up front. The semaphore could
never admit it, and a run blocked on a file it can never start looks exactly
like a stalled remote.

Uploads re-state each object to prove its size before the local copy is gone,
closing a gap where a truncated upload was only noticed by a later verify.

The destination is created before the chat is read. It is also the credentials
check, and doing it first means a bad destination fails in seconds rather than
after a full history walk.

One invocation converges: each item is checked against the index immediately
before download, so there are no passes and re-running is the resume path.
Options that no longer exist say what replaced them instead of failing as
unknown flags.

Verified end to end against the live chat and a scratch remote path: two files
downloaded, uploaded, confirmed present at the right size, staging left empty.
2026-09-06 19:21:46 +07:00
tiennm99 1393abdae2 feat: index a remote and verify a chat against it in-process
Third slice: verify-export.sh, without the subprocess or the python.

One rclone listing builds an in-memory index, and the report is computed from
it. missing-ids.txt and gap.json are gone; so is every python3 heredoc.

Presence is answered from a whole name and never from a message id. The
id-keyed map exists only to tell "absent" apart from "absent, but a stale copy
under an older name is sitting there", and it stays unexported so nothing can
reach for it as an answer. That distinction is the bug this rewrite exists to
remove, so it is enforced by structure rather than by comment.

Names are checked for path containment before use. Storing them verbatim means
a filename chosen by whoever uploaded the file can contain a separator or a
parent reference, and tdl never had to care because its template rewrote those
away. Over-long names are refused for the same reason: the filesystem would
reject them at create time, and a file that can never be written would be
reported absent on every pass forever.

Objects are addressed by the path rclone knows them by, not by the basename
used for matching. The two differ once a remote has directory structure, and
deleting by basename would miss the object or remove a same-named one from the
root. Basenames appearing at more than one path make the snapshot ambiguous, so
they are reported rather than silently resolved.

Indexing runs at full depth with filters cleared. Inheriting RCLONE_MAX_DEPTH
or RCLONE_EXCLUDE would not fail, it would quietly report archived files as
absent and fetch them all again.

Deleting stale copies stays opt-in and confirmed; a non-interactive stdin
declines rather than proceeding. Filenames are quoted wherever they are
printed, so an embedded escape cannot redraw the list an operator approves.

Verified against the live remote: identical to verify-export.sh on the same
state — 18155 expected, 15548 present, 2607 absent, ids 9857-18013, exit 1.
2026-09-06 18:50:53 +07:00
tiennm99 b0c163ed87 feat: resolve chats, walk history, and derive one canonical filename
Second slice: the read path, and the fix for the bug that motivated the
rewrite.

The shell pipeline derived a filename twice. `tdl chat export` wrote the raw
Telegram name into a JSON, while `tdl dl` rendered it through a template whose
default pipes it through filenamify, which rewrites reserved characters and
collapses runs of '!'. A message whose name contained '!!' was therefore looked
up under one name and stored under another; the verifier never found it and
re-fetched it on every pass. Here a single function produces the name, and the
string it returns is used both to test for presence and to write the file, so
the two cannot disagree.

Names are stored exactly as Telegram reports them rather than reproducing
filenamify. That is a deliberate break from what the old pipeline wrote: a file
it stored under a rewritten name is not recognised and will be fetched again.
For the one chat archived so far that is a single file out of 18155, already
removed. Because names are verbatim they are not path-safe, so the code that
turns one into a path must enforce containment.

Walk yields a sequence rather than taking a callback, since the downloader
consumes a pull iterator and range-over-func converts either way without anyone
owning a goroutine. It pages newest first, where the old pipeline went oldest
first, which changes what an interrupted run leaves behind.

Message links are refused rather than guessed at, across every host Telegram
uses and the tg:// forms that carry the message id in a query parameter. A
private channel link and a public message link have the same shape, so the two
are told apart by parsing rather than by pattern.

Verified against the live chat: 18155 media messages, matching the shell
verifier, and every one of the 15548 objects already on the remote is found
under a derived name.
2026-09-06 17:37:18 +07:00
tiennm99 b73b74efd5 feat: add Go binary embedding tdl and rclone as libraries
First slice of replacing the three-script pipeline with one process. The
scripts coordinate tdl and rclone as separate programs, so everything
expensive in them exists to work around the fact that neither can see the
other's state. A single process does not need that machinery.

This slice covers only the foundations: open the session tdl login already
wrote, resolve an rclone destination, and report on both via a doctor
command. Downloading, uploading and verification follow.

The session store is shared with the tdl CLI rather than copied, so the two
cannot run against one namespace at the same time; -n selects another.

AppID and AppHash are read from the store rather than hardcoded, because a
session is bound to the application that created it and tdl records which
one it used.

No middlewares are passed to tclient.New, which already prepends its own
defaults; the DC pool gets them instead, since gotd applies a client's
middlewares only to direct invocations and not to pooled connections.

The rclone config is loaded up front because the lazy path calls os.Exit on
a config it cannot read, which would bypass every defer and exit with the
code this tool reserves for an incomplete run.

Exit codes follow the shell pipeline: 0 ok, 1 incomplete, 2 usage, 3 remote
failure, 130 SIGINT, 143 SIGTERM. Cancellation is checked explicitly after
the Telegram client returns, because gotd reports an interrupted run as
success and a driver would read that as a finished archive.
2026-09-06 17:00:45 +07:00
tiennm99 70394b9763 fix: match exported filenames exactly and report misnamed copies
The verifier reconstructs {DialogID}_{MessageID}_{FileName} from the export
and looks it up in the remote listing. A file stored under any other name is
not the file the export asked for, so it counts as absent however close the
name looks.

That leaves a stale copy behind when the re-download lands beside it, so
absent files whose message id is already present on the remote are now listed
separately with both names, ready to be removed.

Remote paths are reduced to their basename first, so a grouped subdirectory
layout matches the same way a flat one does.
2026-09-06 17:00:23 +07:00
tiennm99 6fb77ce633 chore: tune defaults for a slow async-commit remote
Backends that finish an upload as a server-side task, pikpak among them,
leave it pending long enough that rclone abandons the transfer once
--low-level-retries polls run out. Default RCLONE_TRANSFERS to 2 to keep
that task queue short and RCLONE_LOW_LEVEL_RETRIES to 20 to wait a slow
commit out. Both use the :=default form, so an exported value still wins.

Shorten the --min-age default to 45s: excluding '*.tmp' is what actually
keeps an unfinished download from being uploaded, so the age guard only
needs to cover a rename racing a sweep, and two minutes of dwell made
staging larger than it had to be.

Sweep every 120s rather than 300s in the driver, which shrinks the backlog
each sweep has to move, and raise the pass limit to 30 to suit chats that
need more rounds than 20 to converge.
2026-09-04 11:16:53 +07:00
tiennm99 f77f9c4c1a feat: cap staging size and report progress on end-of-run sweeps
Add -m SIZE to run.sh, passed through by export-until-complete.sh, so the
staging directory cannot outgrow the local disk when tdl downloads faster
than rclone uploads. Staging is measured every 10 seconds, independently of
the sweep interval; at the cap tdl is suspended with SIGSTOP and rclone
sweeps until staging is back under it, then tdl resumes and --continue picks
its .tmp files back up. The drain loops while it frees space and gives up
once only in-flight .tmp files remain, which no sweep can move, so a cap set
below what the concurrent downloads hold warns instead of stalling. The
failure streak still counts only sweeps that actually ran, so a quiet cap
check cannot clear it.

The sweeps that run once at the end now report progress rather than going
quiet for minutes: rclone's redrawn bar on a terminal, one-line stats every
30s when output is redirected. Stats are logged at INFO, so only the stats
level is raised to NOTICE, avoiding the line-per-file output -v would add.

Also send SIGCONT before SIGTERM when stopping tdl, since a suspended
process never sees the term.
2026-09-04 10:14:09 +07:00
tiennm99 913b0d0b7d feat: verify export completeness and drive runs until complete
run.sh finishes when tdl finishes, which does not mean every file arrived: a
dropped session, a stalled remote or an interrupted pass leaves gaps that
nothing reported.

verify-export.sh rebuilds the filename tdl produces for each media message in
the export JSON and checks the remote and staging for it. Messages carrying no
media are skipped rather than counted as gaps. A zero-byte file counts as
missing because rclone overwrites a size-mismatched destination, so a retry
repairs it; files under 1 KiB are reported but not retried, since real media is
sometimes that small and retrying them would never terminate.

export-until-complete.sh loops that check with run.sh, narrowing the export to
the outstanding ids each pass. It stops when the verifier reports complete, when
a pass fetches nothing new, when the remote runs low on space, or on Ctrl-C.
tdl's progress bar is shown on a terminal and suppressed when output is
redirected, so logs stay readable.

Exports, staging and logs are now ignored: they hold real message ids, file
names and text from an account.
2026-09-03 22:42:32 +07:00
tiennm99 40bdec66d0 feat: accept chat ids, usernames and links, and keep exports per chat
-c already passed through to tdl, which resolves a numeric argument as an
MTProto id and anything else through gotd's resolver (@name, name, t.me and
tg:// links). Two forms needed handling: a Bot API '-100' prefixed id, which
MTProto does not use and which tdl never strips, is now converted; and a
message link, which names a message rather than a chat, is rejected with a
clear error instead of a resolver failure.

The export JSON now defaults to export-<chat>.json when -c is given, so
exporting a second chat from the same directory cannot silently download the
first chat again from a stale export.json. An existing file is still reused,
but the script now says which file it is using.
2026-08-26 22:23:51 +07:00
tiennm99 4ae19e32ec feat: target any rclone remote instead of only WebDAV
The pipeline never depended on WebDAV; only the naming, docs and preflight
did. Replace the WebDAV-root listing with a portable check: a named remote
must appear in 'rclone listremotes', and creating the destination proves
reachability and credentials. On-the-fly connection strings have no config
entry, so the name check is skipped for them.

Read the remote list into a variable rather than piping it into 'grep -q':
grep exits on the first match and kills rclone with SIGPIPE, which pipefail
reports as failure, intermittently rejecting a configured remote.

Document that rclone's flags are all settable through their RCLONE_*
environment variables, so the upload side is tunable without new options.
2026-08-25 16:07:03 +07:00
tiennm99 3bc4fd846b docs: rewrite the README around run.sh, and rename the script
The README now covers only what the tool does and how to run it: setup,
options, tdl flag pass-through, subset exports, resume semantics, the
failure modes it guards against, and exit codes.

Flags are documented from tdl's own source rather than assumed: dl -f takes
the exported JSON while dl -i/-e are file-extension filters, and the
expression filter is chat export -f. Chats are addressed by id or domain.
2026-08-25 16:02:11 +07:00
tiennm99 9ff0d259be docs: document the pipeline script and the partial-file caveat
Also warn that the inline rclone loops need --exclude '*.tmp' themselves.
Without it a download stalled by a flood wait is uploaded half-written and
its local copy deleted, which breaks tdl --continue for that file.
2026-08-25 15:48:53 +07:00
tiennm99 7d20f5667c feat: add a tdl-to-webdav rolling pipeline script
tdl has no remote destination, so exporting a group larger than local disk
needs a staging dir plus a concurrent uploader. The script runs tdl into a
small staging dir while rclone moves finished files out to WebDAV, so local
disk only holds the files in flight plus one sync interval of throughput.

Partial downloads are excluded by name rather than by age: tdl writes
<name>.tmp and renames on completion, so a download stalled by a flood wait
stops touching its .tmp and would otherwise age past --min-age and be
uploaded half-written, destroying resume for that file.
--delete-empty-src-dirs runs only in the final sweep, because removing a
directory under a running tdl makes it fail to create its next file. tdl is
stopped on exit, interrupt or termination so no download is orphaned, the
unrestricted sweep runs only after tdl exits 0, and five consecutive rclone
failures abort the run before staging fills the disk.
2026-08-25 15:48:46 +07:00
tiennm99 ec0584e34e chore: churn project - strip implementation, keep tdl + webdav instructions
The Python/Telethon exporter is retired in favor of iyear/tdl, which
covers the same purpose and is actively maintained. README now carries
only migration instructions, including the rclone rolling-pipeline
recipe for exporting to WebDAV. Full implementation remains at d170986.
2026-08-25 15:03:32 +07:00
tiennm99 d1709866e3 docs: correct the dry-run lock and exit-code claims, record the sweep measurement
Two README statements were false rather than merely incomplete. "--dry-run
takes no lock at all, so it can always be run against an export that is
currently in progress" contradicted the paragraph four lines above it
warning that two clients sharing one session is a corruption hazard:
dry-run opens the session file like any other run. Both places now say
what is true - no *export* lock, but a session lock, so an alongside
dry-run needs its own --session.

The exit table promised 1 for a bad argument. argparse exits 2 on its
own, before main()'s try block can map anything, so a bad --limit
collided with the dry-run SHORT verdict. Documented as the shared code
it is instead of a code the tool never returns.

Also notes that both locks are flock-based, and that flock is advisory or
per-client on NFS, so single-instance enforcement is not guaranteed
there. Setup gains the venv creation step it assumed.

Phase 7 gains step 12b: measure the sweep cost under a date filter before
deciding whether to optimize it. With reverse=True and no offset_date,
telethon starts the sweep at message id 1 - verified in the pinned wheel -
so --since walks the whole history discarding messages in keep(), and
--until never terminates early. Neither fix is free: offset_date can
start the sweep mid-album, which would take a post_id that is not the
album's lowest id and break filter-invariant identity, and an --until
break assumes dates rise with ids, false for imported history. The
swept-to-kept ratio decides whether either risk is worth taking.

Reports from the review pass are recorded under plans/reports/, including
the findings left unfixed: resolve_entity catching only ValueError,
_find_in_dialogs being the one network loop outside a flood primitive,
SystemExit in session.py routing around the Abort contract, and
completed_at being cleared before any work begins.
2026-08-22 23:55:00 +07:00
tiennm99 c37598e008 fix: lock the session before opening it, and stop a failed export reading as finished
The session lock was taken after the thing it guards. It lived inside
connected_client, so connect() and _login() had already written the
shared SQLite session by the time the lock existed. Two runs on the
default session each passed their own per-root lock, both opened the same
database, and the loser exited 3 *after* causing the corruption its
message described. The lock now wraps the whole client lifetime, and the
committable-path check runs first so a refused run leaves no lock file
next to a session it was never allowed to create.

--dry-run is inside that lock too. It writes nothing to the export tree,
but it opens the same session file, which is the resource the lock is
about - so running it alongside an export now needs its own --session.

Refusing to mark an export complete required that *nothing* had
succeeded: `failed and not (downloaded or skipped)`. A single
already-present file made skipped non-zero and disabled the guard
outright, so any resume across a partly-complete export could fail every
remaining file and still stamp completed_at with a cursor at
end-of-history. A broken export then answered "did my export finish?"
with a confident yes. The comparison is now against downloaded + skipped;
a legitimate tail of present files still outnumbers its own stray
failures and completes normally.

Also:

- Renewing an expired file reference counted as a failed attempt, so
  expiry on the final attempt burned the last slot and the fresh
  reference was never fetched - reported as "exhausted 3 attempts" after
  two. The refreshed latch already bounds that arm.
- Sidecar repair scanned a fixed window back from EOF for the last
  newline. A partial record larger than the window contains none, so the
  file was truncated to end-window: still unreadable, one megabyte
  shorter. The window now grows until a newline is found.
- --reset-state made the zeroed cursor durable before rotating the old
  sidecar, leaving exactly the mixed-generation log that rotating exists
  to prevent. Rotation now precedes State.open, which is sound only on
  this path because there is no compatibility check to fail.
- Two resets inside one second silently clobbered the first archive
  through os.replace, and the rename was the one here not fsynced.
- title.txt was the only untrusted string written raw. The export root is
  safe because the title never becomes a path component, but cat title.txt
  handed ANSI escapes and a right-to-left override to the operator. It is
  stripped of the same Unicode categories filenames are, from one shared
  set so the two rules cannot drift, and written atomically.

Cost, on the most common path of every resume:

- The post directory was fsynced every post, including posts where every
  file was already present and nothing had been renamed - one fsync per
  post to re-record a directory entry an earlier run had already made
  durable. Gated on an actual download.
- An already-present file was stat'd three times: exists(), stat(), and
  again inside _result. One stat now, reused as the recorded size.

Moving fsync off the loop is not available: the AST scan forbids
to_thread and run_in_executor, deliberately, because parallelism here
buys nothing and escalates flood waits.

Each fix has a test that fails without it, confirmed by reverting the
fix and re-running. 204 tests to 214, coverage unchanged at 93%.
2026-08-22 23:20:00 +07:00
tiennm99 3a51f33068 fix: repair the real export path and close taxonomy and sanitizer gaps
The real run was dead. Removing Config.max_flood_wait left cli._real_run
still passing it, so every non-dry-run export raised an uncaught
TypeError after taking both locks and writing title.txt and the state
file, but before downloading anything. Dry-run was unaffected, which is
why it looked healthy.

That escaped 180 passing tests because every one of them constructed
Config by hand and none invoked main(). test_cli_dispatch.py now drives
the whole dispatch against a fake client - real run, resume, dry run,
--limit 0, exit 7, exit 3, and an unexpected exception still honoring
the exit contract. That test file, not the one-line fix, is the remedy.

Also:

- AuthKeyNotFound subclasses plain Exception and matched no clause, yet
  it is what actually arrives when the connection drops mid-download:
  MTProtoSender sets it on every in-flight request. It now aborts with
  exit 4 like its rarer RPC-level sibling.
- CdnFileTamperedError was swallowed as a network error and retried
  three times. A hash mismatch on CDN-served bytes is a trust event, not
  transient noise, so it gets its own clause and its own message.
- Unicode Cf/Zl/Zp characters survived filename sanitization. U+202E
  RIGHT-TO-LEFT OVERRIDE renders 'a<RLO>gpj.exe' as 'a.exe.jpg' in any
  terminal or file manager - the ANSI vector was closed while the older
  extension-spoofing one stayed open. Stripping by Unicode category
  subsumes the previous control-character table.
- A run where every file failed and none succeeded no longer sets
  completed_at. A dead session looks exactly like that from inside the
  loop, and the field is supposed to answer 'did my export finish?'
  without guessing.
- Network errors reaching main() are no longer labelled 'filesystem',
  and a catch-all keeps any unexpected exception inside the exit-code
  contract.
- The retry ladder no longer sleeps after its final attempt.

Verified unchanged: filenames for ordinary extensions, so no previously
downloaded file is orphaned on resume.
2026-08-12 23:55:00 +07:00
tiennm99 a12ddcb9c0 docs: sync the plan with implementation and review outcomes
Phases 2-6 complete, phase 1 partial, phase 7 pending.

Records what was settled without live credentials: offset_id is an
exclusive lower bound under reverse=True, read from Telethon's iterator
rather than inferred from one sample, so the cursor stores the last
handled id. Also logs the defects found along the way and the two
critical findings from the post-implementation review, plus the one
finding left unresolved rather than decided unilaterally - whether a
file that exhausted its retries should stay abandoned behind a set
completed_at.
2026-08-12 23:40:00 +07:00
tiennm99 30f68b43ff docs: document credentials, resume semantics and flood-wait behavior
Covers why the Bot API structurally cannot do this, what .env may and
may not hold, why the export directory is a chat id, and the sidecar's
union-over-records rule - reading it as last-write-wins reports a
narrower file list than what is on disk after a filter change.

Also states plainly that flood waits have no ceiling by default: the
tool sleeps as long as Telegram demands and logs the computed wake time,
so a long sleep reads as a sleep rather than a hang.
2026-08-12 22:15:00 +07:00
tiennm99 5bf0992fe6 test: cover path safety, resume, flood waits and the concurrency rule
180 tests, all offline against synthetic message stubs: no credentials,
no network, no live group.

The cases that matter are the ones where 'looks like success' and 'is
success' diverge - a cursor ahead of what is on disk, a split album, a
zero-byte file renamed to its target, an empty filter set serialized as
no filter at all.

Two guards are tested for their ability to fail rather than assumed:
the AST scan that forbids parallelism primitives is checked against a
sample that must trip it, and lock contention is proven with a real
second process, since two acquisitions in one process would succeed and
prove nothing.
2026-08-12 21:30:00 +07:00
tiennm99 275ec47576 feat: add tg-export, a resumable Telegram group media exporter
Exports all media from a group to disk, one folder per logical post with
albums collapsed, alongside a messages.jsonl sidecar linking every file
to its message.

Three properties the design is built around:

- Identity is filter-invariant. A post's folder is the lowest message id
  in the full grouped_id group and filenames key on message id, so
  --types photo and --types video address the same files. A positional
  index would shift whenever a member message was deleted.
- The cursor commits only after a complete post and is unreachable from
  every error path, so it can never point past in-flight work.
- Filters are stored verbatim; a mismatch refuses the run and names what
  changed, because the cursor means 'handled through here under these
  filters'.

Exactly one untrusted string becomes a path component: the filename
leaf, sanitized stem and extension both. Every other component is an int
enforced by type. The export directory is the chat id rather than the
group title, which makes traversal via a renamed group structurally
impossible and stops a mid-export rename from orphaning the download.

Downloads are sequential. Telethon does not parallelize transfers, and
concurrency mainly accelerates flood-wait escalation from seconds to
hours.
2026-08-12 20:10:00 +07:00
tiennm99 65f8abdf18 build: add packaging and pin the verified Telethon version
Telethon 2.0 is alpha with a renamed API surface, hence the <2 bound.
requirements.txt records 1.44.0, the version the API probes were run
against, so a later resolver bump cannot silently change the semantics
the traversal depends on.
2026-08-12 15:45:00 +07:00
tiennm99 de9a08796e chore: add gitignore covering sessions, exports, and credentials 2026-08-12 15:01:53 +07:00
tiennm99 4ce369d953 Initial commit 2026-08-11 14:07:26 +07:00