Phases 4 through 6: the two legs and the command that joins them. Downloads go to <name>.part and are renamed only once complete, so a file without the suffix is always whole. That is what lets the upload leg treat "exists" as "finished" — the property run.sh could only approximate with a filename convention plus an age guard, because it could not see inside tdl. Every finished file is checked against the size Telegram reported, and that check rather than the error is the authoritative signal. core's downloader logs a failed transfer and returns nil, and its completion callback is deferred on that named return, so a failure arrives indistinguishable from a success. Trusting it would promote a truncated file and archive it as complete. The disk cap is a semaphore over bytes. A download reserves its own size before starting and releases it only after the upload confirms, so a slow remote stalls downloads by itself. Blocking the iterator is safe because the downloader calls it from its dispatch loop while workers run in a group, so a blocked iterator never stops the uploads that free the space. Gone with it: the du polling, the SIGSTOP and SIGCONT suspension, the min-age guard, the temp-file filter and the sweep-failure counter. A cap smaller than the largest file is refused up front. The semaphore could never admit it, and a run blocked on a file it can never start looks exactly like a stalled remote. Uploads re-state each object to prove its size before the local copy is gone, closing a gap where a truncated upload was only noticed by a later verify. The destination is created before the chat is read. It is also the credentials check, and doing it first means a bad destination fails in seconds rather than after a full history walk. One invocation converges: each item is checked against the index immediately before download, so there are no passes and re-running is the resume path. Options that no longer exist say what replaced them instead of failing as unknown flags. Verified end to end against the live chat and a scratch remote path: two files downloaded, uploaded, confirmed present at the right size, staging left empty.
telegram-exporter
Export a Telegram chat's media to any rclone remote — S3, Google Drive, Dropbox, Backblaze B2, SFTP, WebDAV, or anything else rclone supports — using far less local disk than the chat's total size.
run.sh runs tdl and
rclone as a rolling pipeline: tdl downloads into a small
staging directory while rclone concurrently moves finished files to the remote
and deletes the local copies. Local disk only ever holds the files in flight
plus one sync interval of throughput, so a multi-terabyte chat exports fine on a
small disk. Telegram caps a single file at 2 GB (4 GB from premium uploaders),
so a few dozen GB of staging covers the worst case regardless of chat size.
This is needed because tdl can only write to a local directory — it has no
rclone integration and no remote destination of any kind (tdl dl -d takes a
filesystem path; tdl --storage is its session database, not an output target).
tdl downloads over MTProto with a user account, so Bot API limits do not apply:
full history is readable and there is no 20 MB download cap.
Requirements
- tdl — https://docs.iyear.me/tdl/getting-started/installation/
- rclone — https://rclone.org/install/
- bash. On Windows, run under WSL or Git Bash.
Setup
Both steps are one-time.
# 1) log in to Telegram with your user account (phone + code + 2FA)
tdl login
# 2) configure the destination. The interactive wizard covers every backend:
rclone config
# ...or create one non-interactively, e.g.
rclone config create gdrive drive
rclone config create b2 b2 account=KEY_ID key=APP_KEY
rclone config create dav webdav url=https://dav.example.com/remote.php/dav/files/you \
vendor=other user=YOU pass=SECRET
rclone listremotes # confirm the name you will pass to -r
Any rclone remote form works, including on-the-fly connection strings
(:webdav,url=https://...:/path).
Usage
./run.sh -r gdrive:telegram/media -c @mygroup
That is the whole flow. It exports the chat's message metadata to
export.json, then downloads and uploads concurrently until finished. Progress
and warnings go to stderr; press Ctrl-C at any point and it stops cleanly.
Options
| Flag | Meaning |
|---|---|
-r REMOTE:PATH |
Required. rclone destination, e.g. gdrive:telegram/media, s3:bucket/tg, dav:tg-export |
-c CHAT |
Chat to export when the JSON does not exist yet — id, username, or link (see below) |
-f FILE |
Export JSON to download from (default export-<chat>.json with -c, else export.json) |
-d DIR |
Staging directory (default ./staging) |
-i SECONDS |
Seconds between rclone sweeps (default 60) |
-a AGE |
rclone --min-age, a second guard against moving files still being written (default 45s) |
-m SIZE |
Cap the staging directory at SIZE (K/M/G/T, binary), e.g. 40G. Unset means no cap (see below) |
-h |
Help |
Identifying the chat
-c accepts every form tdl understands, plus one it doesn't:
| Form | Example |
|---|---|
Numeric id, as printed by tdl chat ls |
-c 1697797156 |
Username, with or without @ |
-c @mygroup / -c mygroup |
| Public link | -c https://t.me/mygroup / -c t.me/mygroup |
| Deep link | -c 'tg://resolve?domain=mygroup' |
| Bot API id (converted for you) | -c -1001697797156 → 1697797156 |
tdl resolves a numeric argument as an MTProto id and anything else through
gotd's resolver. MTProto has no -100 prefix, so a Bot API id would otherwise
fail to resolve; the script strips it and logs the conversion.
A message link is rejected — -c names a chat, not a message:
$ ./run.sh -r gdrive:tg -c https://t.me/mygroup/123
error: -c takes a chat, not a message link — pass the chat's username or id
Run tdl chat ls to see ids and usernames side by side.
Each chat gets its own export file by default (export-mygroup.json,
export-1697797156.json), so exporting a second chat from the same directory
never reuses the first one's JSON. When the file already exists it is reused and
the script says so — delete it to re-export.
Anything after -- is passed straight to tdl dl:
./run.sh -r gdrive:telegram/media -- -t 4 -l 1 # calmer parallelism, fewer flood waits
./run.sh -r gdrive:telegram/media -- -i mp4,mkv # only these file extensions
./run.sh -r gdrive:telegram/media -- -e jpg,png # skip these file extensions
tdl defaults to -t 8 -l 4, which is aggressive; lower it if you hit flood
waits on a large export.
Tuning the upload
rclone reads every one of its flags from an environment variable, so the upload side is tunable without touching the script:
RCLONE_TRANSFERS=8 RCLONE_BWLIMIT=20M ./run.sh -r s3:bucket/tg -c @mygroup
run.sh sets two of those itself, and only when the caller has not:
| Variable | Default here | rclone's own default | Why |
|---|---|---|---|
RCLONE_TRANSFERS |
2 |
4 |
Backends that commit an upload as a server-side async task queue those tasks; less parallelism keeps the queue short |
RCLONE_LOW_LEVEL_RETRIES |
20 |
10 |
Each retry re-polls a pending task, so a slow commit is waited out instead of failing the transfer |
Both exist because of one failure mode. On pikpak an upload finishes in two
phases — rclone sends the bytes, then a server-side task must reach
PHASE_TYPE_COMPLETE. rclone waits 500 ms and then polls, giving up after
--low-level-retries attempts with:
ERROR : <file>: Failed to copy: can't verify the task is completed: ... Phase:"PHASE_TYPE_PENDING"
Nothing is lost when that happens — the message is followed by Not deleting source as copy failed, the file stays in staging and the next sweep retries
it. But it wastes the upload, and it counts against a -m cap, since a file
that keeps failing can never be drained. Raise the retries further if you still
see it.
Capping the staging directory
Without -m, staging grows whenever tdl downloads faster than rclone uploads,
which on a fast connection and a slow remote can mean tens of GB between
sweeps. -m puts a ceiling on it:
./run.sh -r s3:bucket/tg -c @mygroup -m 40G
Staging size is checked every 10 seconds, independently of -i. When it
reaches the cap, tdl is suspended with SIGSTOP and rclone sweeps until
staging is back under it, then tdl is resumed — it reconnects on its own and
--continue picks its .tmp files back up. Because the checks are periodic,
the cap is a high-water mark rather than a hard limit: staging can overshoot by
up to ten seconds of download throughput before the gate closes.
Only finished files can be drained, so the cap has to exceed what the
concurrent downloads hold — at most -l times 2 GB (4 GB from premium
uploaders). With the default -l 2 anything from ~10 GB up is safe; below
that, the drain cannot clear the cap and the run logs a warning on every check
instead of throttling.
Exporting a subset
Generate the JSON yourself when you want a narrower export, then point -f at
it:
tdl chat export -c @mygroup -T id -i 1000,5000 --all --with-content -o part.json
./run.sh -r gdrive:telegram/media -f part.json
tdl chat export takes -T time|id|last with -i as the range, and -f as an
expression filter over message fields (-f - lists the available fields).
Sweep output
The periodic sweeps are silent — they run every -i seconds alongside tdl's own
output, and narrating each one would drown it. The sweeps that run once at the
end do report progress, since they can move the whole staging directory with
nothing else on screen:
- the exit sweep on Ctrl-C,
SIGTERM, or a tdl failure (sweeping completed files before exit); - the final sweep after tdl finishes successfully.
On a terminal that is rclone's redrawn --progress bar. When output is
redirected to a log it becomes a one-line stats summary every 30s
(--stats 30s --stats-one-line --stats-log-level NOTICE) — rclone logs stats at
INFO, so raising just the stats to NOTICE avoids the line-per-file spam that
-v would add.
To show progress on every sweep instead, rclone reads its flags from the environment:
RCLONE_PROGRESS=true ./run.sh -r s3:bucket/tg -c @mygroup
Resuming
Re-run the same command. Both legs resume independently and nothing is downloaded or uploaded twice.
Keep the same export.json between runs: --skip-same compares against the
staging directory, which is empty once files have moved to the remote, so
cross-run deduplication rests on tdl's own --continue tracking. If you must
start from a fresh export, narrow it to the missing message-id range
(-T id -i <last>,<max>) rather than re-downloading everything.
Verifying an export
run.sh finishes when tdl finishes, which is not the same as every file having
arrived: a dropped session, a stalled remote, or an interrupted pass all leave
gaps. verify-export.sh settles it by rebuilding the filename tdl produces for
each media message in the export JSON and checking the remote for it.
./verify-export.sh -f export-mygroup.json -r remote:telegram/media
messages in export : 18193
text-only (skip) : 38
media expected : 12000
present and intact : 12000
absent : 0
zero-byte : 0
COMPLETE: every media message is present and non-empty.
Messages with no media are skipped; they carry an empty file and were never
download targets. A zero-byte file counts as missing, because rclone overwrites a
size-mismatched destination and a retry repairs it. Files under 1 KiB are
reported but not retried, since some real media is genuinely that small. Exit
status is 0 when complete and 1 otherwise, with the outstanding message ids
written to missing-ids.txt.
Running until complete
export-until-complete.sh drives run.sh in a loop: verify what is already
there, narrow the export to the ids still missing, run the pipeline on that
subset, and repeat.
./export-until-complete.sh -r remote:telegram/media -c @mygroup
| Flag | Meaning |
|---|---|
-r REMOTE:PATH |
Required. rclone destination |
-c CHAT |
Chat to export metadata for on the first pass |
-f FILE |
Export JSON (default export-<chat>.json) |
-d DIR |
Staging directory (default ./staging) |
-i SECONDS |
rclone sweep interval (default 120) |
-m SIZE |
Staging cap passed through to run.sh, e.g. 40G |
-p N |
Maximum passes (default 30) |
-q GIB |
Stop if remote free space falls below this (default 5) |
It stops when the verifier reports complete (exit 0), when a pass fetches
nothing new (exit 1 — the remaining media is no longer available from
Telegram), when the remote runs low on space (exit 3), or on Ctrl-C (exit
130, after the current pass shuts down cleanly).
tdl's progress bar is shown when stdout is a terminal and suppressed when output is redirected, so a log file stays readable without a flag.
What it guards against
- Partial uploads. tdl writes
<name>.tmpand renames on completion, so every sweep excludes*.tmp. Age alone is not a completion signal: a download stalled by a flood wait stops touching its.tmp, which would then be uploaded half-written and lose its resume point. - Directories vanishing under tdl.
--delete-empty-src-dirsruns only in the final sweep, and the staging directory is recreated after every sweep. Removing a directory under a running tdl makes it fail to create its next file. - Orphaned downloads. tdl is stopped on exit, Ctrl-C, or
SIGTERM, so no download keeps running after the script is gone. - A failed run looking finished. The unrestricted final sweep happens only after tdl exits 0. An interrupted or crashed run gets the age-guarded sweep and keeps staging for the next attempt.
- A dead remote filling the disk. Five consecutive rclone failures abort the run instead of letting staging grow unbounded.
- A fast connection filling the disk. With
-m, tdl is suspended whenever staging reaches the cap and resumed once rclone has drained it, so download throughput cannot outrun the upload leg. - Typos and bad credentials. Before downloading anything, the remote must be
present in
rclone listremotes(skipped for connection strings) and the destination must be creatable, which proves both reachability and auth.
Exit codes: 0 success, 2 usage error, 3 rclone failure, 130/143
interrupted, anything else is tdl's own exit code.
Limits
- Streaming with no staging at all (piping download chunks straight to the remote) is not possible with tdl and would require custom code.
- The script is bash; the two tools it drives are cross-platform, but Windows needs WSL or Git Bash.