tiennm99 aea29826c2 fix: detect truncated uploads, and stop reporting unreachable work as retryable
Verification matched on name and non-zero size, so an upload that died partway
was counted archived permanently. Comparing against the size Telegram reports
found six such objects in the live archive, one of them 221 MiB standing in for
a 2006 MiB video. They are folded into the outstanding set; rclone overwrites a
size mismatch, so another pass repairs them.

A zero Report claimed the archive was complete — nothing expected, nothing
missing — which is the value both commands hold before their Telegram callback
populates it, so any early return printed COMPLETE and exited 0 on an untouched
chat. A Report now knows whether it ran.

A name that can never be written kept the run outstanding forever while the
download set deliberately excluded it, so a driver looping on "incomplete"
walked the whole history and re-indexed the whole remote every pass for work
that could not be done. Such a run now reports STALLED and exits 4.

A destination that stopped accepting uploads was reported and then discarded,
exiting 1. The same driver would retry against a full or unreachable remote
indefinitely, downloading gigabytes each pass to upload none. It exits 3.

Free space is re-checked during the run, not only before it. An archive this
size runs for hours, and the remote can fill in the middle; discovering it
through five failed multi-gigabyte uploads wastes the download for all of them.

Also: parseSize silently wrapped to a negative or zero on a large input, which
reads downstream as "no cap"; list printed attacker-chosen filenames raw, so a
tab shifted the columns and an escape sequence reached the terminal; a missing
backend blamed credentials rather than the build; humanBytes indexed past its
unit table above 1 PiB; a second Init reported success against a config that
never loaded; and sync did not surface the basename collisions verify warned
about, though sync is the command that acts on the verdict.

The env-override test could not observe what it claimed: rclone reads RCLONE_*
at package init, so t.Setenv came too late and the assertion held with the
guard removed. It runs in a subprocess now, as does the new one covering the
index against inherited filters.
2026-09-06 20:44:08 +07:00
2026-08-11 14:07:26 +07:00

telegram-exporter

Archive a Telegram chat's media to any rclone remote — S3, Google Drive, Dropbox, Backblaze B2, SFTP, WebDAV, pikpak, or anything else rclone supports — using far less local disk than the chat's total size.

tgexport embeds tdl and rclone as libraries and runs both halves in one process. Files are downloaded into a small staging directory and uploaded the moment each one finishes, so local disk only ever holds what is in flight. A multi-terabyte chat archives fine on a small disk. Telegram caps a single file at 2 GB (4 GB from premium uploaders), so a few dozen GB of staging covers the worst case regardless of chat size.

This exists because tdl can only write to a local directory — it has no remote destination of any kind (tdl dl -d takes a filesystem path; tdl --storage is its session database, not an output target). tdl downloads over MTProto with a user account, so Bot API limits do not apply: full history is readable and there is no 20 MB download cap.

Requirements

Neither tool is invoked at run time; tgexport reads the session and config they write.

Setup

Both steps are one-time.

tdl login                 # writes the Telegram session tgexport reads
rclone config             # define the destination remote
go build -o tgexport ./cmd/tgexport
./tgexport doctor -r myremote:archive

doctor proves both halves work before a long run: it prints the logged-in account, resolves the destination, and reports free space.

Build variants

Build Backends Size
go build ./cmd/tgexport every rclone backend ~92 MB
go build -tags slim ./cmd/tgexport pikpak only ~49 MB

A backend that is not compiled in does not exist at run time, so use the default build unless the destination will never change.

Usage

./tgexport sync -c CHAT -r REMOTE:PATH [options]

CHAT accepts a numeric id as printed by tdl chat ls, a username with or without @, or a t.me/tg:// link. A Bot API -100… id is converted automatically. A link to a single message is refused — it names a message, not a chat.

# archive a chat, capping staging at 40 GiB
./tgexport sync -c @mychannel -r gdrive:telegram/media -m 40G

# check completeness without downloading anything
./tgexport verify -c @mychannel -r gdrive:telegram/media

# list what the chat holds
./tgexport list -c @mychannel

Options

Flag Default Meaning
-c — chat id, username, or link (required)
-r — rclone destination, REMOTE:PATH (required)
-d ./staging staging directory for files in flight
-m no cap cap staging at a size, e.g. 40G
--threads 4 connections per file
--limit 2 files downloading at once
--uploads 2 files uploading at once
--min-free 5 stop if the remote has fewer than this many GiB free
--limit-items 0 stop after N files; for smoke tests
--confirm true re-state each uploaded file to prove its size
--takeout true use a takeout session
-n default tdl session namespace

Exit codes

Code Meaning
0 complete
1 ran, but files remain
2 usage error
3 remote or Telegram failure
130 / 143 interrupted (SIGINT / SIGTERM)

How it works

Re-running is the resume path. Each item is checked against a listing of the remote immediately before download, so an interrupted run picks up where it left off and a completed one downloads nothing.

Filenames. Every file is stored as {DialogID}_{MessageID}_{FileName}, where FileName is exactly what Telegram reports. One function derives that string, and the same string is used both to ask whether the file is already archived and to write it — so the two can never disagree.

That last point is the reason this program exists. Its predecessor derived the name twice: tdl chat export wrote the raw name into a JSON, while tdl dl rendered it through a template applying filenamify, which rewrites characters a filesystem rejects and collapses runs of !. A file whose name contained !! was looked up under one name and stored under another, so the verifier never found it and re-fetched it on every pass — forever, at 966 MB a time.

Note the consequence: names are not run through filenamify, so they are not byte-compatible with what the old shell pipeline wrote. A file it stored under a rewritten name will not be recognised and gets fetched again.

Disk. -m is a byte budget. A download reserves its own size before starting and releases it only once the upload is confirmed, so when the remote is slow the downloads pause on their own. The cap must exceed the largest single file, and a cap that does not is refused at startup rather than discovered as a hang.

Integrity. A download is written to <name>.part and renamed only once its size matches what Telegram reported, so a file without the suffix is always whole. Uploads are re-stated afterwards to prove they arrived at the right size, before the local copy is gone.

Replacing the shell pipeline

Earlier versions of this repo were three bash scripts — run.sh, export-until-complete.sh and verify-export.sh — driving tdl and rclone as separate processes. Everything expensive in them existed to work around the fact that neither process could see the other's state: a staging directory polled with du -sk, an --min-age guard, a *.tmp exclusion, SIGSTOP/SIGCONT to enforce the disk cap, a sweep-failure counter, and an outer loop that re-verified and re-narrowed a JSON export between passes.

One process needs none of it. Completion is a function returning; the cap is a semaphore. Some hard-won details were worth keeping, and are:

  • pikpak commits uploads as a server-side async task, and rclone abandons a still-pending one when its low-level retries run out. transfers=2 and low-level-retries=20 are the defaults here for that reason. Environment overrides still win.
  • A backend with no quota API is treated as unlimited, so it never blocks a run.
  • Zero-byte files count as missing — rclone overwrites a size-mismatched destination, so re-running repairs them — while files under 1 KiB are reported but trusted, since some real media genuinely is that small.

Flags that disappeared are recognised and explain what replaced them:

Old Why it is gone
-i no sweep interval; uploads start when a download finishes
-a no --min-age; completion is observed, not inferred
-f no export JSON; the chat is read live, so names cannot go stale
-p no passes; one invocation converges
-q renamed --min-free

Notes

  • tgexport and the tdl CLI share one session store and cannot run against the same namespace at once. Use -n for a second namespace if you need both.
  • A partially downloaded file is not resumable across restarts — tdl's library exposes no resume offset — so an interrupted run re-fetches whatever was in flight, bounded by --limit.
  • Everything is read-only against Telegram. Nothing is uploaded, deleted, or marked read.
S
Description
No description provided
Readme Apache-2.0
624 KiB
0 Stars 1 Watchers 0 Forks
Languages
Go 100%