Files
telegram-exporter/internal/tgsource/takeout.go
T
tiennm99 19837ceb85 fix: reject names rclone rewrites, and stop resolving link markers as chats
rclone does not address a file by the bytes os.OpenFile wrote. Names handed to
an Fs go through the backend encoder and names listed back are re-encoded to
the standard set, neither of which the write path performs. A filename holding
one of the rewritten characters was therefore stored under one string and
looked up under another: confirmed against the local backend, where a written
"a‛b.jpg" reports object not found and a written "a\nb.jpg" lists back as
"a␊b.jpg". That is the same two-derivation divergence this program was written
to remove, with rclone's encoder standing where filenamify used to. Such names
are rejected, not encoded, for the same reason every other name is.

The length limit ignored the ".part" suffix that is opened first, so a name
just inside NAME_MAX passed the check and then failed to open on every pass,
stalling the walk on that message forever. The suffix now lives beside the
limit that has to account for it.

filter.NewFilter(nil) does not build a neutral filter; it copies the package
global, which rclone has already filled from RCLONE_*. Indexing inherited the
operator's environment, so a stray RCLONE_MIN_SIZE emptied the index and
re-downloaded the archive. Every narrowing field is now set explicitly and the
result is asserted inactive. A subprocess test covers it, since the env is read
at package init and t.Setenv is too late to observe anything.

An ErrorDirNotFound from a subdirectory was also treated as an empty
destination, returning a partial index as authoritative.

t.me/c/<id> and t.me/s/<name> were passed through whole, and gotd reads the
first path component as the username — resolving "c" or "s", which is a
confusing failure at best and someone else's chat at worst, since
one-character usernames exist. Both now yield the chat, and t.me/s/<name>/<id>
is refused like any other message link. Two tests asserted the old behaviour.

core's dcpool.Takeout deadlocks when takeout init fails: it holds the pool
mutex and recovers by calling Client, which takes the same non-reentrant
mutex. Telegram returns TAKEOUT_INIT_DELAY for a takeout started recently and
takeout is on by default, so two runs in succession hang the process with no
output and no response to cancellation. The session is established once here
instead, falling back to a plain client, and the pool's own Takeout is never
called.
2026-09-06 20:28:09 +07:00

63 lines
2.1 KiB
Go

package tgsource
import (
"context"
"sync"
"github.com/gotd/td/tg"
"github.com/iyear/tdl/core/dcpool"
"github.com/iyear/tdl/core/middlewares/takeout"
)
// safeTakeout wraps a pool to make its takeout path survive a failed init.
//
// core's own dcpool.Takeout deadlocks on that path. It holds the pool's mutex
// for the whole call, and its recovery from a failed init is to return
// p.Client(ctx, dc) — which locks the same mutex again (dcpool.go:113-121 and
// :57-58). sync.Mutex is not reentrant, so the worker blocks forever, then every
// other worker blocks behind it, and the process hangs with no output and no
// response to cancellation, since the goroutine is parked on a mutex rather than
// a select.
//
// That is not an exotic path. Telegram answers account.initTakeoutSession with
// TAKEOUT_INIT_DELAY when a takeout was started recently — tdl's own "ignore
// init delay error" comment shows it expects exactly this — and takeout is on by
// default, so running two exports in succession is enough to trigger it.
//
// Probing before the run is not an alternative: a probe would consume an init
// and make the pool's own init the one that gets the delay error. So the takeout
// session is established here instead, once, and the pool's Takeout is never
// called at all.
type safeTakeout struct {
dcpool.Pool
once sync.Once
id int64
ok bool
}
func withSafeTakeout(p dcpool.Pool) dcpool.Pool { return &safeTakeout{Pool: p} }
// Takeout returns a takeout-scoped client, or an ordinary one if no takeout
// session could be established.
//
// Falling back rather than failing matches what core intended: takeout raises
// rate limits and reaches older history, but a download works without it. The
// difference is that this fallback returns.
func (s *safeTakeout) Takeout(ctx context.Context, dc int) *tg.Client {
base := s.Pool.Client(ctx, dc)
s.once.Do(func() {
id, err := takeout.Takeout(ctx, base.Invoker())
if err != nil {
return // ok stays false; every caller gets a plain client
}
s.id, s.ok = id, true
})
if !s.ok {
return base
}
return tg.NewClient(takeout.Middleware(s.id).Handle(base.Invoker()))
}