mirror of
https://github.com/tiennm99/telegram-exporter.git
synced 2026-10-11 03:13:49 +00:00
rclone does not address a file by the bytes os.OpenFile wrote. Names handed to an Fs go through the backend encoder and names listed back are re-encoded to the standard set, neither of which the write path performs. A filename holding one of the rewritten characters was therefore stored under one string and looked up under another: confirmed against the local backend, where a written "a‛b.jpg" reports object not found and a written "a\nb.jpg" lists back as "a␊b.jpg". That is the same two-derivation divergence this program was written to remove, with rclone's encoder standing where filenamify used to. Such names are rejected, not encoded, for the same reason every other name is. The length limit ignored the ".part" suffix that is opened first, so a name just inside NAME_MAX passed the check and then failed to open on every pass, stalling the walk on that message forever. The suffix now lives beside the limit that has to account for it. filter.NewFilter(nil) does not build a neutral filter; it copies the package global, which rclone has already filled from RCLONE_*. Indexing inherited the operator's environment, so a stray RCLONE_MIN_SIZE emptied the index and re-downloaded the archive. Every narrowing field is now set explicitly and the result is asserted inactive. A subprocess test covers it, since the env is read at package init and t.Setenv is too late to observe anything. An ErrorDirNotFound from a subdirectory was also treated as an empty destination, returning a partial index as authoritative. t.me/c/<id> and t.me/s/<name> were passed through whole, and gotd reads the first path component as the username — resolving "c" or "s", which is a confusing failure at best and someone else's chat at worst, since one-character usernames exist. Both now yield the chat, and t.me/s/<name>/<id> is refused like any other message link. Two tests asserted the old behaviour. core's dcpool.Takeout deadlocks when takeout init fails: it holds the pool mutex and recovers by calling Client, which takes the same non-reentrant mutex. Telegram returns TAKEOUT_INIT_DELAY for a takeout started recently and takeout is on by default, so two runs in succession hang the process with no output and no response to cancellation. The session is established once here instead, falling back to a plain client, and the pool's own Takeout is never called.
63 lines
2.1 KiB
Go
63 lines
2.1 KiB
Go
package tgsource
|
|
|
|
import (
|
|
"context"
|
|
"sync"
|
|
|
|
"github.com/gotd/td/tg"
|
|
|
|
"github.com/iyear/tdl/core/dcpool"
|
|
"github.com/iyear/tdl/core/middlewares/takeout"
|
|
)
|
|
|
|
// safeTakeout wraps a pool to make its takeout path survive a failed init.
|
|
//
|
|
// core's own dcpool.Takeout deadlocks on that path. It holds the pool's mutex
|
|
// for the whole call, and its recovery from a failed init is to return
|
|
// p.Client(ctx, dc) — which locks the same mutex again (dcpool.go:113-121 and
|
|
// :57-58). sync.Mutex is not reentrant, so the worker blocks forever, then every
|
|
// other worker blocks behind it, and the process hangs with no output and no
|
|
// response to cancellation, since the goroutine is parked on a mutex rather than
|
|
// a select.
|
|
//
|
|
// That is not an exotic path. Telegram answers account.initTakeoutSession with
|
|
// TAKEOUT_INIT_DELAY when a takeout was started recently — tdl's own "ignore
|
|
// init delay error" comment shows it expects exactly this — and takeout is on by
|
|
// default, so running two exports in succession is enough to trigger it.
|
|
//
|
|
// Probing before the run is not an alternative: a probe would consume an init
|
|
// and make the pool's own init the one that gets the delay error. So the takeout
|
|
// session is established here instead, once, and the pool's Takeout is never
|
|
// called at all.
|
|
type safeTakeout struct {
|
|
dcpool.Pool
|
|
|
|
once sync.Once
|
|
id int64
|
|
ok bool
|
|
}
|
|
|
|
func withSafeTakeout(p dcpool.Pool) dcpool.Pool { return &safeTakeout{Pool: p} }
|
|
|
|
// Takeout returns a takeout-scoped client, or an ordinary one if no takeout
|
|
// session could be established.
|
|
//
|
|
// Falling back rather than failing matches what core intended: takeout raises
|
|
// rate limits and reaches older history, but a download works without it. The
|
|
// difference is that this fallback returns.
|
|
func (s *safeTakeout) Takeout(ctx context.Context, dc int) *tg.Client {
|
|
base := s.Pool.Client(ctx, dc)
|
|
|
|
s.once.Do(func() {
|
|
id, err := takeout.Takeout(ctx, base.Invoker())
|
|
if err != nil {
|
|
return // ok stays false; every caller gets a plain client
|
|
}
|
|
s.id, s.ok = id, true
|
|
})
|
|
if !s.ok {
|
|
return base
|
|
}
|
|
return tg.NewClient(takeout.Middleware(s.id).Handle(base.Invoker()))
|
|
}
|