The lifecycle tests imported docsgpt.celery_init as a module next to the
file's from-imports, which code scanning flags. Patch the flag by path and send
the signals from celery.signals -- the objects a worker actually fires.
in_worker() read task_join_will_block, which eventlet and gevent leave unset,
and the task's current_worker_task, which they scope to one greenlet -- so a
greenlet a task spawned in those pools still took the web-process branch and
dispatched to its own worker, where it could queue behind its parent and time
out.
The worker's own startup now records it: worker_init fires in every worker's
main process before the pool starts (where solo, threads, eventlet and gevent
run tasks, and what prefork children fork from), and worker_process_init in
each prefork child. worker_ready would be too late -- prefork children are
forked before it fires. The two existing checks stay for anything that runs
tasks without that startup.
Celery records the executing task on the thread that runs it, so a thread that
task starts sees none. The embeddings client and read_document both decided
"am I in a worker?" from that alone, and from any other thread took the
web-process branch: dispatch to the worker they were running in and block on
the result. Celery refuses that get() ("Never call result.get() within a
task!"), so the embed failed and latched the 30s dispatch cooldown for every
caller after it; with joins allowed, read_document would instead wait on a
parsing queue only its own busy process serves.
Threads inside tasks are not hypothetical: per-source retrieval fans out to a
pool, so a scheduled or webhook agent searching several sources embedded from
pool threads. Graph extraction did too, which failed every chunk of a build.
in_worker() in celery_init answers for the whole process. Celery's
task_join_will_block is process-wide and set for every blocking pool (prefork,
solo, threads) -- exactly the condition under which dispatch-and-wait goes
wrong; eventlet/gevent leave it unset, so the task's own thread still counts
through current_worker_task. Verified with real workers on each blocking pool:
from a thread a task started, the old check dispatched and hit the error, the
new one embedded locally.
- The post-task reclaim skip recognises the legacy application.* embed name,
so query embeds queued by the previous release do not pay a full collect.
- The Azure compose file mounts host data on /app/{indexes,inputs,vectors},
where the process actually reads and writes; it mounted /app/application/...
before the rename and /app/docsgpt/... after it, and nothing wrote to either.
- The offline image check triggers on docsgpt/requirements*.txt again;
dependabot's pip entry points at docsgpt/.
- install_hint() and its docstring name docsgpt/requirements-<extra>.txt;
the test asserts the full path.
- application/vectors/ stays ignored: the compose files still mount it.
- The durability QA script quiets the docsgpt logger tree.
- Upgrade note: the three renamed source-sync entries start their timers
from the upgrade.
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
Query embedding moved to the Celery worker, but nothing that ships was
updated to consume the queue it dispatches to.
- Add `embeddings` to every worker `-Q` list (compose x3, k8s, devcontainer,
sandbox README). Without it a search blocked for EMBEDDINGS_DELEGATE_TIMEOUT
and then answered with no retrieved context, because classic_rag swallows the
dispatch error and skips the source -- bad answers, not an error.
- Skip the task_postrun heap reclaim for the embed task. The full gc.collect()
was written for docling/torch parses; on a worker holding the ONNX model it
measured ~86ms against ~8ms for the embed itself, a 9x slowdown of the round
trip for a task that allocates a few kilobytes.
- Resolve the installation pin in the re-embed script. It never imports
application.app, so an install pinned in app_metadata with no EMBEDDINGS_NAME
set -- every stock k8s deployment, whose manifests carry no embedding config
-- would rewrite its whole index with the legacy default and stamp
sources.model to match, then be told by the boot warning to run it again.
- Fail fast for 30s after a failed dispatch. fanout.embed_questions falls back
to letting each store embed its own query, so one dead-worker retrieval paid
the timeout once in the fan-out and again per source.
- Forget the task result. Nothing reads it back: the key is per-dispatch UUID,
not content-addressed, so a repeated query mints another. Left alone every
search leaked ~17KB for result_expires (7 days) into the Redis the broker
shares -- on the bundled k8s manifest (1Gi, no maxmemory policy) that is an
OOMKill that takes the broker with it.
- Release the model ensure_vector_schema loads to read the width of an
unregistered model, in a process that delegates and would never call it.
The width still comes from the model, not the table, so the mismatch check
the hook exists for keeps working.
- Correct the docs that said otherwise: embeddings.md claimed the standard
deployment worked unchanged, upgrading.mdx said no action was needed, and
the settings table listed none of the three delegation settings.