15 Commits
Author SHA1 Message Date
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00
Alex 00be2c05ad fix: make worker-delegated embedding survive the shipped deployments
Query embedding moved to the Celery worker, but nothing that ships was
updated to consume the queue it dispatches to.

- Add `embeddings` to every worker `-Q` list (compose x3, k8s, devcontainer,
  sandbox README). Without it a search blocked for EMBEDDINGS_DELEGATE_TIMEOUT
  and then answered with no retrieved context, because classic_rag swallows the
  dispatch error and skips the source -- bad answers, not an error.

- Skip the task_postrun heap reclaim for the embed task. The full gc.collect()
  was written for docling/torch parses; on a worker holding the ONNX model it
  measured ~86ms against ~8ms for the embed itself, a 9x slowdown of the round
  trip for a task that allocates a few kilobytes.

- Resolve the installation pin in the re-embed script. It never imports
  application.app, so an install pinned in app_metadata with no EMBEDDINGS_NAME
  set -- every stock k8s deployment, whose manifests carry no embedding config
  -- would rewrite its whole index with the legacy default and stamp
  sources.model to match, then be told by the boot warning to run it again.

- Fail fast for 30s after a failed dispatch. fanout.embed_questions falls back
  to letting each store embed its own query, so one dead-worker retrieval paid
  the timeout once in the fan-out and again per source.

- Forget the task result. Nothing reads it back: the key is per-dispatch UUID,
  not content-addressed, so a repeated query mints another. Left alone every
  search leaked ~17KB for result_expires (7 days) into the Redis the broker
  shares -- on the bundled k8s manifest (1Gi, no maxmemory policy) that is an
  OOMKill that takes the broker with it.

- Release the model ensure_vector_schema loads to read the width of an
  unregistered model, in a process that delegates and would never call it.
  The width still comes from the model, not the table, so the mismatch check
  the hook exists for keeps working.

- Correct the docs that said otherwise: embeddings.md claimed the standard
  deployment worked unchanged, upgrading.mdx said no action was needed, and
  the settings table listed none of the three delegation settings.
2026-08-28 14:31:19 +01:00
Alex 47e53a71c4 feat: embed on the worker, and stop batching the ONNX pass
The API embeds every query it serves, so it held its own copy of the model:
~890 MB it never needed. EMBEDDINGS_DELEGATE_TO_WORKER (on by default) sends
the text to the Celery worker instead and gets the vector back, taking an API
process from 1176 MB to 285 MB with no ONNX Runtime imported at all. The client
embeds locally when it finds itself inside a worker task, so the worker never
dispatches to itself -- the same self-deadlock DOCUMENT_PARSE_QUEUE avoids on
the parsing side. EMBEDDINGS_BASE_URL still wins over it, and remains the right
answer for production.

ensure_vector_schema was constructing the embeddings instance purely to read
.dimension off it, loading several hundred MB of ONNX into every API and worker
process at import. For a model the registry describes that is a lookup; only an
unregistered name now falls back to loading.

EMBEDDINGS_BATCH_SIZE was sizing two unrelated things: chunks per store
transaction (and per remote embed request) and documents per ONNX forward pass.
Each pass pads every input up to its longest, and that waste grows with the
square of chunk length, so at the 1250-token default a batch of 32 peaked at
6.6 GB and took 326s where a batch of 1 peaked at 2.9 GB and took 90s. The
forward pass is now sized by EMBEDDINGS_MODEL_BATCH_SIZE, defaulting to 1;
storage and remote batching are unchanged at 32.

reembed embeds in-process: a batch job that walks the whole index should not
round-trip every chunk through a broker, and loading the model there reports a
real failure instead of timing out against an empty queue.

Also drops the mpnet zip download from the docs and the devcontainer, which
pointed at a SentenceTransformers export with no ONNX graph and had been inert
since the FastEmbed swap; corrects the claim that any sentence-transformers
model works; and settles the Configuring/Settings pages on what the registry
and the repository metadata actually decide.
2026-08-28 12:23:18 +01:00
Alex 37d93cbd86 Parse documents on a Celery parsing worker via a read_document tool
Replace the sandbox Docling extractor with read_document, backed by the in-process
backend parser (the same one ingestion uses) and offloaded to a dedicated
'parsing' Celery queue so it can run on GPU-capable workers with predictable RAM.
The tool resolves the input ref under the run-scoped gate, enqueues the parse,
and awaits it with a timeout (degrading to an error rather than hanging); the
worker independently re-resolves the artifact through the same gate and never
trusts a raw path. Untrusted files get the upload path's safeguards (extension
whitelist, size cap, sanitized temp file, cleanup). Options: output
(markdown/text/structured/chunks), ocr, pages, engine, max_chars, include_tables,
persist, json_schema. The workflow native-file 'extract' fallback now uses the
same worker path, so document parsing no longer needs the sandbox and works on
every backend.

Also fixes the branch's periodic-task test (the sandbox reaper made it 12) and
points the dev and e2e Celery workers at the parsing queue.
2026-06-25 13:24:12 +01:00
Alex f8e42cdce1 chore: update docs 2026-06-08 22:12:57 +01:00
Alex b024936ad7 Update devc-welcome.md 2025-02-12 09:48:21 +00:00
Alex 5f42e4ac3f fix: default file codespace 2025-02-11 09:53:26 +00:00
Alex 926ec89f48 Create devc-welcome.md 2025-02-11 09:48:45 +00:00
Alex fbad183d39 fix: post create devcontainers 2025-02-07 18:44:24 +00:00
Alex 7356a2ff07 fix: minor docker fixes 2025-02-07 18:39:07 +00:00
Alex 6ff948c107 fix: dockerfile in devcontainer build dir fix 2025-02-07 14:35:44 +00:00
Alex e3ebce117b fix: devcontainer paths 2 2025-02-07 14:34:19 +00:00
Alex ce69b09730 fix: devcontainer paths 2025-02-07 14:30:15 +00:00
Alex c823cef405 fix: devcontainer codespaces correct api address 2025-02-07 14:25:09 +00:00
Alex d754a43fba feat: devcontainer 2025-02-05 11:54:06 +00:00