Commit Graph
15 Commits
Author SHA1 Message Date
Alex 2b6d4d509e fix(graphrag): return a chunk once from entity_pages
The subject flag was in the GROUP BY, so a chunk two matching entities link --
one the page is about, one merely mentioned in it -- came back as two
identical pages and spent the caller's page budget twice on the same text.
It is aggregated with bool_or now, which is what the ordering wanted anyway.

Covered by a live test against a pgvector-shaped documents table: the graph
tables alone cannot answer this query, so nothing exercised it before.
2026-09-20 10:45:04 +01:00
Alex ecf02d0d60 fix(graphrag): take a source's write lock per chunk, and keep zero weights
Two builds of one source can overlap: the extraction lease is keyed by the
source's updated_at, and enabling a graph updates the source before it
dispatches, so a rebuild started while the last build runs gets a new key and
a lease of its own. Both builds could then pass a chunk's "done" check before
either committed and apply it twice -- doc_freq bumped twice, reproduced with
two live writers. A reset could also land in the middle of a chunk.

apply_chunk and delete_by_source now take a transaction-scoped advisory lock
keyed by the source before touching a row, as the schema bootstrap already
does for DDL. A single build's writes were already serial, so it loses
nothing; overlapping builds take turns chunk by chunk, and the second sees the
first's "done" row and returns (0, 0).

apply_chunk also still defaulted with `rel.get("weight") or 1.0`, turning an
explicit zero into a full-strength edge -- the conversion 3f774d81 removed from
add_edge and the ranker but missed here. Only a missing weight defaults now.
2026-09-19 20:27:27 +01:00
Alex b3d7e63ae3 fix(graphrag): compose graph SQL through psycopg instead of f-strings
Four graph-store queries formatted table and column names into the SQL
string: get_chunk_texts and delete_by_source (Bandit B608, alerts #582/#583
on main) and entity_pages/chunk_similarities from this branch (#661/#662,
dismissed). The names were validated by _safe_identifier, so none was
injectable, but each query was still a string built at runtime.

They are now fixed statements composed with psycopg.sql: identifiers go in as
sql.Identifier, values stay bound. Identifiers are lower-cased before quoting,
because PGVectorStore writes the same names unquoted and Postgres folds those
to lower case -- a quoted mixed-case name would address a different table.
Bandit reports nothing for docsgpt/graphrag now; the queries return the same
rows against a real graph as before.
2026-09-19 20:10:54 +01:00
Alex 5e06470f9e test(graphrag): cover the graph reads, graph tool and vector blend without pgvector
CI has no pgvector, so every live graph test skips there and the queries this
branch added ran in no CI job at all. These pin what holds without a
database: the four new read queries bind every value (the entity name comes
from an LLM tool call) and map their rows; empty input runs no query; a failed
query returns nothing and releases its connection. The graph tool's plumbing,
the sources_have_graph gate and the hybrid path's vector ranking get the same.

Also drops a redundant chained comparison flagged by code scanning.
2026-09-19 14:33:14 +01:00
Alex a83e1dc0af feat(graphrag): seed the walk from what entities are, and rank with passages and vector hits
Graph retrieval tied plain vector search at best and never beat it. Measured
across five corpora, the bottleneck was seeding, not the graph: the walk
started from nodes whose embeddings were computed from bare entity names, and
a whole question shares almost nothing with a name like "Quill".

Extraction now embeds each node from "name (type): description" and each
relationship as the fact it asserts ("Alder streams_to Quill: ..."), stored on
a new nullable graph_edges.fact_embedding column that ensure_vector_schema adds
in place. Entity names are canonicalised (case, punctuation, word breaks and a
cautious plural) so "VECTOR_STORE" and "vector stores" land on one node. Extraction calls run
concurrently (GRAPHRAG_EXTRACTION_WORKERS, default 8) while embedding and graph
writes stay serial on the task thread, so ordering and idempotency are
unchanged; that measured 8.4x faster with identical output.

Retrieval gains per-source options, stored under retrieval.graph and read live
at query time:

- seed_strategy: start from matching entities (default) or matching
  relationships, which can reach an entity the question never names;
- passage_nodes (on): walk the source's passages alongside entities, with
  PageRank damping 0.5 instead of 0.85;
- blend_vector (on): fuse the graph ranking with the source's vector ranking
  by reciprocal rank.

The defaults are the measured-best configuration. Through GraphRAGRetriever,
the new seeding moved recall@4 from 0.41 to 0.68 on a multi-hop corpus and
from 0.50 to 1.00 on the docs corpus, and regressed none of the corpora
measured. Existing graphs keep name-only embeddings until rebuilt.
2026-09-19 14:07:41 +01:00
Alex 3f774d813c fix(graphrag): review fixes — replay-safe writes, strict count, zero weights
Three findings from review, all on code this branch introduced.

apply_chunk was not replay-safe. commit() can report connection loss *after*
Postgres committed, and the reconnect retry then replays the write: _upsert_node
bumps doc_freq a second time and _add_edge inserts another row, since
graph_edges has no uniqueness constraint for a logical edge. The chunk's
graph_ingest_progress row is now written in the same transaction as the rows it
describes, and a replay that finds it already "done" returns (0, 0) without
touching the graph. Extraction drops its separate mark_chunk("done"): the
checkpoint and the graph can no longer disagree.

count_nodes swallows every query failure and answers 0, so extraction's
"fall back to the write count" handler could never run — a failed count after a
successful build reported an empty graph. count_nodes grows a strict mode that
re-raises; retrieval keeps the swallow, which is what routes a source to
ClassicRAG.

A zero edge weight was read as a full-strength link: `or 1.0` rewrote an
explicit 0 before the <= 0 filter. Only missing and null weights default now.
The same coercion sat in _ppr_scores, where it would have kept the ranker's
rule unreachable from the product path, so it is fixed there too.
2026-09-17 16:31:19 +01:00
Alex e3d819d9fd fix(graphrag): stop losing chunks silently during a graph build
Three ways a chunk disappeared from a graph with no way to tell:

A build checks out one connection and then spends minutes per chunk waiting
on the model, so the connection idles long enough for the server or a pooler
to drop it. The pool only validates a connection when it hands one out, and
this one was handed out at the start of the build, so the next write raised
"the connection is lost", the chunk was marked failed, and the build carried
on a chunk short. apply_chunk and mark_chunk now reconnect and retry once;
every statement they run is an idempotent upsert, so a replay cannot
double-write. Only connection loss retries — a bad statement still surfaces.

An unparseable model response marked the chunk failed and logged nothing at
all, so failed_chunks was the only evidence and it named no chunk. Both
failure modes now log the chunk id.

The summary's node count summed per-chunk upserts, so an entity appearing in
ten chunks counted ten times: it reported writes, not graph size. It now
reports the distinct node count, falling back to the write count only if the
count query fails.
2026-09-17 15:44:23 +01:00
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00
Alex 25e07f5cee fix: embeddings registry edge cases and chunk-size accounting
Follow-up to the embeddings work, from a review pass over the branch.

- Route the OpenAI/Azure key handling through the model registry instead of
  matching the canonical name literally, so the `text-embedding-ada-002`
  alias the registry now accepts also reaches the Azure deployment name
  rather than failing every embed with DeploymentNotFound.
- Fall back to a default width where the embeddings model reports no
  dimension. A model outside the registry returns None rather than no
  attribute, so `getattr` with a default did not catch it and the width
  reached the DDL as `vector(None)` / `list_size=None`.
- Point HF_HUB_CACHE at the prefetch directory. Chunking loads the tokenizer
  through `tokenizers`, which reads the hub cache, so a fresh container
  fetched over the network on first ingest and an offline one silently fell
  back to cl100k.
- Charge a token that collapses a long unbroken run by its character span.
  WordPiece emits one [UNK] for any word over its character limit, which made
  base64 and minified content count as near-zero tokens, so nothing split it
  and oversized chunks reached the embedding server.
- Preserve chunk ids and honour --batch-size when rebuilding a FAISS index.
  Fresh uuids orphaned GraphRAG's graph_node_chunks rows, and the whole index
  went out in a single embed call on remote servers.
- Document that granite runs an int8-quantised graph, and scope the
  SentenceTransformer parity claim to mpnet's fp32 graph, which is where it
  was measured.
- Correct the embeddings docs: a matching dimension is not a matching model,
  so a same-width swap raises nothing and silently degrades retrieval.
2026-08-27 13:59:51 +01:00
Alex a72434c6db fix: small pg related fixes for stability 2026-08-20 12:51:31 +01:00
Alex 0b36257202 feat: parser improvements and fixes 2026-08-19 23:49:26 +01:00
Alex d7bbfcfe17 fix: minor graph rag improvements 2026-06-23 20:10:41 +01:00
Alex 7e7f34ee9c test(graphrag): address code-quality nits in graph tests
Meaningful self-loop degree assertion (was always-true); single import style per module; explicit returns in the live-store helper. (CI code-quality bot.)
2026-06-23 03:48:21 +01:00
Alex 2594bb0bed feat(graphrag): graph-view endpoints + network visualization
GraphStore.get_graph_overview (top-N by degree, bounded) + get_node_detail (description + linked chunks). GET /api/sources/<id>/graph + /graph/node/<id> (read-access gated, node scoped to source_id; empty graph -> {nodes:[],edges:[]}). Frontend GraphView (react-force-graph-2d) wired to config.kind=graphrag via card-click/'View graph'; node-click -> description + chunks; tooltip renders untrusted names as text (no innerHTML XSS). All read-only GraphStore methods rollback their txn (no idle-in-transaction lock). Unit G7. (Also a pre-existing prettier fix in WorkflowPreview to keep lint green.)
2026-06-23 02:48:18 +01:00
Alex 6687d85085 feat(graphrag): GraphStore — per-source graph tables in the pgvector DB
On-demand graph_nodes/graph_edges/graph_node_chunks/graph_ingest_progress in the same DB as pgvector documents (no Alembic, no cross-DB FK; Python uuids; parameterized SQL). upsert_node (merge by normalized_name, doc_freq, dedup description), add_edge (+degree), link_node_chunk, search_nodes_by_embedding (cosine), bounded get_subgraph, checkpoint, delete_by_source. Embedding dim derived from the configured model (768 fallback). Unit G2.
2026-06-23 00:51:41 +01:00