diff --git a/AGENTS.md b/AGENTS.md index 90356fc4..6aea7be9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -198,7 +198,7 @@ vale . - Parsers live in `docsgpt/parser/` and handle different document formats in the ingestion stage. - Agents and tools are in `docsgpt/agents/` and `docsgpt/agents/tools/`. - Celery setup/config lives in `docsgpt/celery_init.py` and `docsgpt/celeryconfig.py`. -- Settings and env vars are managed via Pydantic in `docsgpt/core/settings.py`. +- Settings and env vars are managed via Pydantic in `docsgpt/core/settings/` (one module per domain, composed into `Settings`). Every field needs a `description`; regenerate the docs reference with `python -m docsgpt.core.settings.reference --write`. ### Frontend diff --git a/docs/content/Deploying/DocsGPT-Settings.mdx b/docs/content/Deploying/DocsGPT-Settings.mdx index e21e3be3..0e55740e 100644 --- a/docs/content/Deploying/DocsGPT-Settings.mdx +++ b/docs/content/Deploying/DocsGPT-Settings.mdx @@ -27,13 +27,13 @@ API_KEY=YOUR_OPENAI_API_KEY LLM_NAME=gpt-4o ``` -### 2. Configuration via `settings.py` file (Advanced) +### 2. Configuration in code (Advanced) -For more advanced configurations or if you prefer to manage settings directly in code, you can modify the `settings.py` file. This file is located in the `docsgpt/core` directory of your DocsGPT project. +Every setting is defined in the `docsgpt/core/settings/` package, one module per domain (`auth.py`, `llm.py`, `embeddings.py`, ...). If you prefer to manage defaults directly in code, change them there; the `.env` file and the process environment still override whatever the code says. -While modifying `settings.py` offers more flexibility, it's generally recommended to use the `.env` file for basic settings and reserve `settings.py` for more complex adjustments or when you need to configure settings programmatically. +Using the `.env` file is recommended for day-to-day configuration. Reserve code changes for new settings or for defaults you want every deployment of your fork to share. -**Location of `settings.py`:** `docsgpt/core/settings.py` +The [Settings Reference](/Deploying/Settings-Reference) lists every setting with its type, default and description, generated from those definitions. ## Basic Settings Explained @@ -289,7 +289,7 @@ DocsGPT includes a JWT (JSON Web Token) based authentication feature for managin ### `AUTH_TYPE` Overview -The `AUTH_TYPE` setting in your `.env` file or `settings.py` determines the authentication method used by DocsGPT. This allows you to control how users authenticate with your DocsGPT instance. +The `AUTH_TYPE` setting in your `.env` file determines the authentication method used by DocsGPT. This allows you to control how users authenticate with your DocsGPT instance. | Value | Description | | ------------- | ------------------------------------------------------------------------------------------- | @@ -300,7 +300,7 @@ The `AUTH_TYPE` setting in your `.env` file or `settings.py` determines the auth #### How to Configure -Add the following to your `.env` file (or set in `settings.py`): +Add the following to your `.env` file: ```env # Shared signing key (required in production for every authentication mode) @@ -461,7 +461,6 @@ These control how sources are retrieved and whether the advanced RAG features ar | Setting | Default | Description | | --- | --- | --- | -| `RETRIEVERS_ENABLED` | `["classic", "default"]` | Allow-list of retrievers usable instance-wide. Valid keys: `classic`, `default`, `hybrid`, `graphrag`. A per-source `retriever` must be within this list. | | `PER_SOURCE_RETRIEVAL_ENABLED` | `true` | Master switch for per-source retrieval config. When `false`, all sources fall back to the classic retriever regardless of their stored config. | | `GRAPHRAG_ENABLED` | `false` | Enable [GraphRAG](/Sources/GraphRAG). Requires `VECTOR_STORE=pgvector`. | | `GRAPHRAG_EXTRACTION_MODEL` | unset | Model used for ingest-time graph extraction. Unset reuses the instance default model. | @@ -547,11 +546,10 @@ recovers. ## Exploring More Settings -These are just the basic settings to get you started. The `settings.py` file contains many more advanced options that you can explore to further customize DocsGPT, such as: +These are just the basic settings to get you started. DocsGPT has many more advanced options, such as: - Vector store configuration (`VECTOR_STORE`, Qdrant, Milvus, LanceDB settings) If you're looking for an easy way to set up a vector store with pgvector, try [Neon](https://get.neon.com/docsgpt). -- Retriever settings (`RETRIEVERS_ENABLED`) - Cache settings (`CACHE_REDIS_URL`) -- And many more! +- Sandbox, scheduler, guardrails and event-stream tuning -For a complete list of available settings and their descriptions, refer to the `settings.py` file in `docsgpt/core`. Remember to restart your Docker containers after making changes to your `.env` file or `settings.py` for the changes to take effect. +The [Settings Reference](/Deploying/Settings-Reference) lists every setting with its type, default and description. Remember to restart your Docker containers after making changes to your `.env` file for the changes to take effect. diff --git a/docs/content/Deploying/Settings-Reference.mdx b/docs/content/Deploying/Settings-Reference.mdx new file mode 100644 index 00000000..32181adb --- /dev/null +++ b/docs/content/Deploying/Settings-Reference.mdx @@ -0,0 +1,1652 @@ +--- +title: Settings Reference +description: Every DocsGPT setting, grouped by domain, with its type, default and purpose. +--- + +{/* GENERATED FILE. Do not edit by hand: run `python -m docsgpt.core.settings.reference --write`. */} + +# Settings Reference + +Every setting DocsGPT reads, generated from `docsgpt/core/settings/`. Each one is +an environment variable of the same name, set in `.env` or the process +environment; see [App Configuration](/Deploying/DocsGPT-Settings) for how the +file is found and for worked examples. `` below is the data home +described there. + + +## Authentication + +How users authenticate: none, a shared token, per-session JWTs, or OIDC SSO. + +### `AUTH_TYPE` + +Type `"simple_jwt" | "session_jwt" | "oidc"`, default unset. + +Authentication mode: simple_jwt, session_jwt, oidc, or unset (None) for no authentication. + +### `JWT_SECRET_KEY` + +Type `str`, default `""`. + +Signing key for session tokens and other signed capabilities. Required on every replica in production; local development may fall back to a key generated on disk. + +### `ENCRYPTION_SECRET_KEY` + +Type `str`, default `default-docsgpt-encryption-key`. + +Key used to encrypt stored credentials such as tool and connector secrets. + +### `INTERNAL_KEY` + +Type `str`, default unset. + +Internal API key for worker-to-backend authentication. + +### `OIDC_ISSUER` + +Type `str`, default unset. + +OIDC issuer URL with discovery, e.g. https://auth.example.com/application/o/docsgpt/. + +### `OIDC_CLIENT_ID` + +Type `str`, default unset. + +OIDC client id. + +### `OIDC_CLIENT_SECRET` + +Type `str`, default unset. + +OIDC client secret. Optional; PKCE is always used. + +### `OIDC_SCOPES` + +Type `str`, default `openid profile email`. + +Scopes requested from the IdP. + +### `OIDC_USER_ID_CLAIM` + +Type `str`, default `sub`. + +ID-token claim mapped to the DocsGPT user id. + +### `OIDC_FRONTEND_URL` + +Type `str`, default unset. + +Browser-facing app origin, e.g. http://localhost:5173. + +### `OIDC_REDIRECT_URI` + +Type `str`, default unset. + +Override for the callback URL; default is <request host>/api/auth/oidc/callback. + +### `OIDC_SESSION_LIFETIME_SECONDS` + +Type `int`, default `28800`, must be `> 0`. + +Lifetime of the minted session JWT in seconds (8h). + +### `OIDC_PROVIDER_NAME` + +Type `str`, default unset. + +Sign-in button label, e.g. "Acme SSO". + +### `OIDC_ALLOWED_GROUPS` + +Type `str`, default unset. + +Comma-separated group allowlist; unset admits any authenticated user. + +### `OIDC_GROUPS_CLAIM` + +Type `str`, default `groups`. + +ID-token/userinfo claim carrying group membership. + +### `OIDC_ADMIN_GROUPS` + +Type `str`, default unset. + +Comma-separated groups granted admin; unset means no OIDC admin mapping. + +### `LOCAL_MODE_ADMIN` + +Type `bool`, default `false`. + +Grant admin without a database role. Persisted admin grants live in user_roles (AUTH_TYPE=oidc only); this is the only non-DB admin path, for AUTH_TYPE=None self-host. MUST stay False if networked. + +### `SCIM_ENABLED` + +Type `bool`, default `false`. + +Enable SCIM 2.0 provisioning at /scim/v2. + +### `SCIM_TOKEN` + +Type `str`, default unset. + +Bearer token for IdP SCIM clients (required when SCIM is enabled). + + +## LLM providers + +Which model answers, how it is reached, and provider-specific behaviour. + +### `LLM_PROVIDER` + +Type `str`, default `docsgpt`. + +LLM provider key, e.g. openai, anthropic, docsgpt. + +### `LLM_NAME` + +Type `str`, default unset. + +Model name for the provider; with openai, e.g. gpt-4 or gpt-3.5-turbo. + +### `API_KEY` + +Type `str`, default unset. + +LLM API key used by LLM_PROVIDER. + +### `OPENAI_API_KEY` + +Type `str`, default unset. + +OpenAI API key. + +### `ANTHROPIC_API_KEY` + +Type `str`, default unset. + +Anthropic API key. + +### `GOOGLE_API_KEY` + +Type `str`, default unset. + +Google AI API key. + +### `GROQ_API_KEY` + +Type `str`, default unset. + +Groq API key. + +### `HUGGINGFACE_API_KEY` + +Type `str`, default unset. + +Hugging Face API key. + +### `OPEN_ROUTER_API_KEY` + +Type `str`, default unset. + +OpenRouter API key. + +### `NOVITA_API_KEY` + +Type `str`, default unset. + +Novita API key. + +### `OPENAI_API_BASE` + +Type `str`, default unset. + +Azure OpenAI API base URL. + +### `OPENAI_API_VERSION` + +Type `str`, default unset. + +Azure OpenAI API version. + +### `AZURE_DEPLOYMENT_NAME` + +Type `str`, default unset. + +Azure deployment name for answering. + +### `AZURE_EMBEDDINGS_DEPLOYMENT_NAME` + +Type `str`, default unset. + +Azure deployment name for embeddings. + +### `OPENAI_BASE_URL` + +Type `str`, default unset. + +Base URL for OpenAI-compatible model servers. + +### `LLM_PATH` + +Type `str`, default `/models/docsgpt-7b-f16.gguf`. + +Path to the local GGUF model used by the llama.cpp provider. + +### `FALLBACK_LLM_PROVIDER` + +Type `str`, default unset. + +Provider for the fallback LLM. + +### `FALLBACK_LLM_NAME` + +Type `str`, default unset. + +Model name for the fallback LLM. + +### `FALLBACK_LLM_API_KEY` + +Type `str`, default unset. + +API key for the fallback LLM. + +### `TITLE_MODEL_ID` + +Type `str`, default unset. + +Optional cheaper model for conversation titles; unset reuses the answer model. + +### `MODELS_CONFIG_DIR` + +Type `str`, default unset. + +Directory of operator-supplied model YAMLs, loaded after the built-in catalog; later wins on duplicate model id. See docsgpt/core/models/README.md. + +### `DEFAULT_LLM_TOKEN_LIMIT` + +Type `int`, default `128000`. + +Context window assumed when the model is not found in the registry. + +### `RESERVED_TOKENS` + +Type `dict[str, int]`, default `{"system_prompt": 500, "current_query": 500, "safety_buffer": 1000}`. + +Tokens held back from the context window for the system prompt, the query and a safety buffer. + +### `CACHE_REDIS_URL` + +Type `str`, default `redis://localhost:6379/2`. + +Redis URL for the LLM cache. + +### `OPENAI_RESPONSES_STORE` + +Type `bool`, default `false`. + +True persists Responses API calls server-side so previous_response_id can chain turns. False keeps them stateless, carrying reasoning across the tool loop as encrypted items. + +### `OPENAI_RESPONSES_CHAIN_ACROSS_TURNS` + +Type `bool`, default `true`. + +Cross-turn previous_response_id chaining (store mode only). The chained transcript lives on the provider and is invisible to every local guard, so it is bounded: a turn starts from the local history when the previous turn's reported prompt already reached the budget (default: the model's context window) or when the conversation was compressed after that turn was produced. + +### `OPENAI_RESPONSES_CHAIN_BUDGET_TOKENS` + +Type `int`, default unset. + +Prompt-token budget for cross-turn chaining; unset uses the model's context window. + +### `OPENAI_RESPONSES_TRUNCATION_AUTO` + +Type `bool`, default `false`. + +Send truncation: "auto" so the provider drops the oldest input items instead of failing every request once a chain exceeds the model's window. + +### `OPENAI_PROMPT_CACHE_KEY` + +Type `bool`, default `true`. + +Route a user's Responses API calls to the same prompt-cache shard with an opaque per-user key. + +### `OPENAI_PROMPT_CACHE_RETENTION` + +Type `str`, default unset. + +Request extended prompt-cache retention where the provider offers it. + +### `OPENAI_REASONING_SUMMARY` + +Type `str`, default `auto`. + +Reasoning summary mode requested from the Responses API. + + +## Embeddings + +The embedding model, remote or local, and the batching around it. + +### `EMBEDDINGS_NAME` + +Type `str`, default `huggingface_sentence-transformers/all-mpnet-base-v2`. + +Embedding model. The legacy model is the default on purpose: an install that never pinned this has vectors from it, and granite is the same width so a swap would fail silently. New installs get granite from .env-template; existing ones switch by setting this and running docsgpt.scripts.reembed. + +### `EMBEDDINGS_BASE_URL` + +Type `str`, default unset. + +Remote embeddings API URL (OpenAI-compatible). + +### `EMBEDDINGS_KEY` + +Type `str`, default unset. + +API key for embeddings (with OpenAI, the same value as API_KEY). + +### `EMBEDDINGS_MAX_INPUT_TOKENS` + +Type `int`, default unset. + +Truncate each remote embed input to N tokens (overflow is lost). + +### `EMBEDDINGS_BATCH_SIZE` + +Type `int`, default `32`, must be `>= 1`. + +Chunks per store transaction and per remote embed request. + +### `EMBEDDINGS_MODEL_BATCH_SIZE` + +Type `int`, default `1`, must be `>= 1`. + +Documents per local ONNX forward pass. Each pass pads to its longest input, and that waste grows with the square of chunk length: at 1250 tokens, 32 peaked at 6.6 GB, 1 at 2.9 GB. + +### `EMBEDDINGS_THREADS` + +Type `int`, default unset. + +Intra-op threads for the local ONNX runner; unset uses every core. It scales sub-linearly, so several single-threaded workers beat one many-threaded process on the same cores. + +### `EMBEDDINGS_CACHE_DIR` + +Type `str`, default `/models`. + +Where embedding models and their tokenizers are cached. Persistent by default: FastEmbed's own default is the temp dir. + +### `EMBEDDINGS_POOLING` + +Type `"cls" | "mean"`, default unset. + +Pooling strategy ("cls" or "mean"). Read from the model's own repository; set only for a repository that declares none, or to override what it declares. + +### `EMBEDDINGS_NORMALIZE` + +Type `bool`, default unset. + +L2-normalise embeddings. Read from the model's own repository; set only for a repository that declares nothing, or to override what it declares. + +### `EMBEDDINGS_DELEGATE_TO_WORKER` + +Type `bool`, default `true`. + +Embed on the worker so the API holds no model (~890 MB), at one broker round trip per query. Ignored when EMBEDDINGS_BASE_URL is set, which is the better answer for production. + +### `EMBEDDINGS_QUEUE` + +Type `str`, default `embeddings`. + +Celery queue the embed task is routed to. + +### `EMBEDDINGS_DELEGATE_TIMEOUT` + +Type `int`, default `60`, must be `> 0`. + +Seconds the API waits for the worker to return an embedding. + + +## Retrieval + +Which vector store answers searches and how retrieval fans out across sources. + +### `VECTOR_STORE` + +Type `"faiss" | "elasticsearch" | "mongodb" | "qdrant" | "milvus" | "pgvector"`, default `faiss`. + +Vector store backend. + +### `RETRIEVAL_MAX_PARALLEL_SOURCES` + +Type `int`, default `4`, must be `>= 1`. + +Concurrent per-source searches in one retrieval; the query is embedded once and shared. + +### `PER_SOURCE_RETRIEVAL_ENABLED` + +Type `bool`, default `true`. + +Kill-switch for per-source retrieval dispatch; False collapses to a single retriever. + +### `GRAPHRAG_ENABLED` + +Type `bool`, default `false`. + +Gates graph-aware ingestion and retrieval. + +### `GRAPHRAG_EXTRACTION_MODEL` + +Type `str`, default unset. + +Model for ingest-time graph extraction; unset reuses LLM_PROVIDER/LLM_NAME. + +### `GRAPHRAG_MAX_CHUNKS_FOR_EXTRACTION` + +Type `int`, default `2000`, must be `>= 0`. + +Hard cap on chunks extracted per source (cost control); 0 extracts nothing. + + +## Vector stores + +Per-backend connection details; only the backend named by VECTOR_STORE is read. + +### `MONGO_URI` + +Type `str`, default unset. + +Only consulted when VECTOR_STORE=mongodb or when running scripts/db/backfill.py; user data lives in Postgres. + +### `ELASTIC_CLOUD_ID` + +Type `str`, default unset. + +Elastic Cloud id. + +### `ELASTIC_USERNAME` + +Type `str`, default unset. + +Elasticsearch username. + +### `ELASTIC_PASSWORD` + +Type `str`, default unset. + +Elasticsearch password. + +### `ELASTIC_URL` + +Type `str`, default unset. + +Elasticsearch URL. + +### `ELASTIC_INDEX` + +Type `str`, default `docsgpt`. + +Elasticsearch index name. + +### `QDRANT_COLLECTION_NAME` + +Type `str`, default `docsgpt`. + +Qdrant collection name. + +### `QDRANT_LOCATION` + +Type `str`, default unset. + +Qdrant location (':memory:' or a URL). + +### `QDRANT_URL` + +Type `str`, default unset. + +Qdrant server URL. + +### `QDRANT_PORT` + +Type `int`, default `6333`. + +Qdrant REST port. + +### `QDRANT_GRPC_PORT` + +Type `int`, default `6334`. + +Qdrant gRPC port. + +### `QDRANT_PREFER_GRPC` + +Type `bool`, default `false`. + +Use gRPC instead of REST where possible. + +### `QDRANT_HTTPS` + +Type `bool`, default unset. + +Use HTTPS for the Qdrant connection. + +### `QDRANT_API_KEY` + +Type `str`, default unset. + +Qdrant API key. + +### `QDRANT_PREFIX` + +Type `str`, default unset. + +URL prefix for a Qdrant behind a proxy. + +### `QDRANT_TIMEOUT` + +Type `float`, default unset. + +Qdrant request timeout in seconds. + +### `QDRANT_HOST` + +Type `str`, default unset. + +Qdrant host (alternative to QDRANT_URL). + +### `QDRANT_PATH` + +Type `str`, default unset. + +Path for an embedded on-disk Qdrant. + +### `QDRANT_DISTANCE_FUNC` + +Type `str`, default `Cosine`. + +Qdrant distance function. + +### `PGVECTOR_CONNECTION_STRING` + +Type `str`, default unset. + +pgvector connection string. postgres://, postgresql:// and postgresql+psycopg:// are all accepted and normalized internally for psycopg.connect(). Unset falls back to POSTGRES_URI. + +### `PGVECTOR_POOL_MAX_SIZE` + +Type `int`, default `8`, must be `>= 0`. + +Per-process connection pool size; 0 uses one direct connection per store. + +### `PGVECTOR_IVFFLAT_PROBES` + +Type `int`, default unset. + +IVFFlat probes; unset derives sqrt(lists) from the index. Higher means better recall, more scan. + +### `MILVUS_COLLECTION_NAME` + +Type `str`, default `docsgpt`. + +Milvus collection name. + +### `MILVUS_URI` + +Type `str`, default `/milvus_local.db`. + +Milvus server URI. The default is a milvus-lite (embedded) database file under the data home, like the other local stores. + +### `MILVUS_TOKEN` + +Type `str`, default `""`. + +Milvus auth token. + +### `LANCEDB_PATH` + +Type `str`, default `/data/lancedb`. + +LanceDB local data directory. + +### `LANCEDB_TABLE_NAME` + +Type `str`, default `docsgpts`. + +LanceDB table for stored vectors. + + +## User-data database + +The Postgres database holding users, conversations and sources, and what startup may do to it. + +### `POSTGRES_URI` + +Type `str`, default unset. + +User-data Postgres connection URI. + +### `AUTO_MIGRATE` + +Type `bool`, default `true`. + +On startup, apply pending Alembic migrations. Disable if you manage schema out-of-band. + +### `AUTO_CREATE_DB` + +Type `bool`, default `true`. + +On startup, create the target Postgres database if missing (needs CREATEDB privilege). + +### `AUTO_VECTOR_SCHEMA` + +Type `bool`, default `true`. + +On startup, create the pgvector/graph tables and verify the embedding dimension. No Alembic migration covers the vector DB (it may be a separate cluster); set False to manage it yourself. + + +## Workers + +How background tasks are queued and how worker processes are recycled. + +### `CELERY_BROKER_URL` + +Type `str`, default `redis://localhost:6379/0`. + +Celery broker URL. + +### `CELERY_RESULT_BACKEND` + +Type `str`, default `redis://localhost:6379/1`. + +Celery result backend URL. + +### `CELERY_WORKER_PREFETCH_MULTIPLIER` + +Type `int`, default `1`. + +Tasks prefetched per worker process; 1 caps SIGKILL loss to one task. + +### `CELERY_VISIBILITY_TIMEOUT` + +Type `int`, default `3600`, must be `> 0`. + +Broker visibility timeout in seconds. Must exceed the longest legitimate task runtime but stay short enough that SIGKILLed tasks redeliver promptly. + +### `CELERY_WORKER_MAX_MEMORY_PER_CHILD` + +Type `int`, default `4194304`, must be `>= 0`. + +Recycle a prefork child past this resident size in KB; backstops docling/torch heap growth. Checked between tasks, so it does not bound the peak within one. 0 disables. + +### `CELERY_WORKER_MAX_TASKS_PER_CHILD` + +Type `int`, default `0`, must be `>= 0`. + +Recycle a worker child after N tasks; 0 disables. + +### `API_URL` + +Type `str`, default `http://localhost:7091`. + +Backend URL the Celery worker calls back into. + + +## Ingestion and parsing + +Upload limits, the parser engine, and per-format byte caps for ingestion and attachments. + +### `UPLOAD_FOLDER` + +Type `str`, default `inputs`. + +Directory under the data home for uploaded sources. + +### `UPLOAD_MAX_REQUEST_BYTES` + +Type `int`, default `268435456`, must be `> 0`. + +Cap on an upload request body; applied by Flask before multipart parsing. + +### `UPLOAD_MAX_FILE_BYTES` + +Type `int`, default `104857600`, must be `> 0`. + +Cap on a single uploaded file; also enforced while copying. + +### `PARSE_SPEC_MAX_BYTES` + +Type `int`, default `10485760`, must be `> 0`. + +Cap on an OpenAPI/tool spec file accepted for parsing. + +### `UPLOAD_MAX_ARCHIVE_BYTES` + +Type `int`, default `262144000`, must be `> 0`. + +Cap on total bytes extracted from one uploaded archive. + +### `UPLOAD_MAX_ARCHIVE_FILES` + +Type `int`, default `10000`, must be `> 0`. + +Cap on files extracted from one uploaded archive. + +### `UPLOAD_MAX_ARCHIVE_RATIO` + +Type `int`, default `1000`, must be `> 0`. + +Maximum decompressed-to-compressed ratio before an archive is rejected. + +### `UPLOAD_MAX_ARCHIVE_DEPTH` + +Type `int`, default `3`, must be `>= 0`. + +Maximum nesting depth of archives inside archives. + +### `PARSE_PDF_AS_IMAGE` + +Type `bool`, default `false`. + +Render PDF pages to images before parsing. + +### `PARSE_IMAGE_REMOTE` + +Type `bool`, default `false`. + +Send images to a remote parser. + +### `DOC_PARSER_ENGINE` + +Type `"anydoc" | "docling"`, default `anydoc`. + +Document parser for source ingestion, chat attachments and the read_document tool. "anydoc" (default): firecrawl-anydoc, a Rust converter with no ML models; milliseconds per file, ~100 MB peak RSS. "docling": the layout/table-model pipeline (optional install; needed for read_document's structured output and the docling OCR backend). Files anydoc cannot convert (scanned PDFs, malformed input) fall back to docling when it is installed, otherwise to the native OCR parsers (OCR on) or the legacy parsers. Rollback to the previous behaviour is this one variable. + +### `DOCLING_PIPELINE_QUEUE_MAX_SIZE` + +Type `int`, default `2`. + +Pages docling's threaded pipeline buffers in flight; the library default (100) drives worker RSS to ~3 GB on a mid-size PDF. + +### `DOCLING_COMPILE_TORCH_MODELS` + +Type `bool`, default `false`. + +Let docling torch.compile its models (slower start, faster pages). + +### `DOCLING_TABULAR_MAX_BYTES` + +Type `int`, default `2000000`. + +Largest CSV/XLSX docling will parse, in bytes. + +### `DOCLING_MARKUP_MAX_BYTES` + +Type `int`, default `8000000`. + +Largest HTML/XML docling will parse, in bytes. + +### `MARKUP_MAX_BYTES` + +Type `int`, default `8000000`, must be `>= 0`. + +HTML/XHTML larger than this (bytes) are head-truncated before the markdownify parser runs (the anydoc engine's HTML path). The tree that path builds costs ~50x the input (30 MB of HTML measured at 1.6 GB RSS) and the upload cap is 100 MB, so the gate is what keeps one upload from taking the ingest worker down. 0 disables it. + +### `PDF_TRUST_CHECK` + +Type `bool`, default `true`. + +Trust-check anydoc's PDF output (docsgpt/parser/file/pdf_trust.py): flag composite (Type0) fonts without a ToUnicode map, and CJK-declaring PDFs whose extracted text has almost no CJK, the two classes where anydoc drops text silently. A flagged file re-parses on the docling fallback when docling is installed; otherwise the anydoc output is kept and the document gets extra_info["parse_warnings"]. ~30 ms per scanned MB. + +### `ANYDOC_TABLEIZE` + +Type `bool`, default `false`. + +Rewrite dot-leader / whitespace-aligned table runs in anydoc's PDF markdown into GFM tables (docsgpt/parser/file/tableize.py). Off by default: it rewrites content on a heuristic (>=3 uniform label+numbers lines) validated only on a small corpus so far. + +### `ATTACHMENT_PDF_TEXT_FAST_PATH` + +Type `bool`, default `true`. + +Read PDF attachments via their embedded text layer (pypdfium2) instead of docling, falling back to docling when there is no text layer. Attachments go into a prompt, so docling's structural markdown earns far less than the tens of seconds per file it costs; source ingestion is unaffected because chunking and retrieval do depend on that structure. + +### `ATTACHMENT_PDF_TEXT_MIN_MEDIAN_CHARS` + +Type `int`, default `32`. + +Median chars per sampled page below which a PDF attachment is treated as a scan and handed to docling. Measured on real uploads: scans at 0-17 chars/page, text-layer documents at 433-6834. + +### `ATTACHMENT_TEXT_MAX_BYTES` + +Type `int`, default `5000000`. + +Cap on extracted attachment text. + +### `AGENT_IMAGE_MAX_BYTES` + +Type `int`, default `5000000`. + +Cap on an image passed to an agent. + +### `AGENT_IMAGE_MAX_PIXELS` + +Type `int`, default `16777216`. + +Cap on the pixel count of an image passed to an agent. + +### `GITHUB_INGEST_MAX_FILE_BYTES` + +Type `int`, default `1048576`, must be `>= 0`. + +Skip GitHub repo blobs larger than this (0 = no cap). + +### `GITHUB_INGEST_MAX_WORKERS` + +Type `int`, default `8`, must be `>= 1`. + +Parallel file fetches per GitHub repo ingest. + +### `DOCUMENT_PARSE_QUEUE` + +Type `str`, default `parsing`. + +Celery queue the parse_document task is routed to. + +### `DOCUMENT_PARSE_TIMEOUT` + +Type `int`, default `120`. + +Seconds the read_document tool awaits the enqueued parse before degrading. + +### `DOCUMENT_PARSE_TIMEOUT_PER_MB` + +Type `int`, default `60`. + +Extra seconds of parse window per MiB of input. The base timeout is a FLOOR: the window grows with document size because OCR cost scales with pages. Without this a large scan is silently dropped at the base window. + +### `DOCUMENT_PARSE_TIMEOUT_MAX` + +Type `int`, default `900`. + +Absolute ceiling on the size-scaled parse window, in seconds. + +### `DOCUMENT_PARSE_MAX_BYTES` + +Type `int`, default `0`, must be `>= 0`. + +Cap on a parsed document's bytes (0 = reuse SANDBOX_MAX_INPUT_BYTES). + +### `DOCUMENT_MAX_DECOMPRESSED_BYTES` + +Type `int`, default `314572800`. + +Cap on bytes decompressed from an archive handed to read_document. + +### `DOCUMENT_MAX_ARCHIVE_ENTRIES` + +Type `int`, default `10000`. + +Cap on entries in an archive handed to read_document. + + +## OCR + +Whether OCR runs, which stack performs it, and which engine it uses. + +### `OCR_ENABLED` + +Type `bool`, default `false`, also read from `DOCLING_OCR_ENABLED`. + +OCR scanned PDFs and images during source ingestion. + +### `OCR_ATTACHMENTS_ENABLED` + +Type `bool`, default `false`, also read from `DOCLING_OCR_ATTACHMENTS_ENABLED`. + +OCR scanned PDFs and images attached to a chat. + +### `OCR_BACKEND` + +Type `"auto" | "docling" | "native"`, default `auto`. + +Which stack runs OCR when it is on. auto: docling when installed, otherwise native. docling: the layout-model pipeline (hybrid region OCR, reading order, table structure); needs the optional docling extra. native: pypdfium2/Pillow page rendering straight into tesseract or a DeepSeek-OCR endpoint (docsgpt/parser/file/ocr_parser.py); no ML models in the worker, tables come out as text lines under tesseract. + +### `OCR_ENGINE` + +Type `"tesseract" | "deepseek" | "auto" | "ocrmac" | "rapidocr"`, default `tesseract`. + +OCR engine used when OCR is on. Benched 2026-08 on EN/ZH/table/degraded scans (docs/Guides/ocr has the menu). tesseract (recommended): best classic-engine accuracy (perfect EN word recall, 0.000 bilingual CER, 100% table cells), ~35 MB, CPU-only; needs the system binary and language packs, an optional install like every OCR dependency (build with INSTALL_TESSERACT=true, or apt/brew install tesseract-ocr for a local run); both backends. deepseek: DeepSeek-OCR against an Ollama/vLLM endpoint (OCR_DEEPSEEK_*); best table/CJK quality, the worker stays light (no layout models) but each page costs seconds on the model server; both backends. auto: docling's pick, ocrmac on macOS (excellent), rapidocr on Linux (silently shreds some long text lines; avoid as a server default). ocrmac | rapidocr: force one of those. auto/ocrmac/rapidocr exist only inside docling; the native backend runs tesseract for them. An engine that is not installed degrades (docling: to auto) with a warning instead of failing the parse. + +### `OCR_LANGS` + +Type `str`, default `eng`. + +Tesseract language packs, "+"-separated (e.g. "eng+chi_sim+deu"). Other engines keep their own defaults; their language codes differ. + +### `OCR_DEEPSEEK_URL` + +Type `str`, default `http://localhost:11434/v1/chat/completions`. + +Chat-completions URL of the DeepSeek-OCR endpoint (Ollama or vLLM). + +### `OCR_DEEPSEEK_MODEL` + +Type `str`, default `deepseek-ocr:3b`. + +Model name at the DeepSeek-OCR endpoint. + +### `OCR_DEEPSEEK_TIMEOUT` + +Type `float`, default `300.0`. + +Seconds allowed per page request to the DeepSeek endpoint, on both backends (native sends pages one at a time; docling's VLM pipeline keeps its own concurrency). A 3B model on a laptop needs minutes; a vLLM GPU deployment, seconds. + +### `OCR_RENDER_DPI` + +Type `int`, default `200`. + +Native backend only: resolution at which pages without a text layer are rendered before OCR. 200 suits tesseract; clamped to 72-600. + +### `OCR_MIN_CHARS_PER_PAGE` + +Type `int`, default `20`, must be `>= 0`, also read from `DOCLING_OCR_MIN_CHARS_PER_PAGE`. + +Chars-per-page floor below which an OCR'd PDF/image parse is treated as an OCR dropout rather than as content (long-running docling workers were observed returning zero characters for every scanned page after a long scanned PDF, with no error). docling retries once on a fresh full-page-OCR converter; both backends then fail loudly instead of indexing an empty document. 0 disables the guard. + + +## File storage + +Local disk or an S3-compatible bucket, and how download URLs are produced. + +### `STORAGE_TYPE` + +Type `"local" | "s3"`, default `local`. + +File storage backend. + +### `URL_STRATEGY` + +Type `"backend" | "s3"`, default `backend`. + +How download links are produced: backend (streamed through the API) or s3 (presigned URLs). + +### `S3_BUCKET_NAME` + +Type `str`, default `docsgpt-test-bucket`. + +Bucket name. + +### `S3_ENDPOINT_URL` + +Type `str`, default unset. + +Custom endpoint for S3-compatible services (MinIO, R2, B2, Spaces); omit for AWS. + +### `S3_ACCESS_KEY_ID` + +Type `str`, default unset. + +Access key id. + +### `S3_SECRET_ACCESS_KEY` + +Type `str`, default unset. + +Secret access key. + +### `S3_REGION` + +Type `str`, default unset. + +AWS region; use "auto" for Cloudflare R2. + +### `S3_PATH_STYLE` + +Type `bool`, default `false`. + +Path-style addressing (required by most non-AWS services). + +### `SAGEMAKER_REGION` + +Type `str`, default unset. + +**Deprecated.** Set S3_REGION instead; the SAGEMAKER_* fallback will be removed. + +Legacy AWS region from the retired SageMaker provider; deprecated fallback for S3_REGION. + +### `SAGEMAKER_ACCESS_KEY` + +Type `str`, default unset. + +**Deprecated.** Set S3_ACCESS_KEY_ID instead; the SAGEMAKER_* fallback will be removed. + +Legacy AWS access key from the retired SageMaker provider; deprecated fallback for S3_ACCESS_KEY_ID. + +### `SAGEMAKER_SECRET_KEY` + +Type `str`, default unset. + +**Deprecated.** Set S3_SECRET_ACCESS_KEY instead; the SAGEMAKER_* fallback will be removed. + +Legacy AWS secret key from the retired SageMaker provider; deprecated fallback for S3_SECRET_ACCESS_KEY. + + +## Connectors + +Client credentials and callback URLs for Google Drive, Microsoft, Confluence, GitHub and MCP. + +### `GOOGLE_CLIENT_ID` + +Type `str`, default unset. + +Google OAuth client id. + +### `GOOGLE_CLIENT_SECRET` + +Type `str`, default unset. + +Google OAuth client secret. + +### `CONNECTOR_REDIRECT_BASE_URI` + +Type `str`, default `http://127.0.0.1:7091/api/connectors/callback`. + +OAuth callback URL; register it as-is in your provider's console (e.g. GCP). + +### `CONNECTOR_ALLOWED_ORIGINS` + +Type `str`, default unset. + +Comma-separated frontend origins allowed to receive connector OAuth results, e.g. https://docsgpt.example.com. The callback origin and OIDC_FRONTEND_URL are always allowed; a loopback callback also allows localhost:5173. + +### `MICROSOFT_CLIENT_ID` + +Type `str`, default unset. + +Azure AD application (client) id. + +### `MICROSOFT_CLIENT_SECRET` + +Type `str`, default unset. + +Azure AD application client secret. + +### `MICROSOFT_TENANT_ID` + +Type `str`, default `common`. + +Azure AD tenant id, or 'common' for multi-tenant. + +### `MICROSOFT_AUTHORITY` + +Type `str`, default unset. + +Authority URL override; unset derives https://login.microsoftonline.com/<MICROSOFT_TENANT_ID>. + +### `CONFLUENCE_CLIENT_ID` + +Type `str`, default unset. + +Confluence Cloud OAuth client id. + +### `CONFLUENCE_CLIENT_SECRET` + +Type `str`, default unset. + +Confluence Cloud OAuth client secret. + +### `GITHUB_ACCESS_TOKEN` + +Type `str`, default unset. + +GitHub PAT with read access to repositories. + +### `MCP_OAUTH_REDIRECT_URI` + +Type `str`, default unset. + +Public callback URL for MCP OAuth; unset derives it from CONNECTOR_REDIRECT_BASE_URI. + + +## Server + +Serving the UI, public URLs, and process-level knobs of the API server. + +### `DEPLOYMENT_TYPE` + +Type `str`, default unset. + +Deployment class, e.g. cloud or production. A production class refuses to run without a configured JWT_SECRET_KEY instead of generating a local one on disk. + +### `SERVE_UI` + +Type `bool`, default `true`. + +Serve the web UI shipped in the package (docsgpt/static) from the API process. + +### `FLASK_DEBUG_MODE` + +Type `bool`, default `false`. + +Run Flask in debug mode. + +### `VERSION_CHECK` + +Type `bool`, default `true`. + +Anonymous startup version check for security issues. + +### `PUBLIC_API_BASE_URL` + +Type `str`, default unset. + +Public base URL for user-facing endpoint references in prompts. + +### `GRACEFUL_SHUTDOWN_TIMEOUT_SECONDS` + +Type `int`, default `30`. + +Bounds uvicorn's shutdown drain (uvicorn_worker doesn't forward --graceful-timeout). Keep below the gunicorn --timeout (180) watchdog. Used by BoundedDrainUvicornWorker. + +### `WSGI_THREADPOOL_WORKERS` + +Type `int`, default `96`, must be `>= 1`. + +Threads serving the WSGI (Flask) part of the app under the ASGI server. + +### `V1_SESSION_TTL_SECONDS` + +Type `int`, default `86400`. + +Lets OpenAI-compatible clients identify a logical chat by session header, which chat-completions itself has no field for; TTL of that session mapping. + + +## Events and devices + +The internal push channel (notifications and durable replay) and the Redis pool behind it. + +### `ENABLE_SSE_PUSH` + +Type `bool`, default `true`. + +Internal SSE push channel (notifications and durable replay journal). False makes /api/events emit "push_disabled" and return; clients fall back to polling. + +### `EVENTS_STREAM_MAXLEN` + +Type `int`, default `1000`, must be `>= 1`. + +Per-user durable backlog cap in entries; ~24h of replay at typical rates. + +### `SSE_KEEPALIVE_SECONDS` + +Type `int`, default `15`, must be `>= 1`. + +Interval between SSE keepalive comments. + +### `SSE_MAX_CONCURRENT_PER_USER` + +Type `int`, default `8`, must be `>= 0`. + +Simultaneous SSE connections per user; each holds a pooled async Redis connection for its lifetime. 8 covers multi-tab use without one user starving the pool. 0 disables. + +### `ASYNC_REDIS_MAX_CONNECTIONS` + +Type `int`, default `2000`, must be `>= 1`. + +Pool size of the async Redis client behind the event-loop routes, per process. Every open notification tab, chat reconnect and device session holds one connection, so this caps concurrent streams per worker (redis-py's own default is 100). Keep the total across workers below the Redis server's maxclients (10000 by default). + +### `EVENTS_REPLAY_MAX_PER_REQUEST` + +Type `int`, default `200`, must be `>= 1`. + +Backlog entries XRANGE returns per /api/events snapshot. Bounds what one replay moves from Redis to the wire: a client looping Last-Event-ID reconnects enumerates at most this many per round-trip. + +### `EVENTS_REPLAY_MAX_AGE_HOURS` + +Type `int`, default `48`. + +Oldest backlog entry a replay will return. + +### `EVENTS_REPLAY_BUDGET_REQUESTS_PER_WINDOW` + +Type `int`, default `30`. + +Sliding-window cap on snapshot replays per user; exhausting it returns 429 with the cursor pinned so the client backs off until the window rolls over. + +### `EVENTS_REPLAY_BUDGET_WINDOW_SECONDS` + +Type `int`, default `60`. + +Length of the replay budget window. + +### `MESSAGE_EVENTS_RETENTION_DAYS` + +Type `int`, default `14`, must be `> 0`. + +Retention for the message_events journal, enforced by the cleanup_message_events beat task. Replay only needs streams a client could still be tailing. + +### `REMOTE_DEVICE_SESSION_IDLE_SECONDS` + +Type `int`, default `60`, must be `> 0`. + +Seconds without a heartbeat before a remote-device session is considered idle. + +### `REMOTE_DEVICE_REQUIRE_SIGNATURE` + +Type `bool`, default `false`. + +Require signed commands from remote devices. + +### `REMOTE_DEVICE_PAIRING_TTL_SECONDS` + +Type `int`, default `600`, must be `> 0`. + +Lifetime of a pairing code. + +### `REMOTE_DEVICE_CMD_QUEUE_TTL_SECONDS` + +Type `int`, default `900`, must be `> 605`. + +Redis TTL of the per-device command queue, routing invocations cross-process so a scheduled run reaches the web-held device session. Must exceed the max drain deadline (605s) so a command for a briefly-offline device isn't evicted before its own drain gives up. + +### `REMOTE_DEVICE_INVOCATION_TTL_SECONDS` + +Type `int`, default `900`, must be `> 0`. + +Redis TTL of a pending remote-device invocation. + +### `REMOTE_DEVICE_OUTPUT_STREAM_MAXLEN` + +Type `int`, default `10000`. + +Cap on buffered output entries per remote-device invocation stream. + + +## Agents + +What an agent may do per turn and how its context is kept within budget. + +### `AGENT_NAME` + +Type `str`, default `classic`. + +Default agent type for agentless chats. + +### `DEFAULT_AGENT_LIMITS` + +Type `dict[str, int]`, default `{"token_limit": 50000, "request_limit": 500}`. + +Per-agent default quotas: tokens and requests. + +### `DEFAULT_CHAT_TOOLS` + +Type `list[str]`, default `["memory", "read_webpage", "scheduler"]`. + +Config-free tools on by default in agentless chats. scheduler is dual-registered in BUILTIN_AGENT_TOOLS so one synthetic id resolves via defaults or the agent picker. Add code_executor and artifact_generator once a sandbox runner is configured; both execute through it and would fail on every call without one. + +### `ENABLE_TOOL_PREFETCH` + +Type `bool`, default `true`. + +Pre-fetch retrieval before the agent's first turn. + +### `TOOL_RESULT_MAX_TOKENS` + +Type `int`, default `20000`, must be `>= 0`. + +Cap on one tool result entering the LLM context (0 disables); journal and DB keep it whole. + +### `ENABLE_CONVERSATION_COMPRESSION` + +Type `bool`, default `true`. + +Compress long conversations once they approach the context window. + +### `COMPRESSION_THRESHOLD_PERCENTAGE` + +Type `float`, default `0.8`, must be `> 0` and `<= 1`. + +Fraction of the context window at which compression triggers. + +### `COMPRESSION_MODEL_OVERRIDE` + +Type `str`, default unset. + +Use a different model for compression; unset reuses the answer model. + +### `COMPRESSION_PROMPT_VERSION` + +Type `str`, default `v1.0`. + +Tracks compression prompt iterations. + +### `COMPRESSION_MAX_HISTORY_POINTS` + +Type `int`, default `3`. + +Keep only the last N compression points to prevent DB bloat. + +### `COMPRESSION_RECENT_FIELD_MAX_TOKENS` + +Type `int`, default `8000`, must be `>= 0`. + +Per-field cap on the verbatim tail kept after a compression point (0 disables). + +### `WORKFLOW_NODE_NATIVE_MAX_FILES` + +Type `int`, default `5`. + +Files per node passed natively to the LLM; past the cap they are extracted to text or dropped, to bound context and cost. Re-uses SANDBOX_MAX_INPUT_BYTES per file. + +### `WORKFLOW_NODE_EXTRACT_MAX_FILES` + +Type `int`, default `5`. + +Documents per node extracted via the parsing worker. Each issues a separate blocking parse; past the cap they are skipped with a truncation note. + +### `WORKFLOW_NODE_EXTRACT_BUDGET_SECONDS` + +Type `int`, default `900`. + +Wall clock one node may spend on blocking parses, shared across all of them. Without it a node could serialize WORKFLOW_NODE_EXTRACT_MAX_FILES full windows on a web threadpool slot. + +### `WORKFLOW_RUN_STALE_SECONDS` + +Type `int`, default `3600`. + +A run row is pre-created as running; a disconnect or crash can strand it there. The beat reaper fails runs still running past this. Generous so a long run is never cut off. + +### `ARTIFACT_MAX_BYTES` + +Type `int`, default `52428800`. + +Cap on a single stored artifact version's bytes (0 disables). + +### `ARTIFACT_MAX_COUNT_PER_USER` + +Type `int`, default `5000`. + +Cap on artifacts a user may own (0 disables). + +### `ARTIFACT_MAX_TOTAL_BYTES_PER_USER` + +Type `int`, default `5368709120`. + +Cap on a user's total stored artifact bytes (0 disables). + + +## Guardrails + +Input/output checks every agent runs, and the floor no agent may weaken. + +### `GUARDRAILS_ENABLED` + +Type `bool`, default `true`. + +Master switch; False disables every stage. + +### `GUARDRAILS_CHECKS_ENABLED` + +Type `list[str]`, default `[]`. + +Allowlist of GuardrailCreator.checks keys; empty means every registered check. + +### `GUARDRAILS_FLOOR` + +Type `dict[str, Any]`, default `{}`. + +A GuardrailsConfig fragment every agent inherits and cannot weaken; agents may add controls or make an action stricter, never looser. "enabled" is required; without it the floor parses but applies to nothing. Example: \{"enabled": true, "mode": "scan_all", "controls": [\{"check": "secrets", "stage": "output", "action": "redact"\}]\} + +### `GUARDRAILS_JUDGE_MODEL` + +Type `str`, default unset. + +Judge model for the topic/policy checks; unset reuses the request's model. + +### `GUARDRAILS_STORE_SCANNED_TEXT` + +Type `bool`, default `false`. + +Persist scanned text alongside guardrail_events. Off by default: pre-redaction text is exactly the material a PII control exists to keep out of storage. + +### `GUARDRAILS_EVENTS_RETENTION_DAYS` + +Type `int`, default `30`, must be `>= 1`. + +Days guardrail events are kept before the cleanup task removes them. + + +## Scheduler + +Cadence, quotas and timeouts of scheduled runs. + +### `SCHEDULE_DISPATCHER_INTERVAL` + +Type `int`, default `30`. + +Seconds between dispatcher passes that enqueue due schedules. + +### `SCHEDULE_MIN_INTERVAL` + +Type `int`, default `900`. + +Smallest allowed recurrence interval in seconds. + +### `SCHEDULE_MAX_PER_USER` + +Type `int`, default `50`. + +Cap on schedules a user may own. + +### `SCHEDULE_RUN_TIMEOUT` + +Type `int`, default `600`. + +Wall-clock cap on one scheduled run, in seconds. + +### `SCHEDULE_MISFIRE_GRACE` + +Type `int`, default `60`. + +Seconds past the due time within which a missed run still fires. + +### `SCHEDULE_AUTOPAUSE_FAILURES` + +Type `int`, default `3`. + +Consecutive failures after which a schedule is paused automatically. + +### `SCHEDULE_ONCE_MAX_HORIZON` + +Type `int`, default `31536000`. + +How far ahead a one-off run may be scheduled, in seconds (one year). + +### `SCHEDULE_RUN_OUTPUT_RETENTION_DAYS` + +Type `int`, default `90`, must be `> 0`. + +Days scheduled-run output is kept. + + +## Sandbox + +The app is a CLIENT of an always-on runner; defaults are safe so app import never fails unconfigured. + +### `SANDBOX_BACKEND` + +Type `"jupyter" | "daytona"`, default `jupyter`. + +Sandbox backend: jupyter (self-host) or daytona (Daytona Cloud). + +### `SANDBOX_GATEWAY_URL` + +Type `str`, default `http://localhost:8888`. + +URL of the Jupyter Kernel Gateway runner (the docsgpt-sandbox service). + +### `SANDBOX_GATEWAY_AUTH_TOKEN` + +Type `str`, default unset. + +Gateway auth token, if set. + +### `SANDBOX_KERNEL_NAME` + +Type `str`, default `docsgpt-python`. + +Kernelspec per session. The env-scrubbing docsgpt-python spec keeps kernel code from reading the gateway token or operator secrets from os.environ; the stock python3 spec inherits the gateway env verbatim and must not be used with untrusted code. + +### `SANDBOX_MAX_TTL` + +Type `int`, default `1200`. + +Hard cap (s) on agent-selectable keep-alive TTL. + +### `SANDBOX_MAX_SESSIONS` + +Type `int`, default `32`. + +Concurrent live sessions per process, backend-agnostic; at the cap an LRU-idle session is evicted. 0 or negative disables the cap. + +### `SANDBOX_EXEC_TIMEOUT` + +Type `int`, default `60`. + +Default wall-clock cap (s) per exec call. + +### `SANDBOX_HTTP_TIMEOUT` + +Type `int`, default `10`. + +Fixed cap (s) for REST control calls (create/delete/alive/interrupt). + +### `SANDBOX_MAX_OUTPUT_BYTES` + +Type `int`, default `8388608`. + +Cap on buffered stdout+stderr per exec. + +### `SANDBOX_MAX_FILE_BYTES` + +Type `int`, default `10485760`. + +Cap on get_file size routed through stdout. + +### `SANDBOX_MAX_INPUT_BYTES` + +Type `int`, default `26214400`. + +Cap on an input document staged into a sandbox session. + +### `SANDBOX_MEMORY` + +Type `str`, default `1g`. + +Docker mem_limit for the runner container. Consumed by the docsgpt-sandbox compose service, not the app; part of the untrusted-code security boundary. + +### `SANDBOX_CPUS` + +Type `str`, default `1.0`. + +Docker CPU quota for the runner container. Consumed by the docsgpt-sandbox compose service, not the app; part of the untrusted-code security boundary. + +### `DAYTONA_API_KEY` + +Type `str`, default unset. + +Daytona Cloud API key (secret). + +### `DAYTONA_API_URL` + +Type `str`, default unset. + +Override for the Daytona API base URL, if self-targeting. + +### `DAYTONA_TARGET` + +Type `str`, default unset. + +Daytona region/target, e.g. "us". + +### `DAYTONA_SNAPSHOT` + +Type `str`, default unset. + +Image for new sandboxes; render libs via scripts/build_daytona_snapshot.py. + +### `DAYTONA_LANGUAGE` + +Type `str`, default `python`. + +Default runtime language for created sandboxes. + +### `DAYTONA_AUTO_STOP_INTERVAL` + +Type `int`, default `15`, must be `>= 0`. + +Minutes idle before Daytona auto-stops a sandbox (0 disables). + +### `DAYTONA_AUTO_DELETE_INTERVAL` + +Type `int`, default `60`, must be `>= -1`. + +Minutes after stop before Daytona auto-deletes a sandbox (-1 disables). + +### `DAYTONA_MAX_SANDBOXES` + +Type `int`, default `50`. + +Cap on concurrent live Daytona sandboxes (cost-DoS guard). + + +## Speech + +Voice providers and transcription options. + +### `TTS_PROVIDER` + +Type `"google_tts" | "elevenlabs" | "none"`, default `google_tts`. + +Text-to-speech provider; none switches it off. + +### `ELEVENLABS_API_KEY` + +Type `str`, default unset. + +ElevenLabs API key. + +### `STT_PROVIDER` + +Type `"openai" | "faster_whisper" | "none"`, default `openai`. + +Speech-to-text provider; none switches it off. + +### `OPENAI_STT_MODEL` + +Type `str`, default `gpt-4o-mini-transcribe`. + +OpenAI transcription model. + +### `STT_LANGUAGE` + +Type `str`, default unset. + +Language hint for transcription; unset auto-detects. + +### `STT_MAX_FILE_SIZE_MB` + +Type `int`, default `50`. + +Cap on an audio file accepted for transcription. + +### `STT_ENABLE_TIMESTAMPS` + +Type `bool`, default `false`. + +Return word/segment timestamps. + +### `STT_ENABLE_DIARIZATION` + +Type `bool`, default `false`. + +Label speakers in the transcript. diff --git a/docs/content/Deploying/_meta.js b/docs/content/Deploying/_meta.js index 4865a40e..105b9efd 100644 --- a/docs/content/Deploying/_meta.js +++ b/docs/content/Deploying/_meta.js @@ -3,6 +3,10 @@ export default { "title": "⚙️ App Configuration", "href": "/Deploying/DocsGPT-Settings" }, + "Settings-Reference": { + "title": "📖 Settings Reference", + "href": "/Deploying/Settings-Reference" + }, "OIDC-SSO": { "title": "🔐 SSO with OIDC", "href": "/Deploying/OIDC-SSO" diff --git a/docs/content/Guides/Architecture.mdx b/docs/content/Guides/Architecture.mdx index a65f3dcb..a25d34e6 100644 --- a/docs/content/Guides/Architecture.mdx +++ b/docs/content/Guides/Architecture.mdx @@ -251,4 +251,4 @@ The main extension points in [arc53/DocsGPT](https://github.com/arc53/DocsGPT) a | Parsing and workers | [`docsgpt/parser/`](https://github.com/arc53/DocsGPT/tree/main/docsgpt/parser), [`docsgpt/worker.py`](https://github.com/arc53/DocsGPT/blob/main/docsgpt/worker.py), [`docsgpt/api/user/tasks.py`](https://github.com/arc53/DocsGPT/blob/main/docsgpt/api/user/tasks.py) | | Models and vector stores | [`docsgpt/llm/`](https://github.com/arc53/DocsGPT/tree/main/docsgpt/llm), [`docsgpt/vectorstore/`](https://github.com/arc53/DocsGPT/tree/main/docsgpt/vectorstore) | | Storage and events | [`docsgpt/storage/`](https://github.com/arc53/DocsGPT/tree/main/docsgpt/storage), [`docsgpt/streaming/`](https://github.com/arc53/DocsGPT/tree/main/docsgpt/streaming), [`docsgpt/events/`](https://github.com/arc53/DocsGPT/tree/main/docsgpt/events) | -| Configuration, UI, and deployment | [`docsgpt/core/settings.py`](https://github.com/arc53/DocsGPT/blob/main/docsgpt/core/settings.py), [`frontend/`](https://github.com/arc53/DocsGPT/tree/main/frontend), [`deployment/`](https://github.com/arc53/DocsGPT/tree/main/deployment) | +| Configuration, UI, and deployment | [`docsgpt/core/settings/`](https://github.com/arc53/DocsGPT/tree/main/docsgpt/core/settings), [`frontend/`](https://github.com/arc53/DocsGPT/tree/main/frontend), [`deployment/`](https://github.com/arc53/DocsGPT/tree/main/deployment) | diff --git a/docs/content/Guides/compression.md b/docs/content/Guides/compression.md index 14b90c62..53bcea57 100644 --- a/docs/content/Guides/compression.md +++ b/docs/content/Guides/compression.md @@ -19,7 +19,7 @@ The compression system operates on a "summarize and truncate" principle: ## Configuration -You can configure the compression behavior in your `.env` file or `docsgpt/core/settings.py`: +You can configure the compression behavior in your `.env` file or `docsgpt/core/settings/agents.py`: | Setting | Default | Description | | :--- | :--- | :--- | diff --git a/docs/content/Sources/Per-source-configuration.mdx b/docs/content/Sources/Per-source-configuration.mdx index 4c23a710..a66ea285 100644 --- a/docs/content/Sources/Per-source-configuration.mdx +++ b/docs/content/Sources/Per-source-configuration.mdx @@ -115,8 +115,6 @@ Requests are also bounded: `chunks` is clamped to 0–500, and `0` still means Keyword search for the **hybrid** retriever is currently implemented only for the **pgvector** vector store. On other stores (FAISS, Qdrant, Milvus, etc.) the keyword half returns nothing, so `hybrid` quietly behaves like `classic` (vector-only). -Operators can restrict which retrievers are usable instance-wide with the `RETRIEVERS_ENABLED` setting; a per-source `retriever` value must be within that allow-list. - ### Exposure: prefetch vs. agentic tool `exposure` controls *how* a source's content is delivered to the model: diff --git a/docs/content/quickstart.mdx b/docs/content/quickstart.mdx index 3b134fe6..a1702b3e 100644 --- a/docs/content/quickstart.mdx +++ b/docs/content/quickstart.mdx @@ -124,6 +124,6 @@ To work from the source tree, for example to build the images yourself, use `set ## Advanced Configuration -For more advanced customization of DocsGPT settings, such as configuring vector stores, embedding models, and other parameters, please refer to the [DocsGPT Settings documentation](/Deploying/DocsGPT-Settings). This guide explains how to modify the `.env` file or `settings.py` for deeper configuration. +For more advanced customization of DocsGPT settings, such as configuring vector stores, embedding models, and other parameters, please refer to the [DocsGPT Settings documentation](/Deploying/DocsGPT-Settings). This guide explains how to configure DocsGPT through the `.env` file, and links to the full settings reference. Enjoy using DocsGPT! diff --git a/docs/runbooks/sse-notifications.md b/docs/runbooks/sse-notifications.md index 1d100c43..6afaf225 100644 --- a/docs/runbooks/sse-notifications.md +++ b/docs/runbooks/sse-notifications.md @@ -320,7 +320,7 @@ redis-cli -n 2 DEL user::stream ## Settings reference -Everything in `docsgpt/core/settings.py`: +Everything in `docsgpt/core/settings/events.py`: | Setting | Default | Purpose | | --------------------------------------------- | ------- | --------------------------------------------- | diff --git a/docsgpt/agents/base.py b/docsgpt/agents/base.py index 71caaa34..d731a415 100644 --- a/docsgpt/agents/base.py +++ b/docsgpt/agents/base.py @@ -382,7 +382,7 @@ class BaseAgent(ABC): when the conversation was compressed after that turn was produced — the compressed local history is the context then, not the server's. """ - if not getattr(settings, "OPENAI_RESPONSES_CHAIN_ACROSS_TURNS", True): + if not settings.OPENAI_RESPONSES_CHAIN_ACROSS_TURNS: return None if not self.chat_history: return None @@ -422,7 +422,7 @@ class BaseAgent(ABC): # No provider-reported usage on the previous turn (older rows, # estimate-only providers): nothing to bound against. return meta["response_id"] - budget = getattr(settings, "OPENAI_RESPONSES_CHAIN_BUDGET_TOKENS", None) + budget = settings.OPENAI_RESPONSES_CHAIN_BUDGET_TOKENS if not budget: from docsgpt.core.model_utils import get_token_limit diff --git a/docsgpt/agents/tool_executor.py b/docsgpt/agents/tool_executor.py index edec516f..9952e3d2 100644 --- a/docsgpt/agents/tool_executor.py +++ b/docsgpt/agents/tool_executor.py @@ -53,7 +53,7 @@ def _dedupable_tool_names() -> frozenset: """ from docsgpt.core.settings import settings - return frozenset(BUILTIN_AGENT_TOOLS) | frozenset(getattr(settings, "DEFAULT_CHAT_TOOLS", None) or []) + return frozenset(BUILTIN_AGENT_TOOLS) | frozenset(settings.DEFAULT_CHAT_TOOLS or []) def _requires_approval(tool: Dict, action: Dict) -> bool: diff --git a/docsgpt/agents/tools/artifact_generator.py b/docsgpt/agents/tools/artifact_generator.py index 1f3c7007..5c2da920 100644 --- a/docsgpt/agents/tools/artifact_generator.py +++ b/docsgpt/agents/tools/artifact_generator.py @@ -829,7 +829,7 @@ class ArtifactGeneratorTool(Tool): spec_path = f"{token_dir}/spec.json" out_path = f"{token_dir}/out.{_KIND_INFO[kind]['ext']}" program = _RENDERERS[kind].format(spec_path=spec_path, out_path=out_path) - timeout = float(getattr(settings, "SANDBOX_EXEC_TIMEOUT", 60)) + timeout = float(settings.SANDBOX_EXEC_TIMEOUT) manager = SandboxCreator.get_manager() try: diff --git a/docsgpt/agents/tools/attachment_bridge.py b/docsgpt/agents/tools/attachment_bridge.py index 7e1400d1..23c56bfb 100644 --- a/docsgpt/agents/tools/attachment_bridge.py +++ b/docsgpt/agents/tools/attachment_bridge.py @@ -113,7 +113,7 @@ def bridge_attachment( # Reject oversize attachments BEFORE buffering them: the authoritative ``size`` # column lets us avoid pulling a multi-hundred-MB file fully into worker memory, # and the bounded read below backstops a missing/lying ``size``. - max_bytes = int(getattr(settings, "ARTIFACT_MAX_BYTES", 0) or 0) + max_bytes = int(settings.ARTIFACT_MAX_BYTES or 0) declared_size = attachment.get("size") if max_bytes and isinstance(declared_size, (int, float)) and declared_size > max_bytes: raise AttachmentBridgeError( diff --git a/docsgpt/agents/tools/code_executor.py b/docsgpt/agents/tools/code_executor.py index 079d4a78..697411a7 100644 --- a/docsgpt/agents/tools/code_executor.py +++ b/docsgpt/agents/tools/code_executor.py @@ -87,9 +87,9 @@ class CodeExecutorTool(Tool): baked in. Keep the package lists in sync with deployment/sandbox/Dockerfile (jupyter) and scripts/build_daytona_snapshot.py (daytona snapshot). """ - backend = str(getattr(settings, "SANDBOX_BACKEND", "jupyter") or "jupyter").lower() + backend = str(settings.SANDBOX_BACKEND or "jupyter").lower() if backend == "daytona": - if getattr(settings, "DAYTONA_SNAPSHOT", None): + if settings.DAYTONA_SNAPSHOT: return ( "Preinstalled beyond the stdlib: python-pptx, python-docx, openpyxl, " "reportlab, lxml, pillow. pip install anything else from within the code " @@ -330,7 +330,7 @@ class CodeExecutorTool(Tool): # Reject an oversize input BEFORE buffering it: the declared ``size`` # avoids pulling a huge file into worker memory, and the bounded read # below backstops a missing/lying size column. - max_bytes = int(getattr(settings, "SANDBOX_MAX_INPUT_BYTES", 0) or 0) + max_bytes = int(settings.SANDBOX_MAX_INPUT_BYTES or 0) declared_size = version.get("size") if max_bytes and isinstance(declared_size, (int, float)) and declared_size > max_bytes: return {"error": f"input artifact {artifact_id} exceeds the {max_bytes}-byte sandbox input limit."} @@ -484,7 +484,7 @@ class CodeExecutorTool(Tool): @staticmethod def _exec_timeout() -> float: """Return the fixed per-run wall-clock cap (SANDBOX_EXEC_TIMEOUT; not caller-adjustable).""" - return float(getattr(settings, "SANDBOX_EXEC_TIMEOUT", 60)) + return float(settings.SANDBOX_EXEC_TIMEOUT) @staticmethod def _is_timeout(result: ExecResult) -> bool: diff --git a/docsgpt/agents/tools/mcp_tool.py b/docsgpt/agents/tools/mcp_tool.py index e8b4112b..061f4c6d 100644 --- a/docsgpt/agents/tools/mcp_tool.py +++ b/docsgpt/agents/tools/mcp_tool.py @@ -108,11 +108,11 @@ class MCPTool(Tool): if configured_redirect_uri: return configured_redirect_uri.rstrip("/") - explicit = getattr(settings, "MCP_OAUTH_REDIRECT_URI", None) + explicit = settings.MCP_OAUTH_REDIRECT_URI if explicit: return explicit.rstrip("/") - connector_base = getattr(settings, "CONNECTOR_REDIRECT_BASE_URI", None) + connector_base = settings.CONNECTOR_REDIRECT_BASE_URI if connector_base: parsed = urlparse(connector_base) if parsed.scheme and parsed.netloc: diff --git a/docsgpt/agents/tools/read_document.py b/docsgpt/agents/tools/read_document.py index 6e488937..ab9d2294 100644 --- a/docsgpt/agents/tools/read_document.py +++ b/docsgpt/agents/tools/read_document.py @@ -255,7 +255,7 @@ class ReadDocumentTool(Tool): # The task's per-call time limits are raised to match the awaited window: bound to # the base timeout at import, the worker would otherwise self-terminate a large # parse long before this await gives up. - queue = getattr(settings, "DOCUMENT_PARSE_QUEUE", "parsing") + queue = settings.DOCUMENT_PARSE_QUEUE try: async_result = parse_document.apply_async( args=[artifact_id, parent, self.user_id, options], diff --git a/docsgpt/agents/workflow_agent.py b/docsgpt/agents/workflow_agent.py index 6c1f71ff..2c08bc0c 100644 --- a/docsgpt/agents/workflow_agent.py +++ b/docsgpt/agents/workflow_agent.py @@ -331,7 +331,7 @@ class WorkflowAgent(BaseAgent): from docsgpt.storage.storage_creator import StorageCreator storage = StorageCreator.get_storage() - max_bytes = int(getattr(settings, "ARTIFACT_MAX_BYTES", 0) or 0) + max_bytes = int(settings.ARTIFACT_MAX_BYTES or 0) dropped: List[str] = [] if len(self.attachments) > _MAX_INPUT_DOCUMENTS: over = len(self.attachments) - _MAX_INPUT_DOCUMENTS diff --git a/docsgpt/agents/workflows/workflow_engine.py b/docsgpt/agents/workflows/workflow_engine.py index b1189511..54fa6e67 100644 --- a/docsgpt/agents/workflows/workflow_engine.py +++ b/docsgpt/agents/workflows/workflow_engine.py @@ -639,7 +639,7 @@ class WorkflowEngine: raw_ids = self._resolve_input_artifact_ids(inputs) if not raw_ids: return loaded - max_bytes = int(getattr(settings, "SANDBOX_MAX_INPUT_BYTES", 0) or 0) + max_bytes = int(settings.SANDBOX_MAX_INPUT_BYTES or 0) storage = StorageCreator.get_storage() # Two inputs whose current versions share a filename would clobber each other at the # same ``inputs/{name}`` path; track used paths and disambiguate deterministically. @@ -749,15 +749,15 @@ class WorkflowEngine: supported = set(supported_types) supports_images = any(t.startswith("image/") for t in supported) - max_files = int(getattr(settings, "WORKFLOW_NODE_NATIVE_MAX_FILES", 5)) - extract_max = int(getattr(settings, "WORKFLOW_NODE_EXTRACT_MAX_FILES", 5)) + max_files = int(settings.WORKFLOW_NODE_NATIVE_MAX_FILES) + extract_max = int(settings.WORKFLOW_NODE_EXTRACT_MAX_FILES) # One wall clock for every blocking parse this node issues. The cap # above bounds how MANY parses run; this bounds how LONG they take in # total, so N documents cannot serialize N size-scaled windows. parse_deadline = time.monotonic() + float( - getattr(settings, "WORKFLOW_NODE_EXTRACT_BUDGET_SECONDS", 900) + settings.WORKFLOW_NODE_EXTRACT_BUDGET_SECONDS ) - max_bytes = int(getattr(settings, "SANDBOX_MAX_INPUT_BYTES", 25 * 1024 * 1024)) + max_bytes = int(settings.SANDBOX_MAX_INPUT_BYTES) # One read-only connection for the whole batch; the resolved-version # rows are collected, then storage reads happen outside the DB context. @@ -976,7 +976,7 @@ class WorkflowEngine: if not user_id: return None options = {"output": "markdown", "include_tables": False, "persist": False} - queue = getattr(settings, "DOCUMENT_PARSE_QUEUE", "parsing") + queue = settings.DOCUMENT_PARSE_QUEUE # OCR cost scales with pages, so the window grows with the document's size # (floored at DOCUMENT_PARSE_TIMEOUT); the task's per-call time limits are # raised to match, else the worker would self-terminate mid-parse. @@ -1084,7 +1084,7 @@ class WorkflowEngine: """Return the stricter of the node's requested timeout and the sandbox cap.""" from docsgpt.core.settings import settings - cap = float(getattr(settings, "SANDBOX_EXEC_TIMEOUT", 60)) + cap = float(settings.SANDBOX_EXEC_TIMEOUT) if requested is None: return cap try: diff --git a/docsgpt/api/answer/services/compression/service.py b/docsgpt/api/answer/services/compression/service.py index eb47f5d0..6b6b14a5 100644 --- a/docsgpt/api/answer/services/compression/service.py +++ b/docsgpt/api/answer/services/compression/service.py @@ -367,7 +367,7 @@ class CompressionService: never mutated. """ max_tokens = int( - getattr(settings, "COMPRESSION_RECENT_FIELD_MAX_TOKENS", 8000) or 0 + settings.COMPRESSION_RECENT_FIELD_MAX_TOKENS or 0 ) if max_tokens <= 0: return queries diff --git a/docsgpt/api/async_sse.py b/docsgpt/api/async_sse.py index d3fa3437..a5064e2c 100644 --- a/docsgpt/api/async_sse.py +++ b/docsgpt/api/async_sse.py @@ -33,7 +33,6 @@ from docsgpt.streaming.async_event_replay import ( ) from docsgpt.streaming.async_redis import get_async_redis_instance from docsgpt.streaming.event_replay import ( - DEFAULT_KEEPALIVE_SECONDS, DEFAULT_POLL_TIMEOUT_SECONDS, ) from docsgpt.streaming.sse_leases import StreamCapExceeded, acquire_stream_lease @@ -127,7 +126,7 @@ async def stream_message_events(request: Request) -> Response: ) last_event_id = _normalise_last_event_id(raw_cursor) keepalive_seconds = float( - getattr(settings, "SSE_KEEPALIVE_SECONDS", DEFAULT_KEEPALIVE_SECONDS) + settings.SSE_KEEPALIVE_SECONDS ) logger.info( diff --git a/docsgpt/api/user/artifacts/download.py b/docsgpt/api/user/artifacts/download.py index a57ae727..3e23895d 100644 --- a/docsgpt/api/user/artifacts/download.py +++ b/docsgpt/api/user/artifacts/download.py @@ -154,7 +154,7 @@ async def download_artifact(request: Request) -> Response: # URL. If the active backend can't mint one, that's a config error: # surface a 500 rather than silently proxying bytes from a backend # the operator expected to be off the hot path. - if getattr(settings, "URL_STRATEGY", "backend") == "s3": + if settings.URL_STRATEGY == "s3": try: url = await anyio.to_thread.run_sync( partial(storage.generate_presigned_url, storage_path, expires_in=_PRESIGNED_URL_TTL) diff --git a/docsgpt/api/user/base.py b/docsgpt/api/user/base.py index f6cd6f58..9579a73f 100644 --- a/docsgpt/api/user/base.py +++ b/docsgpt/api/user/base.py @@ -235,7 +235,7 @@ def get_vector_store(source_id): store = VectorCreator.create_vectorstore( settings.VECTOR_STORE, source_id=source_id, - embeddings_key=os.getenv("EMBEDDINGS_KEY"), + embeddings_key=settings.EMBEDDINGS_KEY, ) return store diff --git a/docsgpt/api/user/tasks.py b/docsgpt/api/user/tasks.py index d4307b33..82d247e7 100644 --- a/docsgpt/api/user/tasks.py +++ b/docsgpt/api/user/tasks.py @@ -346,9 +346,9 @@ def parse_timeout_for_size(size_bytes: Optional[int]) -> float: """ from docsgpt.core.settings import settings - base = float(getattr(settings, "DOCUMENT_PARSE_TIMEOUT", 120) or 120) - per_mib = float(getattr(settings, "DOCUMENT_PARSE_TIMEOUT_PER_MB", 0) or 0) - ceiling = float(getattr(settings, "DOCUMENT_PARSE_TIMEOUT_MAX", base) or base) + base = float(settings.DOCUMENT_PARSE_TIMEOUT or 120) + per_mib = float(settings.DOCUMENT_PARSE_TIMEOUT_PER_MB or 0) + ceiling = float(settings.DOCUMENT_PARSE_TIMEOUT_MAX or base) size = float(size_bytes) if isinstance(size_bytes, (int, float)) else 0.0 scaled = base + per_mib * max(size, 0.0) / (1024 * 1024) return min(ceiling, max(base, scaled)) diff --git a/docsgpt/app.py b/docsgpt/app.py index 0913218f..4a11937e 100644 --- a/docsgpt/app.py +++ b/docsgpt/app.py @@ -1,5 +1,4 @@ import logging -import os import platform import uuid @@ -170,16 +169,8 @@ def enforce_document_upload_request_size_limit(): # only local development may use the atomic filesystem fallback. settings.JWT_SECRET_KEY = resolve_jwt_secret_key( settings.JWT_SECRET_KEY, - os.getenv("DEPLOYMENT_TYPE"), + settings.DEPLOYMENT_TYPE, ) -if settings.AUTH_TYPE == "oidc": - _missing_oidc = [ - name - for name in ("OIDC_ISSUER", "OIDC_CLIENT_ID", "OIDC_FRONTEND_URL") - if not getattr(settings, name) - ] - if _missing_oidc: - raise RuntimeError(f"AUTH_TYPE=oidc requires settings: {', '.join(_missing_oidc)}") SIMPLE_JWT_TOKEN = None if settings.AUTH_TYPE == "simple_jwt": payload = {"sub": "local"} diff --git a/docsgpt/core/db_uri.py b/docsgpt/core/db_uri.py index 99e93bc3..cb875b4e 100644 --- a/docsgpt/core/db_uri.py +++ b/docsgpt/core/db_uri.py @@ -15,7 +15,7 @@ have to know which driver a given field feeds. Each normalizer also silently upgrades the legacy ``postgresql+psycopg2://`` prefix since psycopg2 is no longer in the project. -This module is deliberately separate from ``docsgpt/core/settings.py`` +This module is deliberately separate from ``docsgpt/core/settings`` so the Settings class stays focused on field declarations, and the URI-rewriting logic can be unit-tested without triggering ``.env`` file loading from importing Settings. diff --git a/docsgpt/core/model_registry.py b/docsgpt/core/model_registry.py index 07a30a02..cc71e206 100644 --- a/docsgpt/core/model_registry.py +++ b/docsgpt/core/model_registry.py @@ -140,7 +140,7 @@ class ModelRegistry: from docsgpt.llm.providers import ALL_PROVIDERS directories = [BUILTIN_MODELS_DIR] - operator_dir = getattr(settings, "MODELS_CONFIG_DIR", None) + operator_dir = settings.MODELS_CONFIG_DIR if operator_dir: op_path = Path(operator_dir) if not op_path.exists(): diff --git a/docsgpt/core/settings.py b/docsgpt/core/settings.py deleted file mode 100644 index 8bf864d0..00000000 --- a/docsgpt/core/settings.py +++ /dev/null @@ -1,596 +0,0 @@ -import os -from typing import Optional - -from pydantic import AliasChoices, Field, field_validator -from pydantic_settings import BaseSettings, SettingsConfigDict - -from docsgpt.core.db_uri import ( - normalize_pgvector_connection_string, - normalize_postgres_uri, -) -from docsgpt.core.paths import env_file, home_dir - -# Runtime data home (DOCSGPT_HOME, the checkout, or cwd); see docsgpt.core.paths. -current_dir = str(home_dir()) - - -class Settings(BaseSettings): - model_config = SettingsConfigDict(extra="ignore") - - AUTH_TYPE: Optional[str] = None # simple_jwt, session_jwt, oidc, or None - - # OIDC SSO (AUTH_TYPE=oidc) — any OpenID Connect IdP with discovery (Authentik, Keycloak, ...) - OIDC_ISSUER: Optional[str] = None # e.g. https://auth.example.com/application/o/docsgpt/ - OIDC_CLIENT_ID: Optional[str] = None - OIDC_CLIENT_SECRET: Optional[str] = None # optional; PKCE is always used - OIDC_SCOPES: str = "openid profile email" - OIDC_USER_ID_CLAIM: str = "sub" # ID-token claim mapped to the DocsGPT user id - OIDC_FRONTEND_URL: Optional[str] = None # browser-facing app origin, e.g. http://localhost:5173 - OIDC_REDIRECT_URI: Optional[str] = None # override; default /api/auth/oidc/callback - OIDC_SESSION_LIFETIME_SECONDS: int = 28800 # minted session JWT lifetime (8h) - OIDC_PROVIDER_NAME: Optional[str] = None # sign-in button label, e.g. "Acme SSO" - OIDC_ALLOWED_GROUPS: Optional[str] = None # comma-separated allowlist; unset = any authenticated user - OIDC_GROUPS_CLAIM: str = "groups" # ID-token/userinfo claim carrying group membership - OIDC_ADMIN_GROUPS: Optional[str] = None # comma-separated groups granted admin; unset = no OIDC admin mapping - - # RBAC: persisted admin grants live in user_roles (AUTH_TYPE=oidc only). This is the - # only non-DB admin path, for AUTH_TYPE=None self-host. MUST stay False if networked. - LOCAL_MODE_ADMIN: bool = False - - # SCIM 2.0 provisioning (IdP-driven user create/deactivate at /scim/v2) - SCIM_ENABLED: bool = False - SCIM_TOKEN: Optional[str] = None # bearer token for IdP SCIM clients (required when enabled) - - LLM_PROVIDER: str = "docsgpt" - LLM_NAME: Optional[str] = None # if LLM_PROVIDER is openai, LLM_NAME can be gpt-4 or gpt-3.5-turbo - # Legacy model on purpose: an install that never pinned this has vectors from it, and - # granite is the same width so a swap would fail silently. New installs get granite from - # .env-template; existing ones switch by setting this and running docsgpt.scripts.reembed. - EMBEDDINGS_NAME: str = "huggingface_sentence-transformers/all-mpnet-base-v2" - EMBEDDINGS_BASE_URL: Optional[str] = None # Remote embeddings API URL (OpenAI-compatible) - EMBEDDINGS_KEY: Optional[str] = None # api key for embeddings (if using openai, just copy API_KEY) - EMBEDDINGS_MAX_INPUT_TOKENS: Optional[int] = None # truncate each remote embed input to N tokens (overflow lost) - EMBEDDINGS_BATCH_SIZE: int = 32 # chunks per store transaction / remote embed request - # Documents per local ONNX forward pass. Each pass pads to its longest input, and that - # waste grows with the square of chunk length: at 1250 tokens, 32 peaked at 6.6 GB, 1 at 2.9 GB. - EMBEDDINGS_MODEL_BATCH_SIZE: int = 1 - # Intra-op threads for the local ONNX runner; None = every core. It scales sub-linearly, - # so several single-threaded workers beat one many-threaded process on the same cores. - EMBEDDINGS_THREADS: Optional[int] = None - # Embedding models and their tokenizers. Persistent by default: FastEmbed's own default is the temp dir. - EMBEDDINGS_CACHE_DIR: Optional[str] = Field(default_factory=lambda: str(home_dir() / "models")) - # Pooling ("cls"/"mean") and L2 normalisation. Read from the model's own repository; - # set these only for a repository that declares neither, or to override what it declares. - EMBEDDINGS_POOLING: Optional[str] = None - EMBEDDINGS_NORMALIZE: Optional[bool] = None - # Embed on the worker so the API holds no model (~890 MB), at one broker round trip per - # query. Ignored when EMBEDDINGS_BASE_URL is set, which is the better answer for production. - EMBEDDINGS_DELEGATE_TO_WORKER: bool = True - EMBEDDINGS_QUEUE: str = "embeddings" # queue the embed task is routed to - EMBEDDINGS_DELEGATE_TIMEOUT: int = 60 # seconds to wait for the worker - GITHUB_INGEST_MAX_FILE_BYTES: int = 1048576 # skip repo blobs larger than this (0 = no cap) - GITHUB_INGEST_MAX_WORKERS: int = 8 # parallel file fetches per GitHub repo ingest - # Operator-supplied model YAMLs, loaded after the built-in catalog; later wins on - # duplicate model id. See docsgpt/core/models/README.md. - MODELS_CONFIG_DIR: Optional[str] = None - - CELERY_BROKER_URL: str = "redis://localhost:6379/0" - CELERY_RESULT_BACKEND: str = "redis://localhost:6379/1" - # Prefetch=1 caps SIGKILL loss to one task. Visibility timeout must exceed the longest - # legitimate task runtime but stay short enough that SIGKILLed tasks redeliver promptly. - CELERY_WORKER_PREFETCH_MULTIPLIER: int = 1 - CELERY_VISIBILITY_TIMEOUT: int = 3600 - # Recycle a prefork child past this resident size in KB; backstops docling/torch heap growth. - # Checked between tasks, so it does not bound the peak within one. 0 disables. - CELERY_WORKER_MAX_MEMORY_PER_CHILD: int = 4194304 - CELERY_WORKER_MAX_TASKS_PER_CHILD: int = 0 # recycle after N tasks; 0 disables - # Only consulted when VECTOR_STORE=mongodb or when running scripts/db/backfill.py; user data lives in Postgres. - MONGO_URI: Optional[str] = None - # User-data Postgres DB. - POSTGRES_URI: Optional[str] = None - # On startup, apply pending Alembic migrations. Disable if you manage schema out-of-band. - AUTO_MIGRATE: bool = True - # On startup, create the target Postgres database if missing (needs CREATEDB privilege). - AUTO_CREATE_DB: bool = True - # On startup, create the pgvector/graph tables and verify the embedding dimension. No Alembic - # migration covers the vector DB (it may be a separate cluster); set False to manage it yourself. - AUTO_VECTOR_SCHEMA: bool = True - LLM_PATH: str = os.path.join(current_dir, "models/docsgpt-7b-f16.gguf") - DEFAULT_MAX_HISTORY: int = 150 - DEFAULT_LLM_TOKEN_LIMIT: int = 128000 # Fallback when model not found in registry - RESERVED_TOKENS: dict = { - "system_prompt": 500, - "current_query": 500, - "safety_buffer": 1000, - } - DEFAULT_AGENT_LIMITS: dict = { - "token_limit": 50000, - "request_limit": 500, - } - UPLOAD_FOLDER: str = "inputs" - # Serve the web UI shipped in the package (docsgpt/static) from the API process. - SERVE_UI: bool = True - # Request cap is applied by Flask before multipart parsing; the per-file cap also while copying. - UPLOAD_MAX_REQUEST_BYTES: int = Field(default=256 * 1024 * 1024, gt=0) - UPLOAD_MAX_FILE_BYTES: int = Field(default=100 * 1024 * 1024, gt=0) - PARSE_SPEC_MAX_BYTES: int = Field(default=10 * 1024 * 1024, gt=0) - # ZIP limits apply cumulatively across nested archives in one extraction. - UPLOAD_MAX_ARCHIVE_BYTES: int = Field(default=250 * 1024 * 1024, gt=0) - UPLOAD_MAX_ARCHIVE_FILES: int = Field(default=10_000, gt=0) - UPLOAD_MAX_ARCHIVE_RATIO: int = Field(default=1000, gt=0) - UPLOAD_MAX_ARCHIVE_DEPTH: int = Field(default=3, ge=0) - PARSE_PDF_AS_IMAGE: bool = False - PARSE_IMAGE_REMOTE: bool = False - # Document parser for source ingestion, chat attachments and the - # read_document tool. "anydoc" (default): firecrawl-anydoc, a Rust - # converter with no ML models — milliseconds per file, ~100 MB peak RSS. - # "docling": the layout/table-model pipeline (optional install; needed - # for read_document's structured output and the docling OCR backend). - # Files anydoc cannot convert (scanned PDFs, malformed input) fall back to - # docling when it is installed, otherwise to the native OCR parsers (OCR - # on) or the legacy parsers. Rollback to the previous behaviour is this - # one variable. - DOC_PARSER_ENGINE: str = "anydoc" - # OCR for scanned PDFs and images. OCR_ENABLED covers source ingestion, - # OCR_ATTACHMENTS_ENABLED chat attachments. Which stack performs it is - # OCR_BACKEND; which engine, OCR_ENGINE. The DOCLING_OCR_* names are the - # pre-2026-09 spellings and stay accepted as aliases. - OCR_ENABLED: bool = Field( - default=False, validation_alias=AliasChoices("OCR_ENABLED", "DOCLING_OCR_ENABLED") - ) - OCR_ATTACHMENTS_ENABLED: bool = Field( - default=False, - validation_alias=AliasChoices("OCR_ATTACHMENTS_ENABLED", "DOCLING_OCR_ATTACHMENTS_ENABLED"), - ) - # Which stack runs OCR when it is on: - # auto — docling when installed, otherwise native. - # docling — the layout-model pipeline (hybrid region OCR, reading order, - # table structure); needs the optional docling extra. - # native — pypdfium2/Pillow page rendering straight into tesseract or a - # DeepSeek-OCR endpoint (docsgpt/parser/file/ocr_parser.py). - # No ML models in the worker; tables come out as text lines - # under tesseract. - OCR_BACKEND: str = "auto" - # Pages docling's threaded pipeline buffers in flight; the library - # default (100) drives worker RSS to ~3 GB on a mid-size PDF. - DOCLING_PIPELINE_QUEUE_MAX_SIZE: int = 2 - DOCLING_COMPILE_TORCH_MODELS: bool = False - DOCLING_TABULAR_MAX_BYTES: int = 2_000_000 - DOCLING_MARKUP_MAX_BYTES: int = 8_000_000 - # HTML/XHTML larger than this (bytes) are head-truncated before the - # markdownify parser runs (the anydoc engine's HTML path). The tree that - # path builds costs ~50x the input — 30 MB of HTML measured at 1.6 GB RSS — - # and the upload cap is 100 MB, so the gate is what keeps one upload from - # taking the ingest worker down. 0 disables it. - MARKUP_MAX_BYTES: int = 8_000_000 - # Trust-check anydoc's PDF output (docsgpt/parser/file/pdf_trust.py): - # flag composite (Type0) fonts without a ToUnicode map, and CJK-declaring - # PDFs whose extracted text has almost no CJK — the two classes where - # anydoc drops text silently. A flagged file re-parses on the docling - # fallback when docling is installed; otherwise the anydoc output is kept - # and the document gets extra_info["parse_warnings"]. ~30 ms per scanned MB. - PDF_TRUST_CHECK: bool = True - # Rewrite dot-leader / whitespace-aligned table runs in anydoc's PDF - # markdown into GFM tables (docsgpt/parser/file/tableize.py). Off by - # default: it rewrites content on a heuristic (>=3 uniform label+numbers - # lines) validated only on a small corpus so far. - ANYDOC_TABLEIZE: bool = False - # OCR engine used when OCR is on (OCR_ENABLED / OCR_ATTACHMENTS_ENABLED). - # Benched 2026-08 on EN/ZH/table/degraded scans (docs/Guides/ocr has the - # menu): - # tesseract — recommended: best classic-engine accuracy (perfect EN word - # recall, 0.000 bilingual CER, 100% table cells), ~35 MB, CPU-only. - # Needs the system binary + language packs: an optional install like - # every OCR dependency (build with INSTALL_TESSERACT=true, or apt/brew - # install tesseract-ocr for a local run). Both backends. - # deepseek — DeepSeek-OCR against an Ollama/vLLM endpoint - # (OCR_DEEPSEEK_*). Best table/CJK quality; the worker stays light - # (no layout models) but each page costs seconds on the model server. - # Both backends. - # auto — docling's pick: ocrmac on macOS (excellent), rapidocr on Linux - # (silently shreds some long text lines — avoid as a server default). - # ocrmac | rapidocr — force one of those. - # auto/ocrmac/rapidocr exist only inside docling; the native backend runs - # tesseract for them. An engine that is not installed degrades (docling: - # to "auto") with a warning instead of failing the parse. - OCR_ENGINE: str = "tesseract" - # Tesseract language packs, "+"-separated (e.g. "eng+chi_sim+deu"). Other - # engines keep their own defaults — their language codes differ. - OCR_LANGS: str = "eng" - OCR_DEEPSEEK_URL: str = "http://localhost:11434/v1/chat/completions" - OCR_DEEPSEEK_MODEL: str = "deepseek-ocr:3b" - # Seconds allowed per page request to the DeepSeek endpoint, on both - # backends (native sends pages one at a time; docling's VLM pipeline - # keeps its own concurrency). A 3B model on a laptop needs minutes; a - # vLLM GPU deployment, seconds. - OCR_DEEPSEEK_TIMEOUT: float = 300.0 - # Native backend only: resolution at which pages without a text layer are - # rendered before OCR. 200 suits tesseract; clamped to 72-600. - OCR_RENDER_DPI: int = 200 - # Chars-per-page floor below which an OCR'd PDF/image parse is treated as an OCR - # dropout (long-running docling workers were observed returning zero characters for - # every scanned page after a long scanned PDF, with no error) rather than as content. - # docling retries once on a fresh full-page-OCR converter; both backends then fail - # loudly instead of indexing an empty document. 0 disables the guard. - OCR_MIN_CHARS_PER_PAGE: int = Field( - default=20, validation_alias=AliasChoices("OCR_MIN_CHARS_PER_PAGE", "DOCLING_OCR_MIN_CHARS_PER_PAGE") - ) - # Read PDF *attachments* via their embedded text layer (pypdfium2) instead - # of docling, falling back to docling when there is no text layer to read. - # Attachments go into a prompt, so docling's structural markdown earns far - # less than the tens of seconds per file it costs; source ingestion is - # unaffected and always uses docling, because chunking and retrieval do - # depend on that structure. - ATTACHMENT_PDF_TEXT_FAST_PATH: bool = True - # Median chars per sampled page below which a PDF is treated as a scan and handed to docling. - # Measured on real uploads: scans at 0-17 chars/page, text-layer documents at 433-6834. - ATTACHMENT_PDF_TEXT_MIN_MEDIAN_CHARS: int = 32 - ATTACHMENT_TEXT_MAX_BYTES: int = 5_000_000 - AGENT_IMAGE_MAX_BYTES: int = 5_000_000 - AGENT_IMAGE_MAX_PIXELS: int = 16_777_216 - VECTOR_STORE: str = "faiss" # "faiss" or "elasticsearch" or "qdrant" or "milvus" or "lancedb" or "pgvector" - # Retriever keys an agent may use; must match RetrieverCreator.retrievers registry keys, - # NOT the legacy ``classic_rag`` label which never matched the registry. - RETRIEVERS_ENABLED: list = ["classic", "default"] - # Concurrent per-source searches in one retrieval; the query is embedded once and shared. - RETRIEVAL_MAX_PARALLEL_SOURCES: int = 4 - # Kill-switch for per-source retrieval dispatch; False collapses to a single retriever. - PER_SOURCE_RETRIEVAL_ENABLED: bool = True - GRAPHRAG_ENABLED: bool = False # gates graph-aware ingestion/retrieval - # Model for ingest-time graph extraction; None reuses LLM_PROVIDER/LLM_NAME. - GRAPHRAG_EXTRACTION_MODEL: Optional[str] = None - # Hard cap on chunks extracted per source (cost control). - GRAPHRAG_MAX_CHUNKS_FOR_EXTRACTION: int = 2000 - AGENT_NAME: str = "classic" - FALLBACK_LLM_PROVIDER: Optional[str] = None # provider for fallback llm - FALLBACK_LLM_NAME: Optional[str] = None # model name for fallback llm - FALLBACK_LLM_API_KEY: Optional[str] = None # api key for fallback llm - - # Google Drive integration - GOOGLE_CLIENT_ID: Optional[str] = None # Replace with your actual Google OAuth client ID - GOOGLE_CLIENT_SECRET: Optional[str] = None # Replace with your actual Google OAuth client secret - CONNECTOR_REDIRECT_BASE_URI: Optional[str] = ( - "http://127.0.0.1:7091/api/connectors/callback" ##add redirect url as it is to your provider's console(gcp) - ) - # Comma-separated frontend origins allowed to receive connector OAuth results, e.g. https://docsgpt.example.com. - # The callback origin and OIDC_FRONTEND_URL are always allowed; a loopback callback also allows localhost:5173. - CONNECTOR_ALLOWED_ORIGINS: Optional[str] = None - - # Microsoft Entra ID (Azure AD) integration - MICROSOFT_CLIENT_ID: Optional[str] = None # Azure AD Application (client) ID - MICROSOFT_CLIENT_SECRET: Optional[str] = None # Azure AD Application client secret - MICROSOFT_TENANT_ID: Optional[str] = "common" # Azure AD Tenant ID (or 'common' for multi-tenant) - MICROSOFT_AUTHORITY: Optional[str] = None # e.g., "https://login.microsoftonline.com/{tenant_id}" - - # Confluence Cloud integration - CONFLUENCE_CLIENT_ID: Optional[str] = None - CONFLUENCE_CLIENT_SECRET: Optional[str] = None - - # GitHub source - GITHUB_ACCESS_TOKEN: Optional[str] = None # PAT token with read repo access - - # LLM Cache - CACHE_REDIS_URL: str = "redis://localhost:6379/2" - - API_URL: str = "http://localhost:7091" # backend url for celery worker - - # Public base URL for user-facing endpoint references in prompts - PUBLIC_API_BASE_URL: Optional[str] = None - MCP_OAUTH_REDIRECT_URI: Optional[str] = None # public callback URL for MCP OAuth - INTERNAL_KEY: Optional[str] = None # internal api key for worker-to-backend auth - - API_KEY: Optional[str] = None # LLM api key (used by LLM_PROVIDER) - - # Provider-specific API keys (for multi-model support) - OPENAI_API_KEY: Optional[str] = None - ANTHROPIC_API_KEY: Optional[str] = None - GOOGLE_API_KEY: Optional[str] = None - GROQ_API_KEY: Optional[str] = None - HUGGINGFACE_API_KEY: Optional[str] = None - OPEN_ROUTER_API_KEY: Optional[str] = None - NOVITA_API_KEY: Optional[str] = None - - OPENAI_API_BASE: Optional[str] = None # azure openai api base url - OPENAI_API_VERSION: Optional[str] = None # azure openai api version - AZURE_DEPLOYMENT_NAME: Optional[str] = None # azure deployment name for answering - AZURE_EMBEDDINGS_DEPLOYMENT_NAME: Optional[str] = None # azure deployment name for embeddings - OPENAI_BASE_URL: Optional[str] = None # openai base url for open ai compatable models - - # elasticsearch - ELASTIC_CLOUD_ID: Optional[str] = None # cloud id for elasticsearch - ELASTIC_USERNAME: Optional[str] = None # username for elasticsearch - ELASTIC_PASSWORD: Optional[str] = None # password for elasticsearch - ELASTIC_URL: Optional[str] = None # url for elasticsearch - ELASTIC_INDEX: Optional[str] = "docsgpt" # index name for elasticsearch - - # Legacy AWS credentials from the retired SageMaker provider. Still read as a deprecated - # fallback by S3 storage; do not use for new deployments. - SAGEMAKER_REGION: Optional[str] = None - SAGEMAKER_ACCESS_KEY: Optional[str] = None - SAGEMAKER_SECRET_KEY: Optional[str] = None - - # Qdrant vectorstore config - QDRANT_COLLECTION_NAME: Optional[str] = "docsgpt" - QDRANT_LOCATION: Optional[str] = None - QDRANT_URL: Optional[str] = None - QDRANT_PORT: Optional[int] = 6333 - QDRANT_GRPC_PORT: int = 6334 - QDRANT_PREFER_GRPC: bool = False - QDRANT_HTTPS: Optional[bool] = None - QDRANT_API_KEY: Optional[str] = None - QDRANT_PREFIX: Optional[str] = None - QDRANT_TIMEOUT: Optional[float] = None - QDRANT_HOST: Optional[str] = None - QDRANT_PATH: Optional[str] = None - QDRANT_DISTANCE_FUNC: str = "Cosine" - - # PGVector config. postgres://, postgresql:// and postgresql+psycopg:// are all accepted - # and normalized internally for psycopg.connect(). - PGVECTOR_CONNECTION_STRING: Optional[str] = None - PGVECTOR_POOL_MAX_SIZE: int = 8 # per-process pool; 0 = one direct connection per store - # IVFFlat probes; None derives sqrt(lists) from the index. Higher = better recall, more scan. - PGVECTOR_IVFFLAT_PROBES: Optional[int] = None - # Milvus vectorstore config - MILVUS_COLLECTION_NAME: Optional[str] = "docsgpt" - # milvus-lite (embedded) database file, under the data home like the other local stores - MILVUS_URI: Optional[str] = Field(default_factory=lambda: str(home_dir() / "milvus_local.db")) - MILVUS_TOKEN: Optional[str] = "" - - # LanceDB vectorstore config - LANCEDB_PATH: str = Field(default_factory=lambda: str(home_dir() / "data" / "lancedb")) # LanceDB local data - LANCEDB_TABLE_NAME: Optional[str] = "docsgpts" # Name of the table to use for storing vectors - - FLASK_DEBUG_MODE: bool = False - STORAGE_TYPE: str = "local" # local or s3 - - # S3-compatible object storage (STORAGE_TYPE=s3): AWS S3, MinIO, R2, B2, Spaces, ... - # For non-AWS, set S3_ENDPOINT_URL and usually S3_PATH_STYLE=true. - S3_BUCKET_NAME: str = "docsgpt-test-bucket" - S3_ENDPOINT_URL: Optional[str] = None # custom endpoint for S3-compatible services; omit for AWS - S3_ACCESS_KEY_ID: Optional[str] = None - S3_SECRET_ACCESS_KEY: Optional[str] = None - S3_REGION: Optional[str] = None # AWS region; use "auto" for Cloudflare R2 - S3_PATH_STYLE: bool = False # path-style addressing (required by most non-AWS services) - - # Anonymous startup version check for security issues. - VERSION_CHECK: bool = True - URL_STRATEGY: str = "backend" # backend or s3 - - JWT_SECRET_KEY: str = "" - - # Encryption settings - ENCRYPTION_SECRET_KEY: str = "default-docsgpt-encryption-key" - - TTS_PROVIDER: str = "google_tts" # google_tts, elevenlabs, or none to switch text-to-speech off - ELEVENLABS_API_KEY: Optional[str] = None - STT_PROVIDER: str = "openai" # openai, faster_whisper, or none to switch speech-to-text off - OPENAI_STT_MODEL: str = "gpt-4o-mini-transcribe" - STT_LANGUAGE: Optional[str] = None - STT_MAX_FILE_SIZE_MB: int = 50 - STT_ENABLE_TIMESTAMPS: bool = False - STT_ENABLE_DIARIZATION: bool = False - - # Tool pre-fetch settings - ENABLE_TOOL_PREFETCH: bool = True - - # True persists Responses API calls server-side so previous_response_id can chain turns. - # False keeps them stateless, carrying reasoning across the tool loop as encrypted items. - OPENAI_RESPONSES_STORE: bool = False - # Cross-turn ``previous_response_id`` chaining (store mode only). The - # chained transcript lives on the provider and is invisible to every - # local guard, so it is bounded: a turn starts from the local history - # when the previous turn's reported prompt already reached the budget - # (default: the model's context window) or when the conversation was - # compressed after that turn was produced. - OPENAI_RESPONSES_CHAIN_ACROSS_TURNS: bool = True - OPENAI_RESPONSES_CHAIN_BUDGET_TOKENS: Optional[int] = None - # ``truncation: "auto"`` lets the provider drop the oldest input items - # instead of failing every request once a chain exceeds the model's window. - OPENAI_RESPONSES_TRUNCATION_AUTO: bool = False - # Prompt-cache hints on the Responses API: route a user's calls to the - # same cache shard (opaque per-user key), and request extended retention - # where offered. - OPENAI_PROMPT_CACHE_KEY: bool = True - OPENAI_PROMPT_CACHE_RETENTION: Optional[str] = None - OPENAI_REASONING_SUMMARY: str = "auto" - - # Lets OpenAI-compatible clients identify a logical chat by session header, which - # chat-completions itself has no field for. - V1_SESSION_TTL_SECONDS: int = 24 * 60 * 60 - # Optional cheaper model for conversation titles; unset reuses the answer model. - TITLE_MODEL_ID: Optional[str] = None - - # Config-free tools on by default in agentless chats. ``scheduler`` is dual-registered in - # BUILTIN_AGENT_TOOLS so one synthetic id resolves via defaults or the agent picker. - # Add "code_executor" and "artifact_generator" once a sandbox runner is configured — both - # execute through it and would fail on every call without one. - DEFAULT_CHAT_TOOLS: list = [ - "memory", - "read_webpage", - "scheduler", - ] - - # Conversation Compression Settings - ENABLE_CONVERSATION_COMPRESSION: bool = True - COMPRESSION_THRESHOLD_PERCENTAGE: float = 0.8 # Trigger at 80% of context - COMPRESSION_MODEL_OVERRIDE: Optional[str] = None # Use different model for compression - COMPRESSION_PROMPT_VERSION: str = "v1.0" # Track prompt iterations - COMPRESSION_MAX_HISTORY_POINTS: int = 3 # Keep only last N compression points to prevent DB bloat - # Per-field cap on the verbatim tail kept after a compression point (0 disables). - COMPRESSION_RECENT_FIELD_MAX_TOKENS: int = 8000 - # Cap on one tool result entering the LLM context (0 disables); journal/DB keep it whole. - TOOL_RESULT_MAX_TOKENS: int = 20000 - - # Agent Guardrails - GUARDRAILS_ENABLED: bool = True # master switch; False disables every stage - # Allowlist of GuardrailCreator.checks keys; empty means every registered check. - GUARDRAILS_CHECKS_ENABLED: list = [] - # A GuardrailsConfig fragment every agent inherits and cannot weaken; agents may add - # controls or make an action stricter, never looser. "enabled" is required — without it - # the floor parses but applies to nothing. Example: - # {"enabled": true, "mode": "scan_all", - # "controls": [{"check": "secrets", "stage": "output", "action": "redact"}]} - GUARDRAILS_FLOOR: dict = {} - # Judge model for the topic/policy checks; None reuses the request's model. - GUARDRAILS_JUDGE_MODEL: Optional[str] = None - # Persist scanned text alongside guardrail_events. Off by default: pre-redaction text is - # exactly the material a PII control exists to keep out of storage. - GUARDRAILS_STORE_SCANNED_TEXT: bool = False - GUARDRAILS_EVENTS_RETENTION_DAYS: int = Field(default=30, ge=1) - - # Internal SSE push channel (notifications + durable replay journal). - # False makes /api/events emit "push_disabled" and return; clients fall back to polling. - ENABLE_SSE_PUSH: bool = True - # Per-user durable backlog cap in entries; ~24h of replay at typical rates. - EVENTS_STREAM_MAXLEN: int = 1000 - # Bounds uvicorn's shutdown drain (uvicorn_worker doesn't forward --graceful-timeout). - # Keep below the gunicorn --timeout (180) watchdog. Used by BoundedDrainUvicornWorker. - GRACEFUL_SHUTDOWN_TIMEOUT_SECONDS: int = 30 - WSGI_THREADPOOL_WORKERS: int = 96 - SSE_KEEPALIVE_SECONDS: int = Field(default=15, ge=1) - # Simultaneous SSE connections per user; each holds a pooled async Redis connection for - # its lifetime. 8 covers multi-tab use without one user starving the pool. 0 disables. - SSE_MAX_CONCURRENT_PER_USER: int = 8 - # Pool size of the async Redis client behind the event-loop routes, per process. Every - # open notification tab, chat reconnect and device session holds one connection, so this - # caps concurrent streams per worker (redis-py's own default is 100). Keep the total - # across workers below the Redis server's maxclients (10000 by default). - ASYNC_REDIS_MAX_CONNECTIONS: int = Field(default=2000, ge=1) - # Backlog entries XRANGE returns per /api/events snapshot. Bounds what one replay moves - # from Redis to the wire: a client looping Last-Event-ID reconnects enumerates at most - # this many per round-trip, and the budget below bounds total throughput. - EVENTS_REPLAY_MAX_PER_REQUEST: int = 200 - EVENTS_REPLAY_MAX_AGE_HOURS: int = 48 - # Sliding-window cap on snapshot replays per user; exhausting it returns 429 with the - # cursor pinned so the client backs off until the window rolls over. - EVENTS_REPLAY_BUDGET_REQUESTS_PER_WINDOW: int = 30 - EVENTS_REPLAY_BUDGET_WINDOW_SECONDS: int = 60 - - # Retention for the message_events journal, enforced by the cleanup_message_events beat - # task. Replay only needs streams a client could still be tailing. - MESSAGE_EVENTS_RETENTION_DAYS: int = 14 - - # Remote Device feature. - REMOTE_DEVICE_SESSION_IDLE_SECONDS: int = 60 - REMOTE_DEVICE_REQUIRE_SIGNATURE: bool = False - REMOTE_DEVICE_PAIRING_TTL_SECONDS: int = 600 - # Redis broker tunables, routing invocations cross-process so a scheduled run reaches the - # web-held device session. The queue TTL must exceed the max drain deadline (605s) so a - # command for a briefly-offline device isn't evicted before its own drain gives up. - REMOTE_DEVICE_CMD_QUEUE_TTL_SECONDS: int = 900 - REMOTE_DEVICE_INVOCATION_TTL_SECONDS: int = 900 - REMOTE_DEVICE_OUTPUT_STREAM_MAXLEN: int = 10_000 - - # Scheduler (see scheduler.md). - SCHEDULE_DISPATCHER_INTERVAL: int = 30 - SCHEDULE_MIN_INTERVAL: int = 900 - SCHEDULE_MAX_PER_USER: int = 50 - SCHEDULE_RUN_TIMEOUT: int = 600 - SCHEDULE_MISFIRE_GRACE: int = 60 - SCHEDULE_AUTOPAUSE_FAILURES: int = 3 - SCHEDULE_ONCE_MAX_HORIZON: int = 31_536_000 - SCHEDULE_RUN_OUTPUT_RETENTION_DAYS: int = 90 - - # Code-execution sandbox. The app is a CLIENT of an always-on runner; defaults are safe so - # app import never fails when the sandbox is unconfigured. - SANDBOX_BACKEND: str = "jupyter" # "jupyter" (self-host) | "daytona" (Daytona Cloud) - # URL of the Jupyter Kernel Gateway runner (the docsgpt-sandbox service). - SANDBOX_GATEWAY_URL: str = "http://localhost:8888" - SANDBOX_GATEWAY_AUTH_TOKEN: Optional[str] = None # gateway auth token, if set - # Kernelspec per session. The env-scrubbing "docsgpt-python" spec keeps kernel code from - # reading the gateway token or operator secrets from os.environ; the stock "python3" spec - # inherits the gateway env verbatim and must not be used with untrusted code. - SANDBOX_KERNEL_NAME: str = "docsgpt-python" - SANDBOX_MAX_TTL: int = 1200 # hard cap (s) on agent-selectable keep-alive TTL - # Concurrent live sessions per process, backend-agnostic; at the cap an LRU-idle session is - # evicted. 0 or negative disables the cap. - SANDBOX_MAX_SESSIONS: int = 32 - SANDBOX_EXEC_TIMEOUT: int = 60 # default wall-clock cap (s) per exec call - SANDBOX_HTTP_TIMEOUT: int = 10 # fixed cap (s) for REST control calls (create/delete/alive/interrupt) - SANDBOX_MAX_OUTPUT_BYTES: int = 8 * 1024 * 1024 # cap on buffered stdout+stderr per exec - SANDBOX_MAX_FILE_BYTES: int = 10 * 1024 * 1024 # cap on get_file size routed through stdout - SANDBOX_MAX_INPUT_BYTES: int = 25 * 1024 * 1024 # cap on an input document staged into a sandbox session - # ``read_document`` parsing on a dedicated Celery ``parsing`` queue (backend parser). - DOCUMENT_PARSE_QUEUE: str = "parsing" # queue the parse_document task is routed to - DOCUMENT_PARSE_TIMEOUT: int = 120 # seconds the tool awaits the enqueued parse before degrading - # The base timeout is a FLOOR: the window grows with document size, because OCR cost scales - # with pages. Without this a large scan is silently dropped at the base window. - DOCUMENT_PARSE_TIMEOUT_PER_MB: int = 60 # extra seconds of parse window per MiB of input - DOCUMENT_PARSE_TIMEOUT_MAX: int = 900 # absolute ceiling on the size-scaled parse window - DOCUMENT_PARSE_MAX_BYTES: int = 0 # cap on a parsed document's bytes (0 = reuse SANDBOX_MAX_INPUT_BYTES) - DOCUMENT_MAX_DECOMPRESSED_BYTES: int = 300 * 1024 * 1024 - DOCUMENT_MAX_ARCHIVE_ENTRIES: int = 10000 - # Files per node passed natively to the LLM; past the cap they are extracted to text or - # dropped, to bound context and cost. Re-uses SANDBOX_MAX_INPUT_BYTES per file. - WORKFLOW_NODE_NATIVE_MAX_FILES: int = 5 - # Documents per node extracted via the parsing worker. Each issues a separate blocking - # parse; past the cap they are skipped with a truncation note. - WORKFLOW_NODE_EXTRACT_MAX_FILES: int = 5 - # Wall clock one node may spend on blocking parses, shared across all of them. Without it a - # node could serialize WORKFLOW_NODE_EXTRACT_MAX_FILES full windows on a web threadpool slot. - WORKFLOW_NODE_EXTRACT_BUDGET_SECONDS: int = 900 - # A run row is pre-created as ``running``; a disconnect or crash can strand it there. The - # beat reaper fails runs still ``running`` past this. Generous so a long run is never cut off. - WORKFLOW_RUN_STALE_SECONDS: int = 3600 - # Runner container caps, consumed by the docsgpt-sandbox compose service, not the app. - # These cgroup limits are part of the untrusted-code security boundary. - SANDBOX_MEMORY: str = "1g" # docker mem_limit for the runner container - SANDBOX_CPUS: str = "1.0" # docker cpu quota for the runner container - # Daytona Cloud backend (SANDBOX_BACKEND="daytona"). All knobs are optional so app import - # never fails when the backend is unused. - DAYTONA_API_KEY: Optional[str] = None # Daytona Cloud API key (secret) - DAYTONA_API_URL: Optional[str] = None # override Daytona API base URL, if self-targeting - DAYTONA_TARGET: Optional[str] = None # Daytona region/target, e.g. "us" - DAYTONA_SNAPSHOT: Optional[str] = None # image for new sandboxes; render libs via scripts/build_daytona_snapshot.py - DAYTONA_LANGUAGE: str = "python" # default runtime language for created sandboxes - DAYTONA_AUTO_STOP_INTERVAL: int = 15 # minutes idle before Daytona auto-stops a sandbox (0 disables) - DAYTONA_AUTO_DELETE_INTERVAL: int = 60 # minutes after stop before Daytona auto-deletes (-1 disables) - DAYTONA_MAX_SANDBOXES: int = 50 # cap on concurrent live Daytona sandboxes (cost-DoS guard) - # Per-user artifact quotas, enforced at persistence time. 0 or negative disables a quota. - ARTIFACT_MAX_BYTES: int = 50 * 1024 * 1024 # cap on a single stored artifact version's bytes - ARTIFACT_MAX_COUNT_PER_USER: int = 5000 # cap on artifacts a user may own - ARTIFACT_MAX_TOTAL_BYTES_PER_USER: int = 5 * 1024 * 1024 * 1024 # cap on a user's total stored bytes - - @field_validator("POSTGRES_URI", mode="before") - @classmethod - def _normalize_postgres_uri_validator(cls, v): - return normalize_postgres_uri(v) - - @field_validator("PGVECTOR_CONNECTION_STRING", mode="before") - @classmethod - def _normalize_pgvector_connection_string_validator(cls, v): - return normalize_pgvector_connection_string(v) - - @field_validator( - "API_KEY", - "OPENAI_API_KEY", - "ANTHROPIC_API_KEY", - "GOOGLE_API_KEY", - "GROQ_API_KEY", - "HUGGINGFACE_API_KEY", - "NOVITA_API_KEY", - "EMBEDDINGS_KEY", - "FALLBACK_LLM_API_KEY", - "QDRANT_API_KEY", - "ELEVENLABS_API_KEY", - "INTERNAL_KEY", - mode="before", - ) - @classmethod - def normalize_api_key(cls, v: Optional[str]) -> Optional[str]: - """ - Normalize API keys: convert 'None', 'none', empty strings, - and whitespace-only strings to actual None. - Handles Pydantic loading 'None' from .env as string "None". - """ - if v is None: - return None - if not isinstance(v, str): - return v - stripped = v.strip() - if stripped == "" or stripped.lower() == "none": - return None - return stripped - - -settings = Settings(_env_file=env_file(), _env_file_encoding="utf-8") diff --git a/docsgpt/core/settings/__init__.py b/docsgpt/core/settings/__init__.py new file mode 100644 index 00000000..2af13cea --- /dev/null +++ b/docsgpt/core/settings/__init__.py @@ -0,0 +1,77 @@ +"""Application settings. + +``settings`` is the process-wide instance, loaded from the environment and the +``.env`` file in the data home (see ``docsgpt.core.paths``). Every setting is a +flat attribute, ``settings.NAME``, matching the environment variable of the +same name. + +The definitions are split by domain into the modules of this package; each +module owns one ``SettingsGroup`` and ``Settings`` composes them all. Add a new +setting to the group it belongs to (or add a group and list it in +``SETTINGS_GROUPS``), with a ``description`` -- the settings reference in the +docs is generated from these definitions. +""" + +from __future__ import annotations + +from typing import Optional + +from docsgpt.core.paths import env_file, home_dir +from docsgpt.core.settings._shared import SettingsGroup, normalize_secret +from docsgpt.core.settings.agents import AgentSettings +from docsgpt.core.settings.auth import AuthSettings +from docsgpt.core.settings.connectors import ConnectorSettings +from docsgpt.core.settings.database import DatabaseSettings +from docsgpt.core.settings.embeddings import EmbeddingsSettings +from docsgpt.core.settings.events import EventsSettings +from docsgpt.core.settings.guardrails import GuardrailSettings +from docsgpt.core.settings.ingestion import IngestionSettings +from docsgpt.core.settings.llm import LLMSettings +from docsgpt.core.settings.ocr import OCRSettings +from docsgpt.core.settings.retrieval import RetrievalSettings +from docsgpt.core.settings.sandbox import SandboxSettings +from docsgpt.core.settings.scheduler import SchedulerSettings +from docsgpt.core.settings.server import ServerSettings +from docsgpt.core.settings.speech import SpeechSettings +from docsgpt.core.settings.storage import StorageSettings +from docsgpt.core.settings.vectorstores import VectorStoreSettings +from docsgpt.core.settings.workers import WorkerSettings + +#: Every settings group, in the order the generated reference lists them. +SETTINGS_GROUPS: tuple[tuple[str, type[SettingsGroup]], ...] = ( + ("Authentication", AuthSettings), + ("LLM providers", LLMSettings), + ("Embeddings", EmbeddingsSettings), + ("Retrieval", RetrievalSettings), + ("Vector stores", VectorStoreSettings), + ("User-data database", DatabaseSettings), + ("Workers", WorkerSettings), + ("Ingestion and parsing", IngestionSettings), + ("OCR", OCRSettings), + ("File storage", StorageSettings), + ("Connectors", ConnectorSettings), + ("Server", ServerSettings), + ("Events and devices", EventsSettings), + ("Agents", AgentSettings), + ("Guardrails", GuardrailSettings), + ("Scheduler", SchedulerSettings), + ("Sandbox", SandboxSettings), + ("Speech", SpeechSettings), +) + +# Runtime data home (DOCSGPT_HOME, the checkout, or cwd); see docsgpt.core.paths. +current_dir = str(home_dir()) + + +class Settings(*(group for _, group in SETTINGS_GROUPS)): + """All settings, composed from the per-domain groups in this package.""" + + @classmethod + def normalize_api_key(cls, v: Optional[str]) -> Optional[str]: + """Normalize a secret the way the per-field validators do; kept for callers that reuse it.""" + return normalize_secret(v) + + +settings = Settings(_env_file=env_file(), _env_file_encoding="utf-8") + +__all__ = ["SETTINGS_GROUPS", "Settings", "SettingsGroup", "current_dir", "settings"] diff --git a/docsgpt/core/settings/_shared.py b/docsgpt/core/settings/_shared.py new file mode 100644 index 00000000..5b8bfec4 --- /dev/null +++ b/docsgpt/core/settings/_shared.py @@ -0,0 +1,80 @@ +"""Building blocks shared by the settings groups. + +Every group in this package is a :class:`SettingsGroup`: a ``BaseSettings`` +subclass that owns one domain's fields. ``docsgpt.core.settings.Settings`` +inherits from all of them, so the composed class keeps the flat +``settings.NAME`` attributes the rest of the codebase reads while each +domain's definitions live in their own module. +""" + +from __future__ import annotations + +import types +import typing +from typing import Any, Optional + +from pydantic import model_validator +from pydantic_settings import BaseSettings, SettingsConfigDict + + +def _is_optional_str(annotation: Any) -> bool: + """``Optional[str]`` or ``Optional[Literal[...]]`` whose choices are all strings.""" + if typing.get_origin(annotation) not in (typing.Union, types.UnionType): + return False + members = set(typing.get_args(annotation)) + if type(None) not in members or len(members) != 2: + return False + (member,) = members - {type(None)} + if member is str: + return True + return typing.get_origin(member) is typing.Literal and all( + isinstance(choice, str) for choice in typing.get_args(member) + ) + + +class SettingsGroup(BaseSettings): + """Base for one domain's settings; groups are composed into ``Settings``. + + Every ``Optional[str]`` field (and optional string ``Literal``) treats the spellings an unset value has in a + ``.env`` file (``KEY=``, ``KEY=None``, whitespace) as ``None``, so a check + like ``if settings.OIDC_ISSUER`` or a fallback like ``settings.X or default`` + sees "unset" rather than a truthy placeholder string. Real values are + stripped. Fields typed ``str`` keep whatever they are given. + """ + + model_config = SettingsConfigDict(extra="ignore") + + @model_validator(mode="before") + @classmethod + def _unset_optional_strings(cls, data: Any) -> Any: + if not isinstance(data, dict): + return data + data = dict(data) + for name, field in cls.model_fields.items(): + if name in data and _is_optional_str(field.annotation): + data[name] = normalize_secret(data[name]) + return data + + +def normalize_choice(value: Any) -> Any: + """Case-fold a closed-choice setting so ``PGVector`` and ``pgvector`` are the same choice.""" + if isinstance(value, str): + return value.strip().lower() + return value + + +def normalize_secret(value: Optional[str]) -> Optional[str]: + """Map the ways an unset secret reaches us from ``.env`` to ``None``. + + ``.env`` files carry ``KEY=None`` and ``KEY=`` for "not set", and pydantic + would otherwise keep those as the strings ``"None"`` and ``""``. Whitespace + around a real value is stripped. + """ + if value is None: + return None + if not isinstance(value, str): + return value + stripped = value.strip() + if stripped == "" or stripped.lower() == "none": + return None + return stripped diff --git a/docsgpt/core/settings/agents.py b/docsgpt/core/settings/agents.py new file mode 100644 index 00000000..b15a4d62 --- /dev/null +++ b/docsgpt/core/settings/agents.py @@ -0,0 +1,91 @@ +"""Agent runtime: default tools, context management, workflows and artifacts.""" + +from __future__ import annotations + +from typing import Optional + +from pydantic import Field + +from docsgpt.core.settings._shared import SettingsGroup + + +class AgentSettings(SettingsGroup): + """What an agent may do per turn and how its context is kept within budget.""" + + AGENT_NAME: str = Field(default="classic", description="Default agent type for agentless chats.") + DEFAULT_AGENT_LIMITS: dict[str, int] = Field( + default={"token_limit": 50000, "request_limit": 500}, + description="Per-agent default quotas: tokens and requests.", + ) + DEFAULT_CHAT_TOOLS: list[str] = Field( + default=["memory", "read_webpage", "scheduler"], + description=( + "Config-free tools on by default in agentless chats. scheduler is dual-registered in " + "BUILTIN_AGENT_TOOLS so one synthetic id resolves via defaults or the agent picker. Add " + "code_executor and artifact_generator once a sandbox runner is configured; both execute through " + "it and would fail on every call without one." + ), + ) + ENABLE_TOOL_PREFETCH: bool = Field(default=True, description="Pre-fetch retrieval before the agent's first turn.") + TOOL_RESULT_MAX_TOKENS: int = Field( + default=20000, + ge=0, + description="Cap on one tool result entering the LLM context (0 disables); journal and DB keep it whole.", + ) + + # Conversation compression. + ENABLE_CONVERSATION_COMPRESSION: bool = Field( + default=True, description="Compress long conversations once they approach the context window." + ) + COMPRESSION_THRESHOLD_PERCENTAGE: float = Field( + default=0.8, gt=0, le=1, description="Fraction of the context window at which compression triggers." + ) + COMPRESSION_MODEL_OVERRIDE: Optional[str] = Field( + default=None, description="Use a different model for compression; unset reuses the answer model." + ) + COMPRESSION_PROMPT_VERSION: str = Field(default="v1.0", description="Tracks compression prompt iterations.") + COMPRESSION_MAX_HISTORY_POINTS: int = Field( + default=3, description="Keep only the last N compression points to prevent DB bloat." + ) + COMPRESSION_RECENT_FIELD_MAX_TOKENS: int = Field( + default=8000, ge=0, description="Per-field cap on the verbatim tail kept after a compression point (0 disables)." + ) + + # Workflows. + WORKFLOW_NODE_NATIVE_MAX_FILES: int = Field( + default=5, + description=( + "Files per node passed natively to the LLM; past the cap they are extracted to text or dropped, to " + "bound context and cost. Re-uses SANDBOX_MAX_INPUT_BYTES per file." + ), + ) + WORKFLOW_NODE_EXTRACT_MAX_FILES: int = Field( + default=5, + description=( + "Documents per node extracted via the parsing worker. Each issues a separate blocking parse; past " + "the cap they are skipped with a truncation note." + ), + ) + WORKFLOW_NODE_EXTRACT_BUDGET_SECONDS: int = Field( + default=900, + description=( + "Wall clock one node may spend on blocking parses, shared across all of them. Without it a node " + "could serialize WORKFLOW_NODE_EXTRACT_MAX_FILES full windows on a web threadpool slot." + ), + ) + WORKFLOW_RUN_STALE_SECONDS: int = Field( + default=3600, + description=( + "A run row is pre-created as running; a disconnect or crash can strand it there. The beat reaper " + "fails runs still running past this. Generous so a long run is never cut off." + ), + ) + + # Per-user artifact quotas, enforced at persistence time. 0 or negative disables a quota. + ARTIFACT_MAX_BYTES: int = Field( + default=50 * 1024 * 1024, description="Cap on a single stored artifact version's bytes (0 disables)." + ) + ARTIFACT_MAX_COUNT_PER_USER: int = Field(default=5000, description="Cap on artifacts a user may own (0 disables).") + ARTIFACT_MAX_TOTAL_BYTES_PER_USER: int = Field( + default=5 * 1024 * 1024 * 1024, description="Cap on a user's total stored artifact bytes (0 disables)." + ) diff --git a/docsgpt/core/settings/auth.py b/docsgpt/core/settings/auth.py new file mode 100644 index 00000000..85c57b84 --- /dev/null +++ b/docsgpt/core/settings/auth.py @@ -0,0 +1,102 @@ +"""Authentication, SSO and provisioning.""" + +from __future__ import annotations + +from typing import Literal, Optional + +from pydantic import Field, field_validator, model_validator + +from docsgpt.core.settings._shared import SettingsGroup, normalize_choice + + +#: Settings an OIDC deployment cannot run without; checked when AUTH_TYPE=oidc. +OIDC_REQUIRED = ("OIDC_ISSUER", "OIDC_CLIENT_ID", "OIDC_FRONTEND_URL") + + +class AuthSettings(SettingsGroup): + """How users authenticate: none, a shared token, per-session JWTs, or OIDC SSO.""" + + AUTH_TYPE: Optional[Literal["simple_jwt", "session_jwt", "oidc"]] = Field( + default=None, + description="Authentication mode: simple_jwt, session_jwt, oidc, or unset (None) for no authentication.", + ) + JWT_SECRET_KEY: str = Field( + default="", + description=( + "Signing key for session tokens and other signed capabilities. Required on every replica in " + "production; local development may fall back to a key generated on disk." + ), + ) + ENCRYPTION_SECRET_KEY: str = Field( + default="default-docsgpt-encryption-key", + description="Key used to encrypt stored credentials such as tool and connector secrets.", + ) + INTERNAL_KEY: Optional[str] = Field( + default=None, description="Internal API key for worker-to-backend authentication." + ) + + # OIDC SSO (AUTH_TYPE=oidc): any OpenID Connect IdP with discovery (Authentik, Keycloak, ...). + OIDC_ISSUER: Optional[str] = Field( + default=None, + description="OIDC issuer URL with discovery, e.g. https://auth.example.com/application/o/docsgpt/.", + ) + OIDC_CLIENT_ID: Optional[str] = Field(default=None, description="OIDC client id.") + OIDC_CLIENT_SECRET: Optional[str] = Field( + default=None, description="OIDC client secret. Optional; PKCE is always used." + ) + OIDC_SCOPES: str = Field(default="openid profile email", description="Scopes requested from the IdP.") + OIDC_USER_ID_CLAIM: str = Field( + default="sub", description="ID-token claim mapped to the DocsGPT user id." + ) + OIDC_FRONTEND_URL: Optional[str] = Field( + default=None, description="Browser-facing app origin, e.g. http://localhost:5173." + ) + OIDC_REDIRECT_URI: Optional[str] = Field( + default=None, description="Override for the callback URL; default is /api/auth/oidc/callback." + ) + OIDC_SESSION_LIFETIME_SECONDS: int = Field( + default=28800, gt=0, description="Lifetime of the minted session JWT in seconds (8h)." + ) + OIDC_PROVIDER_NAME: Optional[str] = Field( + default=None, description='Sign-in button label, e.g. "Acme SSO".' + ) + OIDC_ALLOWED_GROUPS: Optional[str] = Field( + default=None, description="Comma-separated group allowlist; unset admits any authenticated user." + ) + OIDC_GROUPS_CLAIM: str = Field( + default="groups", description="ID-token/userinfo claim carrying group membership." + ) + OIDC_ADMIN_GROUPS: Optional[str] = Field( + default=None, description="Comma-separated groups granted admin; unset means no OIDC admin mapping." + ) + + LOCAL_MODE_ADMIN: bool = Field( + default=False, + description=( + "Grant admin without a database role. Persisted admin grants live in user_roles (AUTH_TYPE=oidc " + "only); this is the only non-DB admin path, for AUTH_TYPE=None self-host. MUST stay False if " + "networked." + ), + ) + + # SCIM 2.0 provisioning (IdP-driven user create/deactivate at /scim/v2). + SCIM_ENABLED: bool = Field(default=False, description="Enable SCIM 2.0 provisioning at /scim/v2.") + SCIM_TOKEN: Optional[str] = Field( + default=None, description="Bearer token for IdP SCIM clients (required when SCIM is enabled)." + ) + + @field_validator("AUTH_TYPE", mode="before") + @classmethod + def _normalize_auth_type(cls, v): + # Unset spellings ("None", "") became None on the group base; this only case-folds a value. + return normalize_choice(v) + + @model_validator(mode="after") + def _require_dependent_settings(self): + if self.AUTH_TYPE == "oidc": + missing = [name for name in OIDC_REQUIRED if not getattr(self, name)] + if missing: + raise ValueError(f"AUTH_TYPE=oidc requires settings: {', '.join(missing)}") + if self.SCIM_ENABLED and not self.SCIM_TOKEN: + raise ValueError("SCIM_ENABLED requires settings: SCIM_TOKEN") + return self diff --git a/docsgpt/core/settings/connectors.py b/docsgpt/core/settings/connectors.py new file mode 100644 index 00000000..b4199300 --- /dev/null +++ b/docsgpt/core/settings/connectors.py @@ -0,0 +1,51 @@ +"""OAuth credentials for external source connectors.""" + +from __future__ import annotations + +from typing import Optional + +from pydantic import Field + +from docsgpt.core.settings._shared import SettingsGroup + + +class ConnectorSettings(SettingsGroup): + """Client credentials and callback URLs for Google Drive, Microsoft, Confluence, GitHub and MCP.""" + + # Google Drive integration. + GOOGLE_CLIENT_ID: Optional[str] = Field(default=None, description="Google OAuth client id.") + GOOGLE_CLIENT_SECRET: Optional[str] = Field(default=None, description="Google OAuth client secret.") + CONNECTOR_REDIRECT_BASE_URI: str = Field( + default="http://127.0.0.1:7091/api/connectors/callback", + description="OAuth callback URL; register it as-is in your provider's console (e.g. GCP).", + ) + CONNECTOR_ALLOWED_ORIGINS: Optional[str] = Field( + default=None, + description=( + "Comma-separated frontend origins allowed to receive connector OAuth results, e.g. " + "https://docsgpt.example.com. The callback origin and OIDC_FRONTEND_URL are always allowed; a " + "loopback callback also allows localhost:5173." + ), + ) + + # Microsoft Entra ID (Azure AD) integration. + MICROSOFT_CLIENT_ID: Optional[str] = Field(default=None, description="Azure AD application (client) id.") + MICROSOFT_CLIENT_SECRET: Optional[str] = Field(default=None, description="Azure AD application client secret.") + MICROSOFT_TENANT_ID: str = Field( + default="common", description="Azure AD tenant id, or 'common' for multi-tenant." + ) + MICROSOFT_AUTHORITY: Optional[str] = Field( + default=None, + description="Authority URL override; unset derives https://login.microsoftonline.com/.", + ) + + # Confluence Cloud integration. + CONFLUENCE_CLIENT_ID: Optional[str] = Field(default=None, description="Confluence Cloud OAuth client id.") + CONFLUENCE_CLIENT_SECRET: Optional[str] = Field(default=None, description="Confluence Cloud OAuth client secret.") + + # GitHub source. + GITHUB_ACCESS_TOKEN: Optional[str] = Field(default=None, description="GitHub PAT with read access to repositories.") + + MCP_OAUTH_REDIRECT_URI: Optional[str] = Field( + default=None, description="Public callback URL for MCP OAuth; unset derives it from CONNECTOR_REDIRECT_BASE_URI." + ) diff --git a/docsgpt/core/settings/database.py b/docsgpt/core/settings/database.py new file mode 100644 index 00000000..e88e03a9 --- /dev/null +++ b/docsgpt/core/settings/database.py @@ -0,0 +1,36 @@ +"""User-data Postgres and schema management at startup.""" + +from __future__ import annotations + +from typing import Optional + +from pydantic import Field, field_validator + +from docsgpt.core.db_uri import normalize_postgres_uri +from docsgpt.core.settings._shared import SettingsGroup + + +class DatabaseSettings(SettingsGroup): + """The Postgres database holding users, conversations and sources, and what startup may do to it.""" + + POSTGRES_URI: Optional[str] = Field(default=None, description="User-data Postgres connection URI.") + AUTO_MIGRATE: bool = Field( + default=True, + description="On startup, apply pending Alembic migrations. Disable if you manage schema out-of-band.", + ) + AUTO_CREATE_DB: bool = Field( + default=True, + description="On startup, create the target Postgres database if missing (needs CREATEDB privilege).", + ) + AUTO_VECTOR_SCHEMA: bool = Field( + default=True, + description=( + "On startup, create the pgvector/graph tables and verify the embedding dimension. No Alembic " + "migration covers the vector DB (it may be a separate cluster); set False to manage it yourself." + ), + ) + + @field_validator("POSTGRES_URI", mode="before") + @classmethod + def _normalize_postgres_uri(cls, v): + return normalize_postgres_uri(v) diff --git a/docsgpt/core/settings/embeddings.py b/docsgpt/core/settings/embeddings.py new file mode 100644 index 00000000..f8c5c130 --- /dev/null +++ b/docsgpt/core/settings/embeddings.py @@ -0,0 +1,88 @@ +"""Embedding model selection and where it runs.""" + +from __future__ import annotations + +from typing import Literal, Optional + +from pydantic import Field, field_validator + +from docsgpt.core.paths import home_dir +from docsgpt.core.settings._shared import SettingsGroup, normalize_choice + + +class EmbeddingsSettings(SettingsGroup): + """The embedding model, remote or local, and the batching around it.""" + + EMBEDDINGS_NAME: str = Field( + default="huggingface_sentence-transformers/all-mpnet-base-v2", + description=( + "Embedding model. The legacy model is the default on purpose: an install that never pinned this " + "has vectors from it, and granite is the same width so a swap would fail silently. New installs " + "get granite from .env-template; existing ones switch by setting this and running " + "docsgpt.scripts.reembed." + ), + ) + EMBEDDINGS_BASE_URL: Optional[str] = Field( + default=None, description="Remote embeddings API URL (OpenAI-compatible)." + ) + EMBEDDINGS_KEY: Optional[str] = Field( + default=None, description="API key for embeddings (with OpenAI, the same value as API_KEY)." + ) + EMBEDDINGS_MAX_INPUT_TOKENS: Optional[int] = Field( + default=None, description="Truncate each remote embed input to N tokens (overflow is lost)." + ) + EMBEDDINGS_BATCH_SIZE: int = Field( + default=32, ge=1, description="Chunks per store transaction and per remote embed request." + ) + EMBEDDINGS_MODEL_BATCH_SIZE: int = Field( + default=1, + ge=1, + description=( + "Documents per local ONNX forward pass. Each pass pads to its longest input, and that waste grows " + "with the square of chunk length: at 1250 tokens, 32 peaked at 6.6 GB, 1 at 2.9 GB." + ), + ) + EMBEDDINGS_THREADS: Optional[int] = Field( + default=None, + description=( + "Intra-op threads for the local ONNX runner; unset uses every core. It scales sub-linearly, so " + "several single-threaded workers beat one many-threaded process on the same cores." + ), + ) + EMBEDDINGS_CACHE_DIR: Optional[str] = Field( + default_factory=lambda: str(home_dir() / "models"), + description=( + "Where embedding models and their tokenizers are cached. Persistent by default: FastEmbed's own " + "default is the temp dir." + ), + ) + EMBEDDINGS_POOLING: Optional[Literal["cls", "mean"]] = Field( + default=None, + description=( + 'Pooling strategy ("cls" or "mean"). Read from the model\'s own repository; set only for a ' + "repository that declares none, or to override what it declares." + ), + ) + EMBEDDINGS_NORMALIZE: Optional[bool] = Field( + default=None, + description=( + "L2-normalise embeddings. Read from the model's own repository; set only for a repository that " + "declares nothing, or to override what it declares." + ), + ) + EMBEDDINGS_DELEGATE_TO_WORKER: bool = Field( + default=True, + description=( + "Embed on the worker so the API holds no model (~890 MB), at one broker round trip per query. " + "Ignored when EMBEDDINGS_BASE_URL is set, which is the better answer for production." + ), + ) + EMBEDDINGS_QUEUE: str = Field(default="embeddings", description="Celery queue the embed task is routed to.") + EMBEDDINGS_DELEGATE_TIMEOUT: int = Field( + default=60, gt=0, description="Seconds the API waits for the worker to return an embedding." + ) + + @field_validator("EMBEDDINGS_POOLING", mode="before") + @classmethod + def _normalize_pooling(cls, v): + return normalize_choice(v) diff --git a/docsgpt/core/settings/events.py b/docsgpt/core/settings/events.py new file mode 100644 index 00000000..3b170be9 --- /dev/null +++ b/docsgpt/core/settings/events.py @@ -0,0 +1,90 @@ +"""Server-sent events, replay journal and remote-device sessions.""" + +from __future__ import annotations + +from pydantic import Field + +from docsgpt.core.settings._shared import SettingsGroup + + +class EventsSettings(SettingsGroup): + """The internal push channel (notifications and durable replay) and the Redis pool behind it.""" + + ENABLE_SSE_PUSH: bool = Field( + default=True, + description=( + "Internal SSE push channel (notifications and durable replay journal). False makes /api/events emit " + '"push_disabled" and return; clients fall back to polling.' + ), + ) + EVENTS_STREAM_MAXLEN: int = Field( + default=1000, ge=1, description="Per-user durable backlog cap in entries; ~24h of replay at typical rates." + ) + SSE_KEEPALIVE_SECONDS: int = Field(default=15, ge=1, description="Interval between SSE keepalive comments.") + SSE_MAX_CONCURRENT_PER_USER: int = Field( + default=8, + ge=0, + description=( + "Simultaneous SSE connections per user; each holds a pooled async Redis connection for its lifetime. " + "8 covers multi-tab use without one user starving the pool. 0 disables." + ), + ) + ASYNC_REDIS_MAX_CONNECTIONS: int = Field( + default=2000, + ge=1, + description=( + "Pool size of the async Redis client behind the event-loop routes, per process. Every open " + "notification tab, chat reconnect and device session holds one connection, so this caps concurrent " + "streams per worker (redis-py's own default is 100). Keep the total across workers below the Redis " + "server's maxclients (10000 by default)." + ), + ) + EVENTS_REPLAY_MAX_PER_REQUEST: int = Field( + default=200, + ge=1, + description=( + "Backlog entries XRANGE returns per /api/events snapshot. Bounds what one replay moves from Redis to " + "the wire: a client looping Last-Event-ID reconnects enumerates at most this many per round-trip." + ), + ) + EVENTS_REPLAY_MAX_AGE_HOURS: int = Field(default=48, description="Oldest backlog entry a replay will return.") + EVENTS_REPLAY_BUDGET_REQUESTS_PER_WINDOW: int = Field( + default=30, + description=( + "Sliding-window cap on snapshot replays per user; exhausting it returns 429 with the cursor pinned " + "so the client backs off until the window rolls over." + ), + ) + EVENTS_REPLAY_BUDGET_WINDOW_SECONDS: int = Field(default=60, description="Length of the replay budget window.") + MESSAGE_EVENTS_RETENTION_DAYS: int = Field( + default=14, + gt=0, + description=( + "Retention for the message_events journal, enforced by the cleanup_message_events beat task. Replay " + "only needs streams a client could still be tailing." + ), + ) + + # Remote Device feature. + REMOTE_DEVICE_SESSION_IDLE_SECONDS: int = Field( + default=60, gt=0, description="Seconds without a heartbeat before a remote-device session is considered idle." + ) + REMOTE_DEVICE_REQUIRE_SIGNATURE: bool = Field( + default=False, description="Require signed commands from remote devices." + ) + REMOTE_DEVICE_PAIRING_TTL_SECONDS: int = Field(default=600, gt=0, description="Lifetime of a pairing code.") + REMOTE_DEVICE_CMD_QUEUE_TTL_SECONDS: int = Field( + default=900, + gt=605, + description=( + "Redis TTL of the per-device command queue, routing invocations cross-process so a scheduled run " + "reaches the web-held device session. Must exceed the max drain deadline (605s) so a command for a " + "briefly-offline device isn't evicted before its own drain gives up." + ), + ) + REMOTE_DEVICE_INVOCATION_TTL_SECONDS: int = Field( + default=900, gt=0, description="Redis TTL of a pending remote-device invocation." + ) + REMOTE_DEVICE_OUTPUT_STREAM_MAXLEN: int = Field( + default=10_000, description="Cap on buffered output entries per remote-device invocation stream." + ) diff --git a/docsgpt/core/settings/guardrails.py b/docsgpt/core/settings/guardrails.py new file mode 100644 index 00000000..871f748b --- /dev/null +++ b/docsgpt/core/settings/guardrails.py @@ -0,0 +1,40 @@ +"""Agent guardrails.""" + +from __future__ import annotations + +from typing import Any, Optional + +from pydantic import Field + +from docsgpt.core.settings._shared import SettingsGroup + + +class GuardrailSettings(SettingsGroup): + """Input/output checks every agent runs, and the floor no agent may weaken.""" + + GUARDRAILS_ENABLED: bool = Field(default=True, description="Master switch; False disables every stage.") + GUARDRAILS_CHECKS_ENABLED: list[str] = Field( + default=[], description="Allowlist of GuardrailCreator.checks keys; empty means every registered check." + ) + GUARDRAILS_FLOOR: dict[str, Any] = Field( + default={}, + description=( + "A GuardrailsConfig fragment every agent inherits and cannot weaken; agents may add controls or " + 'make an action stricter, never looser. "enabled" is required; without it the floor parses but ' + 'applies to nothing. Example: {"enabled": true, "mode": "scan_all", "controls": [{"check": ' + '"secrets", "stage": "output", "action": "redact"}]}' + ), + ) + GUARDRAILS_JUDGE_MODEL: Optional[str] = Field( + default=None, description="Judge model for the topic/policy checks; unset reuses the request's model." + ) + GUARDRAILS_STORE_SCANNED_TEXT: bool = Field( + default=False, + description=( + "Persist scanned text alongside guardrail_events. Off by default: pre-redaction text is exactly the " + "material a PII control exists to keep out of storage." + ), + ) + GUARDRAILS_EVENTS_RETENTION_DAYS: int = Field( + default=30, ge=1, description="Days guardrail events are kept before the cleanup task removes them." + ) diff --git a/docsgpt/core/settings/ingestion.py b/docsgpt/core/settings/ingestion.py new file mode 100644 index 00000000..d44a1360 --- /dev/null +++ b/docsgpt/core/settings/ingestion.py @@ -0,0 +1,152 @@ +"""Uploads, document parsing and the size caps that keep one file from taking a worker down.""" + +from __future__ import annotations + +from typing import Literal + +from pydantic import Field, field_validator + +from docsgpt.core.settings._shared import SettingsGroup, normalize_choice + + +class IngestionSettings(SettingsGroup): + """Upload limits, the parser engine, and per-format byte caps for ingestion and attachments.""" + + UPLOAD_FOLDER: str = Field(default="inputs", description="Directory under the data home for uploaded sources.") + UPLOAD_MAX_REQUEST_BYTES: int = Field( + default=256 * 1024 * 1024, + gt=0, + description="Cap on an upload request body; applied by Flask before multipart parsing.", + ) + UPLOAD_MAX_FILE_BYTES: int = Field( + default=100 * 1024 * 1024, gt=0, description="Cap on a single uploaded file; also enforced while copying." + ) + PARSE_SPEC_MAX_BYTES: int = Field( + default=10 * 1024 * 1024, gt=0, description="Cap on an OpenAPI/tool spec file accepted for parsing." + ) + # ZIP limits apply cumulatively across nested archives in one extraction. + UPLOAD_MAX_ARCHIVE_BYTES: int = Field( + default=250 * 1024 * 1024, gt=0, description="Cap on total bytes extracted from one uploaded archive." + ) + UPLOAD_MAX_ARCHIVE_FILES: int = Field( + default=10_000, gt=0, description="Cap on files extracted from one uploaded archive." + ) + UPLOAD_MAX_ARCHIVE_RATIO: int = Field( + default=1000, gt=0, description="Maximum decompressed-to-compressed ratio before an archive is rejected." + ) + UPLOAD_MAX_ARCHIVE_DEPTH: int = Field( + default=3, ge=0, description="Maximum nesting depth of archives inside archives." + ) + PARSE_PDF_AS_IMAGE: bool = Field(default=False, description="Render PDF pages to images before parsing.") + PARSE_IMAGE_REMOTE: bool = Field(default=False, description="Send images to a remote parser.") + DOC_PARSER_ENGINE: Literal["anydoc", "docling"] = Field( + default="anydoc", + description=( + 'Document parser for source ingestion, chat attachments and the read_document tool. "anydoc" ' + "(default): firecrawl-anydoc, a Rust converter with no ML models; milliseconds per file, ~100 MB " + 'peak RSS. "docling": the layout/table-model pipeline (optional install; needed for ' + "read_document's structured output and the docling OCR backend). Files anydoc cannot convert " + "(scanned PDFs, malformed input) fall back to docling when it is installed, otherwise to the native " + "OCR parsers (OCR on) or the legacy parsers. Rollback to the previous behaviour is this one variable." + ), + ) + DOCLING_PIPELINE_QUEUE_MAX_SIZE: int = Field( + default=2, + description=( + "Pages docling's threaded pipeline buffers in flight; the library default (100) drives worker RSS " + "to ~3 GB on a mid-size PDF." + ), + ) + DOCLING_COMPILE_TORCH_MODELS: bool = Field( + default=False, description="Let docling torch.compile its models (slower start, faster pages)." + ) + DOCLING_TABULAR_MAX_BYTES: int = Field( + default=2_000_000, description="Largest CSV/XLSX docling will parse, in bytes." + ) + DOCLING_MARKUP_MAX_BYTES: int = Field( + default=8_000_000, description="Largest HTML/XML docling will parse, in bytes." + ) + MARKUP_MAX_BYTES: int = Field( + default=8_000_000, + ge=0, + description=( + "HTML/XHTML larger than this (bytes) are head-truncated before the markdownify parser runs (the " + "anydoc engine's HTML path). The tree that path builds costs ~50x the input (30 MB of HTML measured " + "at 1.6 GB RSS) and the upload cap is 100 MB, so the gate is what keeps one upload from taking the " + "ingest worker down. 0 disables it." + ), + ) + PDF_TRUST_CHECK: bool = Field( + default=True, + description=( + "Trust-check anydoc's PDF output (docsgpt/parser/file/pdf_trust.py): flag composite (Type0) fonts " + "without a ToUnicode map, and CJK-declaring PDFs whose extracted text has almost no CJK, the two " + "classes where anydoc drops text silently. A flagged file re-parses on the docling fallback when " + "docling is installed; otherwise the anydoc output is kept and the document gets " + 'extra_info["parse_warnings"]. ~30 ms per scanned MB.' + ), + ) + ANYDOC_TABLEIZE: bool = Field( + default=False, + description=( + "Rewrite dot-leader / whitespace-aligned table runs in anydoc's PDF markdown into GFM tables " + "(docsgpt/parser/file/tableize.py). Off by default: it rewrites content on a heuristic (>=3 uniform " + "label+numbers lines) validated only on a small corpus so far." + ), + ) + ATTACHMENT_PDF_TEXT_FAST_PATH: bool = Field( + default=True, + description=( + "Read PDF attachments via their embedded text layer (pypdfium2) instead of docling, falling back to " + "docling when there is no text layer. Attachments go into a prompt, so docling's structural " + "markdown earns far less than the tens of seconds per file it costs; source ingestion is " + "unaffected because chunking and retrieval do depend on that structure." + ), + ) + ATTACHMENT_PDF_TEXT_MIN_MEDIAN_CHARS: int = Field( + default=32, + description=( + "Median chars per sampled page below which a PDF attachment is treated as a scan and handed to " + "docling. Measured on real uploads: scans at 0-17 chars/page, text-layer documents at 433-6834." + ), + ) + ATTACHMENT_TEXT_MAX_BYTES: int = Field(default=5_000_000, description="Cap on extracted attachment text.") + AGENT_IMAGE_MAX_BYTES: int = Field(default=5_000_000, description="Cap on an image passed to an agent.") + AGENT_IMAGE_MAX_PIXELS: int = Field( + default=16_777_216, description="Cap on the pixel count of an image passed to an agent." + ) + GITHUB_INGEST_MAX_FILE_BYTES: int = Field( + default=1048576, ge=0, description="Skip GitHub repo blobs larger than this (0 = no cap)." + ) + GITHUB_INGEST_MAX_WORKERS: int = Field(default=8, ge=1, description="Parallel file fetches per GitHub repo ingest.") + + # read_document parsing on a dedicated Celery queue (backend parser). + DOCUMENT_PARSE_QUEUE: str = Field(default="parsing", description="Celery queue the parse_document task is routed to.") + DOCUMENT_PARSE_TIMEOUT: int = Field( + default=120, description="Seconds the read_document tool awaits the enqueued parse before degrading." + ) + DOCUMENT_PARSE_TIMEOUT_PER_MB: int = Field( + default=60, + description=( + "Extra seconds of parse window per MiB of input. The base timeout is a FLOOR: the window grows with " + "document size because OCR cost scales with pages. Without this a large scan is silently dropped at " + "the base window." + ), + ) + DOCUMENT_PARSE_TIMEOUT_MAX: int = Field( + default=900, description="Absolute ceiling on the size-scaled parse window, in seconds." + ) + DOCUMENT_PARSE_MAX_BYTES: int = Field( + default=0, ge=0, description="Cap on a parsed document's bytes (0 = reuse SANDBOX_MAX_INPUT_BYTES)." + ) + DOCUMENT_MAX_DECOMPRESSED_BYTES: int = Field( + default=300 * 1024 * 1024, description="Cap on bytes decompressed from an archive handed to read_document." + ) + DOCUMENT_MAX_ARCHIVE_ENTRIES: int = Field( + default=10000, description="Cap on entries in an archive handed to read_document." + ) + + @field_validator("DOC_PARSER_ENGINE", mode="before") + @classmethod + def _normalize_parser_engine(cls, v): + return normalize_choice(v) diff --git a/docsgpt/core/settings/llm.py b/docsgpt/core/settings/llm.py new file mode 100644 index 00000000..173c864a --- /dev/null +++ b/docsgpt/core/settings/llm.py @@ -0,0 +1,106 @@ +"""LLM providers, API keys and per-provider tunables.""" + +from __future__ import annotations + +import os +from typing import Optional + +from pydantic import Field + +from docsgpt.core.paths import home_dir +from docsgpt.core.settings._shared import SettingsGroup + + +class LLMSettings(SettingsGroup): + """Which model answers, how it is reached, and provider-specific behaviour.""" + + LLM_PROVIDER: str = Field(default="docsgpt", description="LLM provider key, e.g. openai, anthropic, docsgpt.") + LLM_NAME: Optional[str] = Field( + default=None, description="Model name for the provider; with openai, e.g. gpt-4 or gpt-3.5-turbo." + ) + API_KEY: Optional[str] = Field(default=None, description="LLM API key used by LLM_PROVIDER.") + + # Provider-specific API keys (for multi-model support). + OPENAI_API_KEY: Optional[str] = Field(default=None, description="OpenAI API key.") + ANTHROPIC_API_KEY: Optional[str] = Field(default=None, description="Anthropic API key.") + GOOGLE_API_KEY: Optional[str] = Field(default=None, description="Google AI API key.") + GROQ_API_KEY: Optional[str] = Field(default=None, description="Groq API key.") + HUGGINGFACE_API_KEY: Optional[str] = Field(default=None, description="Hugging Face API key.") + OPEN_ROUTER_API_KEY: Optional[str] = Field(default=None, description="OpenRouter API key.") + NOVITA_API_KEY: Optional[str] = Field(default=None, description="Novita API key.") + + OPENAI_API_BASE: Optional[str] = Field(default=None, description="Azure OpenAI API base URL.") + OPENAI_API_VERSION: Optional[str] = Field(default=None, description="Azure OpenAI API version.") + AZURE_DEPLOYMENT_NAME: Optional[str] = Field(default=None, description="Azure deployment name for answering.") + AZURE_EMBEDDINGS_DEPLOYMENT_NAME: Optional[str] = Field( + default=None, description="Azure deployment name for embeddings." + ) + OPENAI_BASE_URL: Optional[str] = Field( + default=None, description="Base URL for OpenAI-compatible model servers." + ) + LLM_PATH: str = Field( + default=os.path.join(str(home_dir()), "models/docsgpt-7b-f16.gguf"), + description="Path to the local GGUF model used by the llama.cpp provider.", + ) + + FALLBACK_LLM_PROVIDER: Optional[str] = Field(default=None, description="Provider for the fallback LLM.") + FALLBACK_LLM_NAME: Optional[str] = Field(default=None, description="Model name for the fallback LLM.") + FALLBACK_LLM_API_KEY: Optional[str] = Field(default=None, description="API key for the fallback LLM.") + TITLE_MODEL_ID: Optional[str] = Field( + default=None, description="Optional cheaper model for conversation titles; unset reuses the answer model." + ) + MODELS_CONFIG_DIR: Optional[str] = Field( + default=None, + description=( + "Directory of operator-supplied model YAMLs, loaded after the built-in catalog; later wins on " + "duplicate model id. See docsgpt/core/models/README.md." + ), + ) + DEFAULT_LLM_TOKEN_LIMIT: int = Field( + default=128000, description="Context window assumed when the model is not found in the registry." + ) + RESERVED_TOKENS: dict[str, int] = Field( + default={"system_prompt": 500, "current_query": 500, "safety_buffer": 1000}, + description="Tokens held back from the context window for the system prompt, the query and a safety buffer.", + ) + CACHE_REDIS_URL: str = Field(default="redis://localhost:6379/2", description="Redis URL for the LLM cache.") + + # OpenAI Responses API. + OPENAI_RESPONSES_STORE: bool = Field( + default=False, + description=( + "True persists Responses API calls server-side so previous_response_id can chain turns. False keeps " + "them stateless, carrying reasoning across the tool loop as encrypted items." + ), + ) + OPENAI_RESPONSES_CHAIN_ACROSS_TURNS: bool = Field( + default=True, + description=( + "Cross-turn previous_response_id chaining (store mode only). The chained transcript lives on the " + "provider and is invisible to every local guard, so it is bounded: a turn starts from the local " + "history when the previous turn's reported prompt already reached the budget (default: the model's " + "context window) or when the conversation was compressed after that turn was produced." + ), + ) + OPENAI_RESPONSES_CHAIN_BUDGET_TOKENS: Optional[int] = Field( + default=None, description="Prompt-token budget for cross-turn chaining; unset uses the model's context window." + ) + OPENAI_RESPONSES_TRUNCATION_AUTO: bool = Field( + default=False, + description=( + 'Send truncation: "auto" so the provider drops the oldest input items instead of failing every ' + "request once a chain exceeds the model's window." + ), + ) + OPENAI_PROMPT_CACHE_KEY: bool = Field( + default=True, + description=( + "Route a user's Responses API calls to the same prompt-cache shard with an opaque per-user key." + ), + ) + OPENAI_PROMPT_CACHE_RETENTION: Optional[str] = Field( + default=None, description="Request extended prompt-cache retention where the provider offers it." + ) + OPENAI_REASONING_SUMMARY: str = Field( + default="auto", description="Reasoning summary mode requested from the Responses API." + ) diff --git a/docsgpt/core/settings/ocr.py b/docsgpt/core/settings/ocr.py new file mode 100644 index 00000000..2ffc0432 --- /dev/null +++ b/docsgpt/core/settings/ocr.py @@ -0,0 +1,99 @@ +"""OCR for scanned PDFs and images.""" + +from __future__ import annotations + +from typing import Literal + +from pydantic import AliasChoices, Field, field_validator + +from docsgpt.core.settings._shared import SettingsGroup, normalize_choice + + +class OCRSettings(SettingsGroup): + """Whether OCR runs, which stack performs it, and which engine it uses. + + OCR_ENABLED covers source ingestion, OCR_ATTACHMENTS_ENABLED chat attachments. Which stack performs + it is OCR_BACKEND; which engine, OCR_ENGINE. The DOCLING_OCR_* names are the pre-2026-09 spellings + and stay accepted as aliases. + """ + + OCR_ENABLED: bool = Field( + default=False, + validation_alias=AliasChoices("OCR_ENABLED", "DOCLING_OCR_ENABLED"), + description="OCR scanned PDFs and images during source ingestion.", + ) + OCR_ATTACHMENTS_ENABLED: bool = Field( + default=False, + validation_alias=AliasChoices("OCR_ATTACHMENTS_ENABLED", "DOCLING_OCR_ATTACHMENTS_ENABLED"), + description="OCR scanned PDFs and images attached to a chat.", + ) + OCR_BACKEND: Literal["auto", "docling", "native"] = Field( + default="auto", + description=( + "Which stack runs OCR when it is on. auto: docling when installed, otherwise native. docling: the " + "layout-model pipeline (hybrid region OCR, reading order, table structure); needs the optional " + "docling extra. native: pypdfium2/Pillow page rendering straight into tesseract or a DeepSeek-OCR " + "endpoint (docsgpt/parser/file/ocr_parser.py); no ML models in the worker, tables come out as text " + "lines under tesseract." + ), + ) + OCR_ENGINE: Literal["tesseract", "deepseek", "auto", "ocrmac", "rapidocr"] = Field( + default="tesseract", + description=( + "OCR engine used when OCR is on. Benched 2026-08 on EN/ZH/table/degraded scans (docs/Guides/ocr has " + "the menu). tesseract (recommended): best classic-engine accuracy (perfect EN word recall, 0.000 " + "bilingual CER, 100% table cells), ~35 MB, CPU-only; needs the system binary and language packs, an " + "optional install like every OCR dependency (build with INSTALL_TESSERACT=true, or apt/brew install " + "tesseract-ocr for a local run); both backends. deepseek: DeepSeek-OCR against an Ollama/vLLM " + "endpoint (OCR_DEEPSEEK_*); best table/CJK quality, the worker stays light (no layout models) but " + "each page costs seconds on the model server; both backends. auto: docling's pick, ocrmac on macOS " + "(excellent), rapidocr on Linux (silently shreds some long text lines; avoid as a server default). " + "ocrmac | rapidocr: force one of those. auto/ocrmac/rapidocr exist only inside docling; the native " + "backend runs tesseract for them. An engine that is not installed degrades (docling: to auto) with " + "a warning instead of failing the parse." + ), + ) + OCR_LANGS: str = Field( + default="eng", + description=( + 'Tesseract language packs, "+"-separated (e.g. "eng+chi_sim+deu"). Other engines keep their own ' + "defaults; their language codes differ." + ), + ) + OCR_DEEPSEEK_URL: str = Field( + default="http://localhost:11434/v1/chat/completions", + description="Chat-completions URL of the DeepSeek-OCR endpoint (Ollama or vLLM).", + ) + OCR_DEEPSEEK_MODEL: str = Field(default="deepseek-ocr:3b", description="Model name at the DeepSeek-OCR endpoint.") + OCR_DEEPSEEK_TIMEOUT: float = Field( + default=300.0, + description=( + "Seconds allowed per page request to the DeepSeek endpoint, on both backends (native sends pages one " + "at a time; docling's VLM pipeline keeps its own concurrency). A 3B model on a laptop needs minutes; " + "a vLLM GPU deployment, seconds." + ), + ) + OCR_RENDER_DPI: int = Field( + default=200, + description=( + "Native backend only: resolution at which pages without a text layer are rendered before OCR. 200 " + "suits tesseract; clamped to 72-600." + ), + ) + OCR_MIN_CHARS_PER_PAGE: int = Field( + default=20, + ge=0, + validation_alias=AliasChoices("OCR_MIN_CHARS_PER_PAGE", "DOCLING_OCR_MIN_CHARS_PER_PAGE"), + description=( + "Chars-per-page floor below which an OCR'd PDF/image parse is treated as an OCR dropout rather than " + "as content (long-running docling workers were observed returning zero characters for every " + "scanned page after a long scanned PDF, with no error). docling retries once on a fresh full-page-OCR " + "converter; both backends then fail loudly instead of indexing an empty document. 0 disables the " + "guard." + ), + ) + + @field_validator("OCR_BACKEND", "OCR_ENGINE", mode="before") + @classmethod + def _normalize_ocr_choices(cls, v): + return normalize_choice(v) diff --git a/docsgpt/core/settings/reference.py b/docsgpt/core/settings/reference.py new file mode 100644 index 00000000..c471ee0f --- /dev/null +++ b/docsgpt/core/settings/reference.py @@ -0,0 +1,172 @@ +"""Render the settings reference page from the ``Settings`` definitions. + +The page under ``docs/content/Deploying/Settings-Reference.mdx`` is generated +from the field types, defaults and descriptions in this package, so the model +is the single source of truth. Regenerate it after changing a setting:: + + python -m docsgpt.core.settings.reference --write + +``--check`` exits non-zero when the checked-in page is stale; the test suite +runs the same comparison. +""" + +from __future__ import annotations + +import argparse +import inspect +import json +import sys +import types +import typing +from pathlib import Path +from typing import Any, Literal, Optional, Union + +from pydantic import AliasChoices +from pydantic.fields import FieldInfo + +from docsgpt.core.paths import home_dir +from docsgpt.core.settings import SETTINGS_GROUPS, Settings + +REFERENCE_PATH = Path("docs") / "content" / "Deploying" / "Settings-Reference.mdx" + +HOME_PLACEHOLDER = "" + +_HEADER = """\ +--- +title: Settings Reference +description: Every DocsGPT setting, grouped by domain, with its type, default and purpose. +--- + +{/* GENERATED FILE. Do not edit by hand: run `python -m docsgpt.core.settings.reference --write`. */} + +# Settings Reference + +Every setting DocsGPT reads, generated from `docsgpt/core/settings/`. Each one is +an environment variable of the same name, set in `.env` or the process +environment; see [App Configuration](/Deploying/DocsGPT-Settings) for how the +file is found and for worked examples. `` below is the data home +described there. +""" + + +def _mdx(text: str) -> str: + """Escape prose for MDX, where braces open expressions and ``<`` opens JSX.""" + return text.replace("{", "\\{").replace("}", "\\}").replace("<", "<") + + +def _type_name(annotation: Any) -> str: + origin = typing.get_origin(annotation) + if origin in (Union, types.UnionType): + args = [a for a in typing.get_args(annotation) if a is not type(None)] + return " | ".join(_type_name(a) for a in args) + if origin is Literal: + return " | ".join(json.dumps(v) for v in typing.get_args(annotation)) + if origin is not None: + name = getattr(origin, "__name__", str(origin)) + args = typing.get_args(annotation) + return f"{name}[{', '.join(_type_name(arg) for arg in args)}]" if args else name + return getattr(annotation, "__name__", str(annotation)) + + +def _default_text(field: FieldInfo) -> str: + value = field.default_factory() if field.default_factory is not None else field.default + if value is None: + return "unset" + if isinstance(value, bool): + return "`true`" if value else "`false`" + if isinstance(value, str): + value = value.replace(str(home_dir()), HOME_PLACEHOLDER) + return "`\"\"`" if value == "" else f"`{value}`" + if isinstance(value, (list, dict)): + return f"`{json.dumps(value)}`" + return f"`{value}`" + + +def _constraints(field: FieldInfo) -> list[str]: + out = [] + for item in field.metadata: + for attr, symbol in (("gt", ">"), ("ge", ">="), ("lt", "<"), ("le", "<=")): + if hasattr(item, attr): + out.append(f"{symbol} {getattr(item, attr)}") + return out + + +def _aliases(name: str, field: FieldInfo) -> list[str]: + alias = field.validation_alias + if isinstance(alias, AliasChoices): + return [str(c) for c in alias.choices if str(c) != name] + if isinstance(alias, str) and alias != name: + return [alias] + return [] + + +def _render_field(name: str, field: FieldInfo) -> str: + facts = [f"Type `{_type_name(field.annotation)}`", f"default {_default_text(field)}"] + constraints = _constraints(field) + if constraints: + # Code spans: a bare ``<=`` in MDX prose is parsed as the start of a JSX tag. + facts.append("must be " + " and ".join(f"`{c}`" for c in constraints)) + aliases = _aliases(name, field) + if aliases: + facts.append("also read from " + ", ".join(f"`{a}`" for a in aliases)) + # The facts are code spans, which MDX leaves alone; only prose needs escaping. + lines = [f"### `{name}`", "", ", ".join(facts) + "."] + if field.deprecated: + note = field.deprecated if isinstance(field.deprecated, str) else "This setting is deprecated." + lines += ["", f"**Deprecated.** {_mdx(str(note))}"] + if field.description: + lines += ["", _mdx(field.description)] + return "\n".join(lines) + + +def _group_intro(group: type) -> str: + doc = inspect.getdoc(group) or "" + return doc.split("\n\n", 1)[0].replace("\n", " ").strip() + + +def render_reference() -> str: + """The full reference page as MDX text.""" + parts = [_HEADER] + for title, group in SETTINGS_GROUPS: + parts.append(f"\n## {title}\n") + intro = _group_intro(group) + if intro: + parts.append(_mdx(intro) + "\n") + for name in group.model_fields: + parts.append(_render_field(name, Settings.model_fields[name]) + "\n") + return "\n".join(parts).rstrip("\n") + "\n" + + +def reference_path(root: Optional[Path] = None) -> Path: + """Where the generated page lives in a checkout; ``root`` defaults to the repository root.""" + if root is None: + root = Path(__file__).resolve().parents[3] + return root / REFERENCE_PATH + + +def main(argv: Optional[list[str]] = None) -> int: + parser = argparse.ArgumentParser(description=__doc__.split("\n\n", 1)[0]) + action = parser.add_mutually_exclusive_group() + action.add_argument("--write", action="store_true", help="write the page into the docs tree") + action.add_argument("--check", action="store_true", help="exit 1 if the checked-in page is stale") + args = parser.parse_args(argv) + + rendered = render_reference() + path = reference_path() + if args.write: + path.write_text(rendered, encoding="utf-8") + print(f"wrote {path}") + return 0 + if args.check: + current = path.read_text(encoding="utf-8") if path.exists() else "" + if current != rendered: + print(f"{path} is stale; run: python -m docsgpt.core.settings.reference --write", file=sys.stderr) + return 1 + print(f"{path} is up to date") + return 0 + sys.stdout.write(rendered) + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/docsgpt/core/settings/retrieval.py b/docsgpt/core/settings/retrieval.py new file mode 100644 index 00000000..651d3f3d --- /dev/null +++ b/docsgpt/core/settings/retrieval.py @@ -0,0 +1,38 @@ +"""Retrieval strategy and GraphRAG.""" + +from __future__ import annotations + +from typing import Literal, Optional + +from pydantic import Field, field_validator + +from docsgpt.core.settings._shared import SettingsGroup, normalize_choice + + +class RetrievalSettings(SettingsGroup): + """Which vector store answers searches and how retrieval fans out across sources.""" + + VECTOR_STORE: Literal["faiss", "elasticsearch", "mongodb", "qdrant", "milvus", "pgvector"] = Field( + default="faiss", description="Vector store backend." + ) + RETRIEVAL_MAX_PARALLEL_SOURCES: int = Field( + default=4, + ge=1, + description="Concurrent per-source searches in one retrieval; the query is embedded once and shared.", + ) + PER_SOURCE_RETRIEVAL_ENABLED: bool = Field( + default=True, + description="Kill-switch for per-source retrieval dispatch; False collapses to a single retriever.", + ) + GRAPHRAG_ENABLED: bool = Field(default=False, description="Gates graph-aware ingestion and retrieval.") + GRAPHRAG_EXTRACTION_MODEL: Optional[str] = Field( + default=None, description="Model for ingest-time graph extraction; unset reuses LLM_PROVIDER/LLM_NAME." + ) + GRAPHRAG_MAX_CHUNKS_FOR_EXTRACTION: int = Field( + default=2000, ge=0, description="Hard cap on chunks extracted per source (cost control); 0 extracts nothing." + ) + + @field_validator("VECTOR_STORE", mode="before") + @classmethod + def _normalize_vector_store(cls, v): + return normalize_choice(v) diff --git a/docsgpt/core/settings/sandbox.py b/docsgpt/core/settings/sandbox.py new file mode 100644 index 00000000..96e4d6e0 --- /dev/null +++ b/docsgpt/core/settings/sandbox.py @@ -0,0 +1,93 @@ +"""Code-execution sandbox: the Jupyter gateway runner or Daytona Cloud.""" + +from __future__ import annotations + +from typing import Literal, Optional + +from pydantic import Field, field_validator + +from docsgpt.core.settings._shared import SettingsGroup, normalize_choice + + +class SandboxSettings(SettingsGroup): + """The app is a CLIENT of an always-on runner; defaults are safe so app import never fails unconfigured.""" + + SANDBOX_BACKEND: Literal["jupyter", "daytona"] = Field( + default="jupyter", description="Sandbox backend: jupyter (self-host) or daytona (Daytona Cloud)." + ) + SANDBOX_GATEWAY_URL: str = Field( + default="http://localhost:8888", + description="URL of the Jupyter Kernel Gateway runner (the docsgpt-sandbox service).", + ) + SANDBOX_GATEWAY_AUTH_TOKEN: Optional[str] = Field(default=None, description="Gateway auth token, if set.") + SANDBOX_KERNEL_NAME: str = Field( + default="docsgpt-python", + description=( + "Kernelspec per session. The env-scrubbing docsgpt-python spec keeps kernel code from reading the " + "gateway token or operator secrets from os.environ; the stock python3 spec inherits the gateway env " + "verbatim and must not be used with untrusted code." + ), + ) + SANDBOX_MAX_TTL: int = Field(default=1200, description="Hard cap (s) on agent-selectable keep-alive TTL.") + SANDBOX_MAX_SESSIONS: int = Field( + default=32, + description=( + "Concurrent live sessions per process, backend-agnostic; at the cap an LRU-idle session is evicted. " + "0 or negative disables the cap." + ), + ) + SANDBOX_EXEC_TIMEOUT: int = Field(default=60, description="Default wall-clock cap (s) per exec call.") + SANDBOX_HTTP_TIMEOUT: int = Field( + default=10, description="Fixed cap (s) for REST control calls (create/delete/alive/interrupt)." + ) + SANDBOX_MAX_OUTPUT_BYTES: int = Field( + default=8 * 1024 * 1024, description="Cap on buffered stdout+stderr per exec." + ) + SANDBOX_MAX_FILE_BYTES: int = Field( + default=10 * 1024 * 1024, description="Cap on get_file size routed through stdout." + ) + SANDBOX_MAX_INPUT_BYTES: int = Field( + default=25 * 1024 * 1024, description="Cap on an input document staged into a sandbox session." + ) + # Runner container caps, consumed by the docsgpt-sandbox compose service, not the app. + SANDBOX_MEMORY: str = Field( + default="1g", + description=( + "Docker mem_limit for the runner container. Consumed by the docsgpt-sandbox compose service, not " + "the app; part of the untrusted-code security boundary." + ), + ) + SANDBOX_CPUS: str = Field( + default="1.0", + description=( + "Docker CPU quota for the runner container. Consumed by the docsgpt-sandbox compose service, not " + "the app; part of the untrusted-code security boundary." + ), + ) + + # Daytona Cloud backend (SANDBOX_BACKEND=daytona). All knobs are optional so app import never fails + # when the backend is unused. + DAYTONA_API_KEY: Optional[str] = Field(default=None, description="Daytona Cloud API key (secret).") + DAYTONA_API_URL: Optional[str] = Field( + default=None, description="Override for the Daytona API base URL, if self-targeting." + ) + DAYTONA_TARGET: Optional[str] = Field(default=None, description='Daytona region/target, e.g. "us".') + DAYTONA_SNAPSHOT: Optional[str] = Field( + default=None, + description="Image for new sandboxes; render libs via scripts/build_daytona_snapshot.py.", + ) + DAYTONA_LANGUAGE: str = Field(default="python", description="Default runtime language for created sandboxes.") + DAYTONA_AUTO_STOP_INTERVAL: int = Field( + default=15, ge=0, description="Minutes idle before Daytona auto-stops a sandbox (0 disables)." + ) + DAYTONA_AUTO_DELETE_INTERVAL: int = Field( + default=60, ge=-1, description="Minutes after stop before Daytona auto-deletes a sandbox (-1 disables)." + ) + DAYTONA_MAX_SANDBOXES: int = Field( + default=50, description="Cap on concurrent live Daytona sandboxes (cost-DoS guard)." + ) + + @field_validator("SANDBOX_BACKEND", mode="before") + @classmethod + def _normalize_sandbox_backend(cls, v): + return normalize_choice(v) diff --git a/docsgpt/core/settings/scheduler.py b/docsgpt/core/settings/scheduler.py new file mode 100644 index 00000000..b86d9995 --- /dev/null +++ b/docsgpt/core/settings/scheduler.py @@ -0,0 +1,28 @@ +"""Scheduled agent runs (see scheduler.md).""" + +from __future__ import annotations + +from pydantic import Field + +from docsgpt.core.settings._shared import SettingsGroup + + +class SchedulerSettings(SettingsGroup): + """Cadence, quotas and timeouts of scheduled runs.""" + + SCHEDULE_DISPATCHER_INTERVAL: int = Field( + default=30, description="Seconds between dispatcher passes that enqueue due schedules." + ) + SCHEDULE_MIN_INTERVAL: int = Field(default=900, description="Smallest allowed recurrence interval in seconds.") + SCHEDULE_MAX_PER_USER: int = Field(default=50, description="Cap on schedules a user may own.") + SCHEDULE_RUN_TIMEOUT: int = Field(default=600, description="Wall-clock cap on one scheduled run, in seconds.") + SCHEDULE_MISFIRE_GRACE: int = Field( + default=60, description="Seconds past the due time within which a missed run still fires." + ) + SCHEDULE_AUTOPAUSE_FAILURES: int = Field( + default=3, description="Consecutive failures after which a schedule is paused automatically." + ) + SCHEDULE_ONCE_MAX_HORIZON: int = Field( + default=31_536_000, description="How far ahead a one-off run may be scheduled, in seconds (one year)." + ) + SCHEDULE_RUN_OUTPUT_RETENTION_DAYS: int = Field(default=90, gt=0, description="Days scheduled-run output is kept.") diff --git a/docsgpt/core/settings/server.py b/docsgpt/core/settings/server.py new file mode 100644 index 00000000..407ef4aa --- /dev/null +++ b/docsgpt/core/settings/server.py @@ -0,0 +1,46 @@ +"""The API process itself.""" + +from __future__ import annotations + +from typing import Optional + +from pydantic import Field + +from docsgpt.core.settings._shared import SettingsGroup + + +class ServerSettings(SettingsGroup): + """Serving the UI, public URLs, and process-level knobs of the API server.""" + + DEPLOYMENT_TYPE: Optional[str] = Field( + default=None, + description=( + "Deployment class, e.g. cloud or production. A production class refuses to run without a " + "configured JWT_SECRET_KEY instead of generating a local one on disk." + ), + ) + SERVE_UI: bool = Field( + default=True, description="Serve the web UI shipped in the package (docsgpt/static) from the API process." + ) + FLASK_DEBUG_MODE: bool = Field(default=False, description="Run Flask in debug mode.") + VERSION_CHECK: bool = Field(default=True, description="Anonymous startup version check for security issues.") + PUBLIC_API_BASE_URL: Optional[str] = Field( + default=None, description="Public base URL for user-facing endpoint references in prompts." + ) + GRACEFUL_SHUTDOWN_TIMEOUT_SECONDS: int = Field( + default=30, + description=( + "Bounds uvicorn's shutdown drain (uvicorn_worker doesn't forward --graceful-timeout). Keep below the " + "gunicorn --timeout (180) watchdog. Used by BoundedDrainUvicornWorker." + ), + ) + WSGI_THREADPOOL_WORKERS: int = Field( + default=96, ge=1, description="Threads serving the WSGI (Flask) part of the app under the ASGI server." + ) + V1_SESSION_TTL_SECONDS: int = Field( + default=24 * 60 * 60, + description=( + "Lets OpenAI-compatible clients identify a logical chat by session header, which chat-completions " + "itself has no field for; TTL of that session mapping." + ), + ) diff --git a/docsgpt/core/settings/speech.py b/docsgpt/core/settings/speech.py new file mode 100644 index 00000000..f25380b0 --- /dev/null +++ b/docsgpt/core/settings/speech.py @@ -0,0 +1,33 @@ +"""Text-to-speech and speech-to-text.""" + +from __future__ import annotations + +from typing import Literal, Optional + +from pydantic import Field, field_validator + +from docsgpt.core.settings._shared import SettingsGroup, normalize_choice + + +class SpeechSettings(SettingsGroup): + """Voice providers and transcription options.""" + + TTS_PROVIDER: Literal["google_tts", "elevenlabs", "none"] = Field( + default="google_tts", description="Text-to-speech provider; none switches it off." + ) + ELEVENLABS_API_KEY: Optional[str] = Field(default=None, description="ElevenLabs API key.") + STT_PROVIDER: Literal["openai", "faster_whisper", "none"] = Field( + default="openai", description="Speech-to-text provider; none switches it off." + ) + OPENAI_STT_MODEL: str = Field(default="gpt-4o-mini-transcribe", description="OpenAI transcription model.") + STT_LANGUAGE: Optional[str] = Field(default=None, description="Language hint for transcription; unset auto-detects.") + STT_MAX_FILE_SIZE_MB: int = Field(default=50, description="Cap on an audio file accepted for transcription.") + STT_ENABLE_TIMESTAMPS: bool = Field(default=False, description="Return word/segment timestamps.") + STT_ENABLE_DIARIZATION: bool = Field(default=False, description="Label speakers in the transcript.") + + @field_validator("TTS_PROVIDER", "STT_PROVIDER", mode="before") + @classmethod + def _normalize_speech_providers(cls, v): + # An empty value has always meant "off"; keep that spelling working. + v = normalize_choice(v) + return "none" if v == "" else v diff --git a/docsgpt/core/settings/storage.py b/docsgpt/core/settings/storage.py new file mode 100644 index 00000000..0212c050 --- /dev/null +++ b/docsgpt/core/settings/storage.py @@ -0,0 +1,57 @@ +"""Where uploaded files and generated artifacts are stored.""" + +from __future__ import annotations + +from typing import Literal, Optional + +from pydantic import Field, field_validator + +from docsgpt.core.settings._shared import SettingsGroup, normalize_choice + + +class StorageSettings(SettingsGroup): + """Local disk or an S3-compatible bucket, and how download URLs are produced.""" + + STORAGE_TYPE: Literal["local", "s3"] = Field(default="local", description="File storage backend.") + URL_STRATEGY: Literal["backend", "s3"] = Field( + default="backend", + description="How download links are produced: backend (streamed through the API) or s3 (presigned URLs).", + ) + + # S3-compatible object storage (STORAGE_TYPE=s3): AWS S3, MinIO, R2, B2, Spaces, ... + # For non-AWS, set S3_ENDPOINT_URL and usually S3_PATH_STYLE=true. + S3_BUCKET_NAME: str = Field(default="docsgpt-test-bucket", description="Bucket name.") + S3_ENDPOINT_URL: Optional[str] = Field( + default=None, description="Custom endpoint for S3-compatible services (MinIO, R2, B2, Spaces); omit for AWS." + ) + S3_ACCESS_KEY_ID: Optional[str] = Field(default=None, description="Access key id.") + S3_SECRET_ACCESS_KEY: Optional[str] = Field(default=None, description="Secret access key.") + S3_REGION: Optional[str] = Field(default=None, description='AWS region; use "auto" for Cloudflare R2.') + S3_PATH_STYLE: bool = Field( + default=False, description="Path-style addressing (required by most non-AWS services)." + ) + + # Legacy AWS credentials from the retired SageMaker provider. + SAGEMAKER_REGION: Optional[str] = Field( + default=None, + deprecated="Set S3_REGION instead; the SAGEMAKER_* fallback will be removed.", + description="Legacy AWS region from the retired SageMaker provider; deprecated fallback for S3_REGION.", + ) + SAGEMAKER_ACCESS_KEY: Optional[str] = Field( + default=None, + deprecated="Set S3_ACCESS_KEY_ID instead; the SAGEMAKER_* fallback will be removed.", + description="Legacy AWS access key from the retired SageMaker provider; deprecated fallback for S3_ACCESS_KEY_ID.", + ) + SAGEMAKER_SECRET_KEY: Optional[str] = Field( + default=None, + deprecated="Set S3_SECRET_ACCESS_KEY instead; the SAGEMAKER_* fallback will be removed.", + description=( + "Legacy AWS secret key from the retired SageMaker provider; deprecated fallback for " + "S3_SECRET_ACCESS_KEY." + ), + ) + + @field_validator("STORAGE_TYPE", "URL_STRATEGY", mode="before") + @classmethod + def _normalize_storage_choices(cls, v): + return normalize_choice(v) diff --git a/docsgpt/core/settings/vectorstores.py b/docsgpt/core/settings/vectorstores.py new file mode 100644 index 00000000..010eefa3 --- /dev/null +++ b/docsgpt/core/settings/vectorstores.py @@ -0,0 +1,84 @@ +"""Connection settings for each vector store backend.""" + +from __future__ import annotations + +from typing import Optional + +from pydantic import Field, field_validator + +from docsgpt.core.db_uri import normalize_pgvector_connection_string +from docsgpt.core.paths import home_dir +from docsgpt.core.settings._shared import SettingsGroup + + +class VectorStoreSettings(SettingsGroup): + """Per-backend connection details; only the backend named by VECTOR_STORE is read.""" + + MONGO_URI: Optional[str] = Field( + default=None, + description=( + "Only consulted when VECTOR_STORE=mongodb or when running scripts/db/backfill.py; user data lives " + "in Postgres." + ), + ) + + # Elasticsearch. + ELASTIC_CLOUD_ID: Optional[str] = Field(default=None, description="Elastic Cloud id.") + ELASTIC_USERNAME: Optional[str] = Field(default=None, description="Elasticsearch username.") + ELASTIC_PASSWORD: Optional[str] = Field(default=None, description="Elasticsearch password.") + ELASTIC_URL: Optional[str] = Field(default=None, description="Elasticsearch URL.") + ELASTIC_INDEX: str = Field(default="docsgpt", description="Elasticsearch index name.") + + # Qdrant. + QDRANT_COLLECTION_NAME: str = Field(default="docsgpt", description="Qdrant collection name.") + QDRANT_LOCATION: Optional[str] = Field(default=None, description="Qdrant location (':memory:' or a URL).") + QDRANT_URL: Optional[str] = Field(default=None, description="Qdrant server URL.") + QDRANT_PORT: int = Field(default=6333, description="Qdrant REST port.") + QDRANT_GRPC_PORT: int = Field(default=6334, description="Qdrant gRPC port.") + QDRANT_PREFER_GRPC: bool = Field(default=False, description="Use gRPC instead of REST where possible.") + QDRANT_HTTPS: Optional[bool] = Field(default=None, description="Use HTTPS for the Qdrant connection.") + QDRANT_API_KEY: Optional[str] = Field(default=None, description="Qdrant API key.") + QDRANT_PREFIX: Optional[str] = Field(default=None, description="URL prefix for a Qdrant behind a proxy.") + QDRANT_TIMEOUT: Optional[float] = Field(default=None, description="Qdrant request timeout in seconds.") + QDRANT_HOST: Optional[str] = Field(default=None, description="Qdrant host (alternative to QDRANT_URL).") + QDRANT_PATH: Optional[str] = Field(default=None, description="Path for an embedded on-disk Qdrant.") + QDRANT_DISTANCE_FUNC: str = Field(default="Cosine", description="Qdrant distance function.") + + # PGVector. + PGVECTOR_CONNECTION_STRING: Optional[str] = Field( + default=None, + description=( + "pgvector connection string. postgres://, postgresql:// and postgresql+psycopg:// are all accepted " + "and normalized internally for psycopg.connect(). Unset falls back to POSTGRES_URI." + ), + ) + PGVECTOR_POOL_MAX_SIZE: int = Field( + default=8, ge=0, description="Per-process connection pool size; 0 uses one direct connection per store." + ) + PGVECTOR_IVFFLAT_PROBES: Optional[int] = Field( + default=None, + description="IVFFlat probes; unset derives sqrt(lists) from the index. Higher means better recall, more scan.", + ) + + # Milvus. + MILVUS_COLLECTION_NAME: str = Field(default="docsgpt", description="Milvus collection name.") + MILVUS_URI: Optional[str] = Field( + default_factory=lambda: str(home_dir() / "milvus_local.db"), + description=( + "Milvus server URI. The default is a milvus-lite (embedded) database file under the data home, " + "like the other local stores." + ), + ) + MILVUS_TOKEN: str = Field(default="", description="Milvus auth token.") + + # LanceDB. + LANCEDB_PATH: str = Field( + default_factory=lambda: str(home_dir() / "data" / "lancedb"), + description="LanceDB local data directory.", + ) + LANCEDB_TABLE_NAME: str = Field(default="docsgpts", description="LanceDB table for stored vectors.") + + @field_validator("PGVECTOR_CONNECTION_STRING", mode="before") + @classmethod + def _normalize_pgvector_connection_string(cls, v): + return normalize_pgvector_connection_string(v) diff --git a/docsgpt/core/settings/workers.py b/docsgpt/core/settings/workers.py new file mode 100644 index 00000000..96efc8f2 --- /dev/null +++ b/docsgpt/core/settings/workers.py @@ -0,0 +1,39 @@ +"""Celery broker, result backend and worker process limits.""" + +from __future__ import annotations + +from pydantic import Field + +from docsgpt.core.settings._shared import SettingsGroup + + +class WorkerSettings(SettingsGroup): + """How background tasks are queued and how worker processes are recycled.""" + + CELERY_BROKER_URL: str = Field(default="redis://localhost:6379/0", description="Celery broker URL.") + CELERY_RESULT_BACKEND: str = Field(default="redis://localhost:6379/1", description="Celery result backend URL.") + CELERY_WORKER_PREFETCH_MULTIPLIER: int = Field( + default=1, description="Tasks prefetched per worker process; 1 caps SIGKILL loss to one task." + ) + CELERY_VISIBILITY_TIMEOUT: int = Field( + default=3600, + gt=0, + description=( + "Broker visibility timeout in seconds. Must exceed the longest legitimate task runtime but stay " + "short enough that SIGKILLed tasks redeliver promptly." + ), + ) + CELERY_WORKER_MAX_MEMORY_PER_CHILD: int = Field( + default=4194304, + ge=0, + description=( + "Recycle a prefork child past this resident size in KB; backstops docling/torch heap growth. " + "Checked between tasks, so it does not bound the peak within one. 0 disables." + ), + ) + CELERY_WORKER_MAX_TASKS_PER_CHILD: int = Field( + default=0, ge=0, description="Recycle a worker child after N tasks; 0 disables." + ) + API_URL: str = Field( + default="http://localhost:7091", description="Backend URL the Celery worker calls back into." + ) diff --git a/docsgpt/devices/broker.py b/docsgpt/devices/broker.py index 09d55360..79ac2287 100644 --- a/docsgpt/devices/broker.py +++ b/docsgpt/devices/broker.py @@ -631,15 +631,15 @@ class DeviceBroker: @staticmethod def _inv_ttl() -> int: - return int(getattr(settings, "REMOTE_DEVICE_INVOCATION_TTL_SECONDS", 900)) + return int(settings.REMOTE_DEVICE_INVOCATION_TTL_SECONDS) @staticmethod def _cmd_ttl() -> int: - return int(getattr(settings, "REMOTE_DEVICE_CMD_QUEUE_TTL_SECONDS", 900)) + return int(settings.REMOTE_DEVICE_CMD_QUEUE_TTL_SECONDS) @staticmethod def _out_maxlen() -> int: - return int(getattr(settings, "REMOTE_DEVICE_OUTPUT_STREAM_MAXLEN", 10_000)) + return int(settings.REMOTE_DEVICE_OUTPUT_STREAM_MAXLEN) def _to_int(value: Optional[str]) -> Optional[int]: diff --git a/docsgpt/graphrag/store.py b/docsgpt/graphrag/store.py index e9ee39b9..818004c0 100644 --- a/docsgpt/graphrag/store.py +++ b/docsgpt/graphrag/store.py @@ -77,11 +77,9 @@ class GraphStore: """Stores and queries a per-source knowledge graph in the pgvector DB.""" def __init__(self, connection_string: Optional[str] = None): - self._connection_string = connection_string or getattr( - settings, "PGVECTOR_CONNECTION_STRING", None - ) + self._connection_string = connection_string or settings.PGVECTOR_CONNECTION_STRING - if not self._connection_string and getattr(settings, "POSTGRES_URI", None): + if not self._connection_string and settings.POSTGRES_URI: from docsgpt.core.db_uri import normalize_pgvector_connection_string self._connection_string = normalize_pgvector_connection_string( diff --git a/docsgpt/guardrails/guardrail_creator.py b/docsgpt/guardrails/guardrail_creator.py index d891e065..f3357f67 100644 --- a/docsgpt/guardrails/guardrail_creator.py +++ b/docsgpt/guardrails/guardrail_creator.py @@ -52,7 +52,7 @@ class GuardrailCreator: does not require an operator to also edit their env. """ cls._ensure_builtin() - allowlist = getattr(settings, "GUARDRAILS_CHECKS_ENABLED", None) or [] + allowlist = settings.GUARDRAILS_CHECKS_ENABLED or [] if not allowlist: return sorted(cls.checks) return sorted(k for k in cls.checks if k in set(allowlist)) diff --git a/docsgpt/guardrails/runtime.py b/docsgpt/guardrails/runtime.py index 16c20638..3c73a53b 100644 --- a/docsgpt/guardrails/runtime.py +++ b/docsgpt/guardrails/runtime.py @@ -38,7 +38,7 @@ def _merge_mode(agent_mode: str, floor_mode: str) -> str: def instance_floor() -> Optional[GuardrailsConfig]: """The operator-set minimum, or None when unset/invalid.""" - raw = getattr(settings, "GUARDRAILS_FLOOR", None) + raw = settings.GUARDRAILS_FLOOR if not raw: return None try: @@ -104,7 +104,7 @@ def floor_keys() -> set: def resolve_config(raw_agent_config: Optional[dict]) -> GuardrailsConfig: """Parse ``agents.config`` and apply the instance floor.""" - if not getattr(settings, "GUARDRAILS_ENABLED", True): + if not settings.GUARDRAILS_ENABLED: return GuardrailsConfig() agent = AgentConfig.parse(raw_agent_config).guardrails return merge_floor(agent, instance_floor()) @@ -123,7 +123,7 @@ def _judge_factory(agent): decoded_token=agent.decoded_token, model_id=( model_override - or getattr(settings, "GUARDRAILS_JUDGE_MODEL", None) + or settings.GUARDRAILS_JUDGE_MODEL or agent.upstream_model_id ), agent_id=agent.agent_id, @@ -169,7 +169,7 @@ class GuardrailRecorder: self._seen: set = set() def __call__(self, decision: StageDecision) -> None: - store_text = bool(getattr(settings, "GUARDRAILS_STORE_SCANNED_TEXT", False)) + store_text = bool(settings.GUARDRAILS_STORE_SCANNED_TEXT) for verdict in decision.verdicts: if not verdict.outcome.triggered and verdict.outcome.evaluated: continue diff --git a/docsgpt/llm/handlers/base.py b/docsgpt/llm/handlers/base.py index 0cd4c84c..039d7007 100644 --- a/docsgpt/llm/handlers/base.py +++ b/docsgpt/llm/handlers/base.py @@ -35,7 +35,7 @@ def _bound_tool_response_for_llm(tool_response: Any) -> Any: from docsgpt.core.settings import settings from docsgpt.utils import num_tokens_from_string - max_tokens = int(getattr(settings, "TOOL_RESULT_MAX_TOKENS", 20000) or 0) + max_tokens = int(settings.TOOL_RESULT_MAX_TOKENS or 0) if max_tokens <= 0: return tool_response text = tool_response if isinstance(tool_response, str) else str(tool_response) diff --git a/docsgpt/llm/openai.py b/docsgpt/llm/openai.py index a901f0a1..ad2f1f4b 100644 --- a/docsgpt/llm/openai.py +++ b/docsgpt/llm/openai.py @@ -1307,14 +1307,14 @@ class OpenAILLM(BaseLLM): params["include"] = ["reasoning.encrypted_content"] # Backstop against a chain that outgrows the model's native window: # the provider drops the oldest input items instead of failing. - if getattr(settings, "OPENAI_RESPONSES_TRUNCATION_AUTO", False): + if settings.OPENAI_RESPONSES_TRUNCATION_AUTO: params["truncation"] = "auto" # Prompt-cache hints. The key pins a conversation to one cache shard; # retention asks for the extended tier where the deployment offers it. cache_key = getattr(self, "_prompt_cache_key", None) - if cache_key and getattr(settings, "OPENAI_PROMPT_CACHE_KEY", False): + if cache_key and settings.OPENAI_PROMPT_CACHE_KEY: params["prompt_cache_key"] = str(cache_key) - retention = getattr(settings, "OPENAI_PROMPT_CACHE_RETENTION", None) + retention = settings.OPENAI_PROMPT_CACHE_RETENTION if retention: params["prompt_cache_retention"] = retention return params diff --git a/docsgpt/parser/connectors/share_point/auth.py b/docsgpt/parser/connectors/share_point/auth.py index 9a4264f6..ec006740 100644 --- a/docsgpt/parser/connectors/share_point/auth.py +++ b/docsgpt/parser/connectors/share_point/auth.py @@ -41,7 +41,7 @@ class SharePointAuth(BaseConnectorAuth): self.redirect_uri = settings.CONNECTOR_REDIRECT_BASE_URI self.tenant_id = settings.MICROSOFT_TENANT_ID - self.authority = getattr(settings, "MICROSOFT_AUTHORITY", f"https://login.microsoftonline.com/{self.tenant_id}") + self.authority = settings.MICROSOFT_AUTHORITY or f"https://login.microsoftonline.com/{self.tenant_id}" self.auth_app = ConfidentialClientApplication( client_id=self.client_id, diff --git a/docsgpt/parser/document_reader.py b/docsgpt/parser/document_reader.py index 3a031d19..ed3446e7 100644 --- a/docsgpt/parser/document_reader.py +++ b/docsgpt/parser/document_reader.py @@ -97,10 +97,10 @@ def bound_parse_payload(payload: Dict[str, Any], max_chars: Optional[int] = None def _max_input_bytes() -> int: """Return the size cap for a parsed document (its own setting, else the sandbox cap).""" - explicit = int(getattr(settings, "DOCUMENT_PARSE_MAX_BYTES", 0) or 0) + explicit = int(settings.DOCUMENT_PARSE_MAX_BYTES or 0) if explicit > 0: return explicit - return int(getattr(settings, "SANDBOX_MAX_INPUT_BYTES", 25 * 1024 * 1024)) + return int(settings.SANDBOX_MAX_INPUT_BYTES) # Every zip-packaged format a parser map can route: OOXML and its macro/ @@ -127,8 +127,8 @@ def _zip_bomb_reason(source: Union[bytes, str, Path], suffix: str) -> Optional[s """ if suffix not in _ZIP_CONTAINER_EXTENSIONS: return None - max_entries = int(getattr(settings, "DOCUMENT_MAX_ARCHIVE_ENTRIES", 10000)) - cap = int(getattr(settings, "DOCUMENT_MAX_DECOMPRESSED_BYTES", 300 * 1024 * 1024)) + max_entries = int(settings.DOCUMENT_MAX_ARCHIVE_ENTRIES) + cap = int(settings.DOCUMENT_MAX_DECOMPRESSED_BYTES) opened = io.BytesIO(source) if isinstance(source, bytes) else source try: with zipfile.ZipFile(opened) as zf: @@ -169,14 +169,14 @@ def _resolve_ocr_enabled(ocr: str) -> bool: return True if ocr == "off": return False - return bool(getattr(settings, "OCR_ENABLED", False)) + return bool(settings.OCR_ENABLED) def _effective_engine(engine: str) -> str: """Resolve ``auto`` to the server's ``DOC_PARSER_ENGINE``; other values pass through.""" if engine != "auto": return engine - configured = getattr(settings, "DOC_PARSER_ENGINE", None) or "anydoc" + configured = settings.DOC_PARSER_ENGINE or "anydoc" return str(configured).strip().lower() diff --git a/docsgpt/parser/embedding_pipeline.py b/docsgpt/parser/embedding_pipeline.py index a0dbd02b..893c77f6 100755 --- a/docsgpt/parser/embedding_pipeline.py +++ b/docsgpt/parser/embedding_pipeline.py @@ -53,7 +53,7 @@ def _resolve_batch_size() -> int: Returns: Chunks per embed request, always >= 1. """ - raw = getattr(settings, "EMBEDDINGS_BATCH_SIZE", None) + raw = settings.EMBEDDINGS_BATCH_SIZE # Explicit type check rather than a bare ``int(raw)``: ``int(MagicMock())`` # succeeds and yields 1, which would silently drop ingest back to the # per-chunk behaviour this batching replaces. @@ -326,7 +326,7 @@ def embed_and_store_documents( store = VectorCreator.create_vectorstore( settings.VECTOR_STORE, source_id=source_id, - embeddings_key=os.getenv("EMBEDDINGS_KEY"), + embeddings_key=settings.EMBEDDINGS_KEY, ) loop_start = resume_index else: @@ -336,7 +336,7 @@ def embed_and_store_documents( settings.VECTOR_STORE, docs_init=[docs[0]], source_id=source_id, - embeddings_key=os.getenv("EMBEDDINGS_KEY"), + embeddings_key=settings.EMBEDDINGS_KEY, ) # Record the seeded chunk so single-doc ingests don't fail # ``assert_index_complete`` — the loop never runs for @@ -351,7 +351,7 @@ def embed_and_store_documents( store = VectorCreator.create_vectorstore( settings.VECTOR_STORE, source_id=source_id, - embeddings_key=os.getenv("EMBEDDINGS_KEY"), + embeddings_key=settings.EMBEDDINGS_KEY, ) # Only wipe the index on a fresh run — a resume must keep the # chunks that earlier attempts already embedded. diff --git a/docsgpt/parser/file/bulk.py b/docsgpt/parser/file/bulk.py index 4828f893..10ccfb1f 100644 --- a/docsgpt/parser/file/bulk.py +++ b/docsgpt/parser/file/bulk.py @@ -344,7 +344,7 @@ def get_default_file_extractor( """ if ocr_enabled is None: ocr_enabled = settings.OCR_ENABLED - selected = (engine or getattr(settings, "DOC_PARSER_ENGINE", None) or "anydoc") + selected = (engine or settings.DOC_PARSER_ENGINE or "anydoc") selected = str(selected).strip().lower() if selected == "docling": return _docling_file_extractor(ocr_enabled, pdf_text_fast_path) diff --git a/docsgpt/parser/file/docling_parser.py b/docsgpt/parser/file/docling_parser.py index e1efd546..577f43d3 100644 --- a/docsgpt/parser/file/docling_parser.py +++ b/docsgpt/parser/file/docling_parser.py @@ -335,11 +335,7 @@ def _ocr_min_chars_per_page() -> int: try: return int( - getattr( - settings, - "OCR_MIN_CHARS_PER_PAGE", - _DEFAULT_OCR_MIN_CHARS_PER_PAGE, - ) + settings.OCR_MIN_CHARS_PER_PAGE ) except (TypeError, ValueError): return _DEFAULT_OCR_MIN_CHARS_PER_PAGE diff --git a/docsgpt/parser/file/ocr_parser.py b/docsgpt/parser/file/ocr_parser.py index 28b67a7c..3a72f049 100644 --- a/docsgpt/parser/file/ocr_parser.py +++ b/docsgpt/parser/file/ocr_parser.py @@ -124,7 +124,7 @@ def resolve_ocr_backend(requested: Optional[str] = None) -> str: """ from docsgpt.core.settings import settings - backend = str(requested or getattr(settings, "OCR_BACKEND", None) or "auto").strip().lower() + backend = str(requested or settings.OCR_BACKEND or "auto").strip().lower() if backend not in VALID_OCR_BACKENDS: logger.warning(f"Unknown OCR_BACKEND {backend!r}; using auto") backend = "auto" @@ -154,7 +154,7 @@ def resolve_native_ocr_engine(requested: Optional[str] = None) -> str: """ from docsgpt.core.settings import settings - engine = str(requested or getattr(settings, "OCR_ENGINE", None) or "tesseract").strip().lower() + engine = str(requested or settings.OCR_ENGINE or "tesseract").strip().lower() if engine in NATIVE_OCR_ENGINES: return engine if engine in VALID_OCR_ENGINES: @@ -172,7 +172,7 @@ def ocr_min_chars_per_page() -> int: from docsgpt.core.settings import settings try: - return int(getattr(settings, "OCR_MIN_CHARS_PER_PAGE", _DEFAULT_MIN_CHARS_PER_PAGE)) + return int(settings.OCR_MIN_CHARS_PER_PAGE) except (TypeError, ValueError): return _DEFAULT_MIN_CHARS_PER_PAGE @@ -182,7 +182,7 @@ def render_dpi() -> int: from docsgpt.core.settings import settings try: - dpi = int(getattr(settings, "OCR_RENDER_DPI", _DEFAULT_RENDER_DPI)) + dpi = int(settings.OCR_RENDER_DPI) except (TypeError, ValueError): dpi = _DEFAULT_RENDER_DPI return max(_MIN_RENDER_DPI, min(_MAX_RENDER_DPI, dpi)) @@ -299,7 +299,7 @@ class TesseractEngine: else: from docsgpt.core.settings import settings - configured = str(getattr(settings, "OCR_LANGS", "") or "eng") + configured = str(settings.OCR_LANGS or "eng") langs = [lang.strip() for lang in configured.split("+") if lang.strip()] return "+".join(langs) or "eng" @@ -388,7 +388,7 @@ class DeepseekOcrEngine: self.url = url or settings.OCR_DEEPSEEK_URL self.model = model or settings.OCR_DEEPSEEK_MODEL - self.timeout = float(timeout if timeout is not None else getattr(settings, "OCR_DEEPSEEK_TIMEOUT", 300)) + self.timeout = float(timeout if timeout is not None else settings.OCR_DEEPSEEK_TIMEOUT) self.prompt = prompt self.max_tokens = max_tokens diff --git a/docsgpt/parser/remote/github_loader.py b/docsgpt/parser/remote/github_loader.py index dbf75d65..3ba4a55c 100644 --- a/docsgpt/parser/remote/github_loader.py +++ b/docsgpt/parser/remote/github_loader.py @@ -129,7 +129,7 @@ class GitHubLoader(BaseRemote): def _max_file_bytes(self) -> int: """Resolve the per-blob size cap; ``0`` disables it.""" - raw = getattr(settings, "GITHUB_INGEST_MAX_FILE_BYTES", None) + raw = settings.GITHUB_INGEST_MAX_FILE_BYTES if isinstance(raw, bool) or not isinstance(raw, (int, str)): return 1048576 try: @@ -139,7 +139,7 @@ class GitHubLoader(BaseRemote): def _max_workers(self) -> int: """Resolve the parallel-fetch width, clamped to a sane range.""" - raw = getattr(settings, "GITHUB_INGEST_MAX_WORKERS", None) + raw = settings.GITHUB_INGEST_MAX_WORKERS if isinstance(raw, bool) or not isinstance(raw, (int, str)): return 8 try: diff --git a/docsgpt/parser/tokenization.py b/docsgpt/parser/tokenization.py index 6391461d..0f45475f 100644 --- a/docsgpt/parser/tokenization.py +++ b/docsgpt/parser/tokenization.py @@ -267,7 +267,7 @@ def get_token_counter(embeddings_name: Optional[str] = None) -> TokenCounter: A :class:`HuggingFaceCounter` for a model whose tokenizer could be loaded, else a :class:`TiktokenCounter`. """ - name = embeddings_name or getattr(settings, "EMBEDDINGS_NAME", None) + name = embeddings_name or settings.EMBEDDINGS_NAME key = name or "__default__" with _cache_lock: if key in _cache: diff --git a/docsgpt/retriever/dispatcher.py b/docsgpt/retriever/dispatcher.py index d96b0037..a14fc02f 100644 --- a/docsgpt/retriever/dispatcher.py +++ b/docsgpt/retriever/dispatcher.py @@ -344,6 +344,6 @@ def build_dispatcher(create_classic: Callable[[], BaseRetriever], **kwargs): Returns: A ``Dispatcher`` or the legacy retriever from ``create_classic``. """ - if not getattr(settings, "PER_SOURCE_RETRIEVAL_ENABLED", True): + if not settings.PER_SOURCE_RETRIEVAL_ENABLED: return create_classic() return Dispatcher(**kwargs) diff --git a/docsgpt/sandbox/artifacts_capture.py b/docsgpt/sandbox/artifacts_capture.py index 9315f87d..697ca718 100644 --- a/docsgpt/sandbox/artifacts_capture.py +++ b/docsgpt/sandbox/artifacts_capture.py @@ -277,7 +277,7 @@ def _cleanup_orphan(storage: Any, saved_key: Optional[str]) -> None: def _check_single_artifact_size(size: int) -> None: """Reject a single artifact version whose byte size exceeds ``ARTIFACT_MAX_BYTES``.""" - max_bytes = int(getattr(settings, "ARTIFACT_MAX_BYTES", 0) or 0) + max_bytes = int(settings.ARTIFACT_MAX_BYTES or 0) if max_bytes > 0 and size > max_bytes: raise QuotaExceeded(f"artifact is too large: {size} bytes exceeds the {max_bytes}-byte per-file cap") @@ -292,10 +292,10 @@ def _enforce_user_quota(repo: ArtifactsRepository, user_id: str, added_bytes: in existing identity. """ _check_single_artifact_size(added_bytes) - max_count = int(getattr(settings, "ARTIFACT_MAX_COUNT_PER_USER", 0) or 0) + max_count = int(settings.ARTIFACT_MAX_COUNT_PER_USER or 0) if new_artifact and max_count > 0 and repo.count_for_user(user_id) >= max_count: raise QuotaExceeded(f"artifact count quota reached ({max_count}); delete artifacts to free space") - max_total = int(getattr(settings, "ARTIFACT_MAX_TOTAL_BYTES_PER_USER", 0) or 0) + max_total = int(settings.ARTIFACT_MAX_TOTAL_BYTES_PER_USER or 0) if max_total > 0 and repo.total_bytes_for_user(user_id) + added_bytes > max_total: raise QuotaExceeded(f"artifact storage quota reached ({max_total} bytes); delete artifacts to free space") diff --git a/docsgpt/sandbox/sandbox_creator.py b/docsgpt/sandbox/sandbox_creator.py index 4cf46b7e..0b34628d 100644 --- a/docsgpt/sandbox/sandbox_creator.py +++ b/docsgpt/sandbox/sandbox_creator.py @@ -64,7 +64,7 @@ class SandboxCreator: def get_manager(cls) -> SandboxManager: """Return the process-wide ``SandboxManager``, building it on first use.""" if cls._instance is None: - backend = cls.create_backend(getattr(settings, "SANDBOX_BACKEND", "jupyter")) + backend = cls.create_backend(settings.SANDBOX_BACKEND) cls._instance = SandboxManager( backend=backend, max_ttl=float(settings.SANDBOX_MAX_TTL), diff --git a/docsgpt/scripts/reembed.py b/docsgpt/scripts/reembed.py index 63c6e56e..eae8c71c 100644 --- a/docsgpt/scripts/reembed.py +++ b/docsgpt/scripts/reembed.py @@ -309,7 +309,7 @@ def reembed_pgvector(source_id: str, batch_size: int, dry_run: bool) -> Tuple[in # The graph seeds every traversal from its own vectors, so leaving them # in the old model's space is the same silent mismatch this script # exists to remove -- and at equal widths nothing would report it. - if getattr(settings, "GRAPHRAG_ENABLED", False): + if settings.GRAPHRAG_ENABLED: nodes = reembed_graph_nodes(store, conn, source_id, batch_size, dry_run) if nodes: logger.info( @@ -512,7 +512,7 @@ def main(argv: Optional[Sequence[str]] = None) -> int: # latency and a dependency on a worker running. Loading the model here also # means the script reports a real failure for a model it cannot load, # instead of timing out against an empty queue. - if getattr(settings, "EMBEDDINGS_DELEGATE_TO_WORKER", False): + if settings.EMBEDDINGS_DELEGATE_TO_WORKER: logger.info("Embedding in-process; worker delegation does not apply here.") settings.EMBEDDINGS_DELEGATE_TO_WORKER = False diff --git a/docsgpt/storage/db/bootstrap.py b/docsgpt/storage/db/bootstrap.py index 3ebf677e..fcfac3f6 100644 --- a/docsgpt/storage/db/bootstrap.py +++ b/docsgpt/storage/db/bootstrap.py @@ -96,7 +96,7 @@ def _release_boot_only_embeddings(log: logging.Logger) -> None: if settings.EMBEDDINGS_BASE_URL: return - if getattr(settings, "EMBEDDINGS_DELEGATE_TO_WORKER", False) is not True: + if settings.EMBEDDINGS_DELEGATE_TO_WORKER is not True: return import gc @@ -144,8 +144,8 @@ def ensure_vector_schema(*, logger: Optional[logging.Logger] = None) -> None: ) return - dsn = getattr(settings, "PGVECTOR_CONNECTION_STRING", None) - if not dsn and getattr(settings, "POSTGRES_URI", None): + dsn = settings.PGVECTOR_CONNECTION_STRING + if not dsn and settings.POSTGRES_URI: from docsgpt.core.db_uri import normalize_pgvector_connection_string dsn = normalize_pgvector_connection_string(settings.POSTGRES_URI) @@ -172,7 +172,7 @@ def ensure_vector_schema(*, logger: Optional[logging.Logger] = None) -> None: dim: Optional[int] = dimension_for(settings.EMBEDDINGS_NAME) - graph_enabled = bool(getattr(settings, "GRAPHRAG_ENABLED", False)) + graph_enabled = bool(settings.GRAPHRAG_ENABLED) started = time.monotonic() # A plain connection, never the store's pool: this can run pre-fork under # ``gunicorn --preload``, and an inherited pooled socket is a broken one. diff --git a/docsgpt/storage/s3.py b/docsgpt/storage/s3.py index 790b869d..38e20552 100644 --- a/docsgpt/storage/s3.py +++ b/docsgpt/storage/s3.py @@ -4,6 +4,7 @@ import io import logging import os import posixpath +import warnings from typing import BinaryIO, Callable, List, Optional, Tuple import boto3 @@ -30,9 +31,13 @@ class S3Storage(BaseStorage): secret_key = settings.S3_SECRET_ACCESS_KEY region = settings.S3_REGION - legacy_access = getattr(settings, "SAGEMAKER_ACCESS_KEY", None) - legacy_secret = getattr(settings, "SAGEMAKER_SECRET_KEY", None) - legacy_region = getattr(settings, "SAGEMAKER_REGION", None) + # The SAGEMAKER_* fields are marked deprecated on the model and warn on every read; + # this is the one sanctioned reader, and it raises its own operator-facing warning. + with warnings.catch_warnings(): + warnings.simplefilter("ignore", DeprecationWarning) + legacy_access = settings.SAGEMAKER_ACCESS_KEY + legacy_secret = settings.SAGEMAKER_SECRET_KEY + legacy_region = settings.SAGEMAKER_REGION used_legacy = ( (not access_key and legacy_access) diff --git a/docsgpt/storage/storage_creator.py b/docsgpt/storage/storage_creator.py index 1c57db64..8604dbf1 100644 --- a/docsgpt/storage/storage_creator.py +++ b/docsgpt/storage/storage_creator.py @@ -18,7 +18,7 @@ class StorageCreator: @classmethod def get_storage(cls) -> BaseStorage: if cls._instance is None: - storage_type = getattr(settings, "STORAGE_TYPE", "local") + storage_type = settings.STORAGE_TYPE cls._instance = cls.create_storage(storage_type) return cls._instance diff --git a/docsgpt/utils.py b/docsgpt/utils.py index 233e3f7b..2b7496f5 100644 --- a/docsgpt/utils.py +++ b/docsgpt/utils.py @@ -347,7 +347,7 @@ def generate_agent_image_capability( agent_id: object, image_path: object, user_id: object ) -> str: """Create an HMAC capability for one agent's current internal image.""" - secret = getattr(settings, "JWT_SECRET_KEY", "") + secret = settings.JWT_SECRET_KEY if not isinstance(secret, str) or not secret: return "" try: @@ -389,7 +389,7 @@ def generate_image_url(image_path, agent_id=None, user_id=None): if not capability: return "" canonical_agent_id = str(uuid.UUID(str(agent_id))) - base_url = getattr(settings, "API_URL", "http://localhost:7091").rstrip("/") + base_url = settings.API_URL.rstrip("/") return f"{base_url}/api/images/{canonical_agent_id}/{capability}" diff --git a/docsgpt/vectorstore/base.py b/docsgpt/vectorstore/base.py index cad18833..d0c63c04 100644 --- a/docsgpt/vectorstore/base.py +++ b/docsgpt/vectorstore/base.py @@ -277,7 +277,7 @@ def _delegation_enabled() -> bool: with a ``MagicMock``, whose every attribute is a truthy object, and ``bool()`` on that would silently route them through the broker. """ - return getattr(settings, "EMBEDDINGS_DELEGATE_TO_WORKER", False) is True + return settings.EMBEDDINGS_DELEGATE_TO_WORKER is True def get_embeddings( diff --git a/docsgpt/vectorstore/embeddings_delegated.py b/docsgpt/vectorstore/embeddings_delegated.py index 0677c850..75f730b8 100644 --- a/docsgpt/vectorstore/embeddings_delegated.py +++ b/docsgpt/vectorstore/embeddings_delegated.py @@ -173,8 +173,8 @@ class DelegatedEmbeddings: the full timeout at once -- at the shipped 60s and 96 WSGI threads, an API that serves nothing at all, health checks included. """ - queue = getattr(settings, "EMBEDDINGS_QUEUE", "embeddings") - timeout = getattr(settings, "EMBEDDINGS_DELEGATE_TIMEOUT", 60) + queue = settings.EMBEDDINGS_QUEUE + timeout = settings.EMBEDDINGS_DELEGATE_TIMEOUT remaining = self._cooldown_remaining() if remaining > 0: diff --git a/docsgpt/vectorstore/embeddings_local.py b/docsgpt/vectorstore/embeddings_local.py index 259240d4..5a236e3b 100644 --- a/docsgpt/vectorstore/embeddings_local.py +++ b/docsgpt/vectorstore/embeddings_local.py @@ -188,8 +188,8 @@ def _describe_from_repo(repo: str) -> Optional[EmbeddingModel]: def _apply_overrides(spec: EmbeddingModel) -> EmbeddingModel: """Let ``EMBEDDINGS_POOLING``/``EMBEDDINGS_NORMALIZE`` win over any source.""" - pooling = getattr(settings, "EMBEDDINGS_POOLING", None) - normalize = getattr(settings, "EMBEDDINGS_NORMALIZE", None) + pooling = settings.EMBEDDINGS_POOLING + normalize = settings.EMBEDDINGS_NORMALIZE changes = {} if isinstance(pooling, str) and pooling.strip().lower() in ("cls", "mean"): changes["pooling"] = pooling.strip().lower() @@ -289,10 +289,10 @@ class EmbeddingsWrapper: try: _register(self.spec) init_kwargs = {"model_name": self.spec.repo} - threads = getattr(settings, "EMBEDDINGS_THREADS", None) + threads = settings.EMBEDDINGS_THREADS if isinstance(threads, int) and threads > 0: init_kwargs["threads"] = threads - cache_dir = getattr(settings, "EMBEDDINGS_CACHE_DIR", None) + cache_dir = settings.EMBEDDINGS_CACHE_DIR if cache_dir: init_kwargs["cache_dir"] = cache_dir self.model = TextEmbedding(**init_kwargs) @@ -331,7 +331,7 @@ class EmbeddingsWrapper: if not documents: return [] batch_size: Optional[int] = None - raw = getattr(settings, "EMBEDDINGS_MODEL_BATCH_SIZE", None) + raw = settings.EMBEDDINGS_MODEL_BATCH_SIZE if isinstance(raw, int) and not isinstance(raw, bool) and raw > 0: batch_size = raw diff --git a/docsgpt/vectorstore/pgconn.py b/docsgpt/vectorstore/pgconn.py index 5164629d..7c4aa33f 100644 --- a/docsgpt/vectorstore/pgconn.py +++ b/docsgpt/vectorstore/pgconn.py @@ -59,7 +59,7 @@ def resolve_pool_max_size() -> int: """ from docsgpt.core.settings import settings - value = getattr(settings, "PGVECTOR_POOL_MAX_SIZE", DEFAULT_POOL_MAX_SIZE) + value = settings.PGVECTOR_POOL_MAX_SIZE if isinstance(value, int) and not isinstance(value, bool) and value >= 0: return value return DEFAULT_POOL_MAX_SIZE diff --git a/docsgpt/vectorstore/pgvector.py b/docsgpt/vectorstore/pgvector.py index d586073d..60bb5262 100644 --- a/docsgpt/vectorstore/pgvector.py +++ b/docsgpt/vectorstore/pgvector.py @@ -54,9 +54,9 @@ class PGVectorStore(BaseVectorStore): # Use provided connection string or fall back to settings. # If PGVECTOR_CONNECTION_STRING is not set but POSTGRES_URI is, # reuse the same cluster — normalize from SQLAlchemy dialect to libpq form. - self._connection_string = connection_string or getattr(settings, 'PGVECTOR_CONNECTION_STRING', None) + self._connection_string = connection_string or settings.PGVECTOR_CONNECTION_STRING - if not self._connection_string and getattr(settings, 'POSTGRES_URI', None): + if not self._connection_string and settings.POSTGRES_URI: from docsgpt.core.db_uri import normalize_pgvector_connection_string self._connection_string = normalize_pgvector_connection_string(settings.POSTGRES_URI) diff --git a/tests/core/test_settings.py b/tests/core/test_settings.py new file mode 100644 index 00000000..1074fe63 --- /dev/null +++ b/tests/core/test_settings.py @@ -0,0 +1,248 @@ +"""Contract tests for ``docsgpt.core.settings``. + +``Settings`` is composed from one ``SettingsGroup`` per domain; these tests pin +the properties that composition must keep (flat names, no duplicate fields, +validators from every group applied) and that the generated reference page +tracks the definitions. +""" + +import types +import typing +import warnings +from pathlib import Path + +import pytest +from pydantic import ValidationError + +from docsgpt.core.settings import SETTINGS_GROUPS, Settings, settings +from docsgpt.core.settings.reference import reference_path, render_reference + +SECRET_FIELDS = ( + "API_KEY", + "OPENAI_API_KEY", + "ANTHROPIC_API_KEY", + "GOOGLE_API_KEY", + "GROQ_API_KEY", + "HUGGINGFACE_API_KEY", + "NOVITA_API_KEY", + "OPEN_ROUTER_API_KEY", + "EMBEDDINGS_KEY", + "FALLBACK_LLM_API_KEY", + "QDRANT_API_KEY", + "ELASTIC_PASSWORD", + "ELEVENLABS_API_KEY", + "INTERNAL_KEY", + "SCIM_TOKEN", + "OIDC_ISSUER", + "GITHUB_ACCESS_TOKEN", + "MICROSOFT_AUTHORITY", + "MCP_OAUTH_REDIRECT_URI", + "S3_ACCESS_KEY_ID", + "S3_SECRET_ACCESS_KEY", + "SANDBOX_GATEWAY_AUTH_TOKEN", + "DAYTONA_API_KEY", +) + + +@pytest.mark.unit +class TestComposition: + def test_every_group_field_is_a_flat_settings_attribute(self): + with warnings.catch_warnings(): + warnings.simplefilter("ignore", DeprecationWarning) # reading a deprecated field warns + for _, group in SETTINGS_GROUPS: + for name in group.model_fields: + assert name in Settings.model_fields, name + assert hasattr(settings, name), name + + def test_deprecated_fields_warn_on_read(self): + with pytest.warns(DeprecationWarning, match="S3_REGION"): + _ = Settings(_env_file=None).SAGEMAKER_REGION + + def test_no_field_is_defined_in_two_groups(self): + owners: dict[str, str] = {} + for title, group in SETTINGS_GROUPS: + for name in group.model_fields: + assert name not in owners, f"{name} is defined in both {owners[name]} and {title}" + owners[name] = title + assert len(owners) == len(Settings.model_fields) + + def test_every_field_has_a_description(self): + missing = [name for name, field in Settings.model_fields.items() if not field.description] + assert not missing, f"settings without a description: {missing}" + + def test_defaults_load_without_an_env_file(self, monkeypatch): + for name in Settings.model_fields: + monkeypatch.delenv(name, raising=False) + fresh = Settings(_env_file=None) + assert fresh.LLM_PROVIDER == "docsgpt" + assert fresh.VECTOR_STORE == "faiss" + + +@pytest.mark.unit +class TestValidators: + """Validators live on the group that owns the field; composition must keep all of them. + + Pydantic collects validators by method name across the MRO, so two groups + defining one under the same name would silently keep only one; the checks + below span every group. + """ + + def test_every_optional_string_treats_unset_spellings_as_none(self): + names = [ + name + for name, field in Settings.model_fields.items() + if typing.get_origin(field.annotation) in (typing.Union, types.UnionType) + and type(None) in typing.get_args(field.annotation) + and all(a is str or typing.get_origin(a) is typing.Literal for a in typing.get_args(field.annotation) if a is not type(None)) + ] + assert len(names) > 60 + assert {"EMBEDDINGS_POOLING", "AUTH_TYPE", "OIDC_ISSUER"} <= set(names) + loaded = Settings.model_validate({name: " None " for name in names}) + with warnings.catch_warnings(): + warnings.simplefilter("ignore", DeprecationWarning) + assert [name for name in names if getattr(loaded, name) is not None] == [] + + def test_plain_strings_keep_empty_values(self): + assert Settings.model_validate({"MILVUS_TOKEN": "", "JWT_SECRET_KEY": ""}).MILVUS_TOKEN == "" + + @pytest.mark.parametrize("name", SECRET_FIELDS) + @pytest.mark.parametrize("raw", ["None", "none", "", " "]) + def test_unset_secret_spellings_become_none(self, name, raw): + assert getattr(Settings.model_validate({name: raw}), name) is None + + @pytest.mark.parametrize("name", SECRET_FIELDS) + def test_secret_is_stripped(self, name): + assert getattr(Settings.model_validate({name: " k3y "}), name) == "k3y" + + def test_normalize_api_key_classmethod_is_kept(self): + assert Settings.normalize_api_key("None") is None + assert Settings.normalize_api_key(" x ") == "x" + assert Settings.normalize_api_key(42) == 42 + + def test_postgres_uris_are_normalized(self): + loaded = Settings.model_validate( + {"POSTGRES_URI": "postgres://u:p@h/db", "PGVECTOR_CONNECTION_STRING": "postgresql+psycopg://u:p@h/v"} + ) + assert loaded.POSTGRES_URI.startswith("postgresql+psycopg://") + assert loaded.PGVECTOR_CONNECTION_STRING.startswith("postgresql://") + + def test_legacy_docling_ocr_aliases_are_read(self): + loaded = Settings.model_validate({"DOCLING_OCR_ENABLED": "true", "DOCLING_OCR_MIN_CHARS_PER_PAGE": "7"}) + assert loaded.OCR_ENABLED is True + assert loaded.OCR_MIN_CHARS_PER_PAGE == 7 + + +@pytest.mark.unit +class TestReference: + def test_reference_lists_every_setting_once(self): + page = render_reference() + for name in Settings.model_fields: + assert page.count(f"### `{name}`") == 1, name + + def test_reference_prose_has_no_bare_angle_brackets_or_braces(self): + """MDX parses ``<`` and ``{`` in prose as JSX; only code spans may carry them raw.""" + for lineno, line in enumerate(render_reference().splitlines(), 1): + if line.startswith(("{/*", "---")): + continue + prose = "".join(line.split("`")[::2]) # drop the inside of every code span + prose = prose.replace("\\{", "").replace("\\}", "") # escaped braces are fine + assert "<" not in prose and "{" not in prose and "}" not in prose, f"line {lineno}: {line}" + + def test_checked_in_reference_is_current(self): + path: Path = reference_path() + if not path.exists(): + pytest.skip("docs tree not present (installed package, not a checkout)") + assert path.read_text(encoding="utf-8") == render_reference(), ( + "docs/content/Deploying/Settings-Reference.mdx is stale; " + "run: python -m docsgpt.core.settings.reference --write" + ) + + +@pytest.mark.unit +class TestCrossFieldRules: + OIDC = {"OIDC_ISSUER": "https://idp.example/", "OIDC_CLIENT_ID": "docsgpt", "OIDC_FRONTEND_URL": "http://app"} + + def test_oidc_requires_issuer_client_and_frontend(self): + with pytest.raises(ValidationError, match="AUTH_TYPE=oidc requires settings: OIDC_CLIENT_ID, OIDC_FRONTEND_URL"): + Settings.model_validate({"AUTH_TYPE": "oidc", "OIDC_ISSUER": self.OIDC["OIDC_ISSUER"]}) + + @pytest.mark.parametrize("raw", ["", "None", " "]) + def test_oidc_unset_spellings_do_not_satisfy_the_requirement(self, raw): + with pytest.raises(ValidationError, match="OIDC_CLIENT_ID"): + Settings.model_validate({"AUTH_TYPE": "oidc", **self.OIDC, "OIDC_CLIENT_ID": raw}) + + def test_oidc_with_required_settings_loads(self): + assert Settings.model_validate({"AUTH_TYPE": "OIDC", **self.OIDC}).AUTH_TYPE == "oidc" + + def test_oidc_settings_are_not_required_for_other_modes(self): + assert Settings.model_validate({"AUTH_TYPE": "session_jwt"}).OIDC_ISSUER is None + + @pytest.mark.parametrize("raw", [None, "", "None"]) + def test_scim_enabled_requires_a_token(self, raw): + with pytest.raises(ValidationError, match="SCIM_ENABLED requires settings: SCIM_TOKEN"): + Settings.model_validate({"SCIM_ENABLED": True, "SCIM_TOKEN": raw}) + + def test_scim_disabled_needs_no_token(self): + assert Settings.model_validate({"SCIM_ENABLED": False}).SCIM_TOKEN is None + + +@pytest.mark.unit +class TestClosedChoices: + """Enum-like settings are Literal types: a typo fails at startup instead of falling through.""" + + @pytest.mark.parametrize("name", ["AUTH_TYPE", "EMBEDDINGS_POOLING"]) + @pytest.mark.parametrize("raw", ["None", "none", "", " "]) + def test_optional_choice_unset_spellings(self, name, raw): + assert getattr(Settings.model_validate({name: raw}), name) is None + + @pytest.mark.parametrize( + ("name", "raw", "expected"), + [ + ("AUTH_TYPE", " Session_JWT ", "session_jwt"), + ("VECTOR_STORE", "PGVector", "pgvector"), + ("STORAGE_TYPE", "S3", "s3"), + ("URL_STRATEGY", "Backend", "backend"), + ("OCR_BACKEND", "Native", "native"), + ("OCR_ENGINE", "Tesseract ", "tesseract"), + ("SANDBOX_BACKEND", "Daytona", "daytona"), + ("DOC_PARSER_ENGINE", "Docling", "docling"), + ("TTS_PROVIDER", "ElevenLabs", "elevenlabs"), + ("STT_PROVIDER", "", "none"), + ("TTS_PROVIDER", "NONE", "none"), + ("EMBEDDINGS_POOLING", "CLS", "cls"), + ], + ) + def test_choices_are_case_insensitive(self, name, raw, expected): + assert getattr(Settings.model_validate({name: raw}), name) == expected + + @pytest.mark.parametrize( + ("name", "raw"), + [ + ("AUTH_TYPE", "basic"), + ("VECTOR_STORE", "lancedb"), + ("STORAGE_TYPE", "gcs"), + ("OCR_BACKEND", "paddle"), + ("SANDBOX_BACKEND", "docker"), + ("DOC_PARSER_ENGINE", "fast"), + ("STT_PROVIDER", "whisper"), + ("EMBEDDINGS_POOLING", "max"), + ], + ) + def test_unknown_choice_is_rejected(self, name, raw): + with pytest.raises(ValidationError): + Settings.model_validate({name: raw}) + + @pytest.mark.parametrize( + ("name", "raw"), + [ + ("EMBEDDINGS_BATCH_SIZE", 0), + ("COMPRESSION_THRESHOLD_PERCENTAGE", 1.5), + ("UPLOAD_MAX_FILE_BYTES", 0), + ("MESSAGE_EVENTS_RETENTION_DAYS", 0), + ("REMOTE_DEVICE_CMD_QUEUE_TTL_SECONDS", 605), + ("GRAPHRAG_MAX_CHUNKS_FOR_EXTRACTION", -1), + ], + ) + def test_out_of_range_numbers_are_rejected(self, name, raw): + with pytest.raises(ValidationError): + Settings.model_validate({name: raw}) diff --git a/tests/llm/test_openai_responses.py b/tests/llm/test_openai_responses.py index 620d3172..c02e08f2 100644 --- a/tests/llm/test_openai_responses.py +++ b/tests/llm/test_openai_responses.py @@ -29,6 +29,9 @@ def _make_llm(monkeypatch, capabilities=None, store_responses=False): AZURE_DEPLOYMENT_NAME="dep", OPENAI_RESPONSES_STORE=store_responses, OPENAI_REASONING_SUMMARY="auto", + OPENAI_RESPONSES_TRUNCATION_AUTO=False, + OPENAI_PROMPT_CACHE_KEY=False, + OPENAI_PROMPT_CACHE_RETENTION=None, ), ) from docsgpt.llm.openai import OpenAILLM diff --git a/tests/llm/test_responses_chain_budget.py b/tests/llm/test_responses_chain_budget.py index 12025f61..d41ef822 100644 --- a/tests/llm/test_responses_chain_budget.py +++ b/tests/llm/test_responses_chain_budget.py @@ -24,18 +24,19 @@ def _make_llm(monkeypatch, store_responses=True, **extra_settings): "docsgpt.llm.openai.StorageCreator", types.SimpleNamespace(get_storage=lambda: None), ) - monkeypatch.setattr( - "docsgpt.llm.openai.settings", - types.SimpleNamespace( - OPENAI_API_KEY="k", - API_KEY="k", - OPENAI_BASE_URL="", - AZURE_DEPLOYMENT_NAME="dep", - OPENAI_RESPONSES_STORE=store_responses, - OPENAI_REASONING_SUMMARY="auto", - **extra_settings, - ), - ) + # Every setting the Responses path reads, with the hints off; tests opt in per case. + stub = { + "OPENAI_API_KEY": "k", + "API_KEY": "k", + "OPENAI_BASE_URL": "", + "AZURE_DEPLOYMENT_NAME": "dep", + "OPENAI_RESPONSES_STORE": store_responses, + "OPENAI_REASONING_SUMMARY": "auto", + "OPENAI_RESPONSES_TRUNCATION_AUTO": False, + "OPENAI_PROMPT_CACHE_KEY": False, + "OPENAI_PROMPT_CACHE_RETENTION": None, + } + monkeypatch.setattr("docsgpt.llm.openai.settings", types.SimpleNamespace(**{**stub, **extra_settings})) from docsgpt.llm.openai import OpenAILLM llm = OpenAILLM(api_key="k") @@ -142,7 +143,7 @@ def _params(llm, **kwargs): @pytest.mark.unit -def test_build_responses_params_defaults_omit_truncation_and_cache_hints(monkeypatch): +def test_build_responses_params_omits_truncation_and_cache_hints_when_off(monkeypatch): llm = _make_llm(monkeypatch) llm._prompt_cache_key = "conv-123" params = _params(llm) diff --git a/tests/parser/connectors/test_share_point_auth.py b/tests/parser/connectors/test_share_point_auth.py index feab4c6b..ecf25a7d 100644 --- a/tests/parser/connectors/test_share_point_auth.py +++ b/tests/parser/connectors/test_share_point_auth.py @@ -14,8 +14,8 @@ def mock_settings(): s.MICROSOFT_TENANT_ID = "tenant-id-123" s.CONNECTOR_REDIRECT_BASE_URI = "https://redirect.example.com/callback" s.MONGO_DB_NAME = "test_db" - # Delete MICROSOFT_AUTHORITY so getattr falls back to default - del s.MICROSOFT_AUTHORITY + # Unset, as in a real Settings object, so the tenant-derived authority is used. + s.MICROSOFT_AUTHORITY = None return s diff --git a/tests/test_remaining_coverage.py b/tests/test_remaining_coverage.py index b251a6b1..89fbded2 100644 --- a/tests/test_remaining_coverage.py +++ b/tests/test_remaining_coverage.py @@ -421,7 +421,7 @@ class TestBaseLLMAbstractRawGen: # --------------------------------------------------------------------------- -# docsgpt/core/settings.py (line 184 - clean_none_string) +# docsgpt/core/settings (normalize_api_key) # --------------------------------------------------------------------------- @pytest.mark.unit class TestSettingsNormalizeApiKey: