The previous commit claimed both handlers, but only the first one
changed: the search text expected `if (...) return;` on a single line
while prettier had already wrapped it across two, so the second
replacement silently matched nothing. A ctrl/cmd/shift-click on the
collapsed-list entry therefore still closed the mobile nav and cleared
the selected agent in the tab it left behind.
- Opening a folder dropped the active filter. agentsListPath() always
built off the bare /agents/manage root, so from /agents/manage/mine it
navigated to /agents/manage?folder=..., and filterFromPath() then read
back `all`. A regression from making the filters routable: they used
to be component state, which a replace-navigation left alone. The
helper now takes the filter and builds off agentsFilterPath(), which
also fixes the other half — switching filter while inside a folder
re-navigates to the filtered path with the folder kept, instead of the
sync effect pulling it back to the unfiltered URL.
- Both manage-agents handlers cleared the selected agent and closed the
mobile nav before returning early for a modifier-click, so opening the
list in a new tab mutated the tab you stayed on. The modifier check
and preventDefault() now run first.
- The detail breadcrumbs had the name precedence inverted. Tools.tsx
shows `customName || displayName` throughout, so a renamed tool kept
its old name in the breadcrumb alone. ToolConfig now matches, and
RemoteDeviceConfig leads with device?.name as its own heading does.
Boxes is the densest glyph in the settings nav by some way — nine
elements and 43 path commands against a median of about twelve. Three
cubes with their internal facets inside 20px puts the strokes almost on
top of each other, so at the nav's 1.75 weight it renders as a darker
blob than the icons above and below it, even though the stroke width is
identical. Blocks sits at the median, matching Sources, Tools and Teams,
and stays distinct from the database glyph two rows up.
AgentSection computed its folder breadcrumb in a useMemo placed below two
early returns for the empty states. A hook after an early return is
skipped on the render that takes it, which React rejects outright with
"rendered fewer hooks than expected" — the page showed only the error
boundary. Reachable now that each filter is its own route: landing
straight on an empty one renders once while the data loads and again
once it arrives empty, taking the early return the second time. The
memo moves above them.
Two problems in the chat's entrance animations, found while looking for
the reported flicker:
- .fade-in-bubble carried `opacity: 0` on the element and reached full
opacity only by running `fadeInUp` to completion with `forwards`. The
answer text was therefore visible *because* an animation had finished,
so anything that stopped one running left it blank — including the
obvious reduced-motion reset, which is presumably why the rule covers
only .shimmer-text today. The start state moves into the keyframes,
the element rests visible, and both entrances now honour
prefers-reduced-motion.
- The timings were long for their jobs. An expanded tool call reaches
its full height at once, so fading its content over half a second read
as the content lagging the layout rather than as a reveal; it is now
0.16s. The answer entrance goes to 0.26s with a 6px rise instead of
0.5s and 10px.
The sidebar cross-faded between two states, which stopped describing
what was happening once sections could nest: entering a section slid,
but opening an agent from the agent list swapped in place with no
motion at all, so going deeper and going sideways looked identical.
Panels are now positioned from a single number — their depth relative
to the level on screen. A panel above the current level waits off to
the right, the current one sits at rest, and ones below park just off
to the left, so push and pop fall out of the same rule and no direction
has to be tracked. The panel behind travels a quarter of the width and
dims rather than sliding out with the one in front, and the arriving
panel carries a shadow off its leading edge that the container clips
once it lands, so the two read as stacked rather than adjacent.
The motion was also starting far too late. Mounting a section's page
costs a single ~170ms blocking frame in a production build, and the
sidebar's own class change rode along in that same commit: measured
from the click, the panels did not begin moving for ~290ms, so the
animation played to an audience that had stopped expecting it. The two
updates are now split by priority. The level lands as an urgent update
touching nothing but the sidebar, so React can commit and paint it
straight away; the route change goes through startTransition, which
renders the page at low priority and yields instead of blocking that
paint. The target's section is resolved from the path up front, so the
incoming panel arrives with its content already in place. The style
change now lands ~53ms after the click. Only translate and opacity are
animated, so the compositor keeps the motion smooth across the frames
the page render still costs.
Timing is tuned against where the travel actually lands rather than by
feel: half the distance by ~65ms so the panel tracks the click, 90% by
~180ms so the movement reads as movement, settled by ~300ms.
Routes under /agents covered two different things: using an agent (a
conversation) and managing them (the list, editor, logs, schedules).
Sharing the prefix left no way to tell them apart from the pathname, so
the sidebar could not react to one without also reacting to the other.
Management now lives under /agents/manage, and every caller builds its
links from agents/paths.ts rather than from a literal. Pre-split URLs
redirect, keeping their query.
With the prefixes distinct, agent management becomes a section like
settings and admin:
- The list's five filters are routes rather than component state, so a
filtered view is linkable and survives a reload. The hand-rolled pill
row is gone from the content on desktop.
- An agent's own pages (overview, logs, schedules) get a nav titled
after the agent, replacing the breadcrumb-and-underline sub-nav.
Sections nest to support it: `parentPath` makes back mean "up one
level", so leaving an agent lands on the agent list rather than the
chat.
- Sections named after a record are built per route, since the title
comes from the store rather than the path.
- `pageTitle` distinguishes sections whose destinations are separate
pages from ones whose destinations are views of a single page; the
latter keep their own heading instead of flipping to "All".
Below lg, where the sidebar is an overlay, those view-style
destinations appear as a pill row on the page — bouncing out to an
index page to change a filter would be worse than a row of pills.
The workflow builder keeps its full-screen canvas and its own header,
the one place a content-owned nav still earns its keep.
Page shells converge on the shared padding, max width and header, and
the agent list's folder trail uses the shared breadcrumb primitives.
Settings and admin both drove their pages from a horizontal tab strip.
Seven tabs no longer fit: settings had grown scroll arrows, gradient
masks and a hiddenGradient state machine just to survive on mobile, and
the active tab was resolved by comparing translated labels against the
URL. Teams and admin had no home in the strip at all — admin was
reachable only from the Help popover.
Replace both strips with a vertical nav that takes over the sidebar
while you are inside a section, plus a back button that returns to the
app. A declarative registry in navigation/sections.ts holds each
section's destinations; the active item is resolved from the route by
longest path match, so a detail route like /settings/tools/slack keeps
Tools active. Section state is derived from the route rather than
stored, so deep links and browser back keep working.
- Below lg the sidebar is an overlay, so /settings and /admin render
their destination list as page content and each page carries a back
link to it.
- Entering settings no longer clears the conversation; the back button
returns to the route you came from and the chat list stays mounted,
keeping its scroll position.
- Collapsing the sidebar inside a section shows the same destinations
as icons instead of stranding you on one page.
- Detail views nested in a section page (a tool's config, a team) now
use a breadcrumb rather than a second back arrow, so only the section
nav means "leave".
- Teams and admin join the settings nav; admin-only entries are hidden
from non-admins along with the group heading they leave empty.
- check_usage treats a request through a keyless (draft) agent as agent
traffic, matching how its usage rows are bucketed and the headless rule
- dashboard edits carry the stored enabled flag instead of re-enabling the
policy; disabled policies are labelled in the Quotas tab and the editor
- quota 429s send x-should-retry: false so OpenAI SDK clients do not retry
a refusal that cannot succeed before the reset
- cached-input and cache-write rates for Anthropic, OpenRouter and Groq
gpt-oss-120b; refresh OpenRouter deepseek-v3.2 list prices
- UsageQuota reuses usagePercent; docs note that a user override needs an
existing user
- A tool continuation refused for usage now releases the resume claim it
took; before, retries got a 409 until the stale claim was reverted.
- Agent traffic is any row with an agent key or an agent id, so keyless
agents and workflow nodes count toward the agent bucket, not direct.
- The user quota modal discards responses for a previously opened user.
- The usage meter shows every limited bucket, not only 'all'.
- Restore the class separator on the analytics stat card that a formatter
run removed, and align the OpenRouter DeepSeek description with its rates.
The admin dashboard gets a Quotas tab for the instance default, team
allowances and user overrides, with a notice listing models that cost
limits cannot see. A Quota action on the Users tab shows a user's effective
limits, the layer each comes from and their usage, next to the editor for
their override. Each budget is either not set at that layer, a limit, or
unlimited.
Users with a quota see a usage meter with the reset time on the Analytics
settings page, and a refused chat request shows the used amount, the limit
and the reset time in the user's language.
Each token gets a Regenerate action: a confirmation with the new expiration
(preselecting the lifetime the token was issued with), then the one-time
secret view.
The "restrict to specific resources" pickers could not be scrolled and did
not close on an outside click. Inside a Modal a non-modal popover is
portalled outside the dialog, so the dialog's scroll lock swallowed the wheel,
and Radix defers its outside-click dismissal to the document click, which
Modal stops from propagating. MultiSelect takes a `modal` prop for that case.
A resource restriction was checked on ids in the request, not on what the
addressed row pulls in or belongs to. Closed:
- Workflow writes for tokens restricted on sources, tools or prompts (a graph
names those inside its nodes), and attaching a workflow to an agent unless
the token is restricted on workflows too.
- Chat for tools-restricted tokens (chat executes tools; rejected at token
creation as well), and agent-less chat for tokens restricted on prompts or
workflows.
- conversation_id on chat: it must belong to the agent being run, or to no
agent for agent-less chat. Otherwise the server continued, appended to, or
resumed pending tool calls of another agent's conversation.
- Schedules for tokens restricted on anything but agents; schedule-id routes
for every restricted token.
- Conversations and analytics for every restricted token, not only
agent-restricted ones.
Also: create, first publish and adopt return the agent API key masked to a
token without agents:keys; token ids must be canonical UUIDs (urn:uuid: gave
a 500); an expired token is reported as expired; token creation takes a
per-user advisory lock so the cap cannot be raced; admin revoke-sessions
writes a pat_revoked event per token. The UI drops a row whose revoke returns
404 and does not offer a tools restriction next to chat:run.
i18next HTML-escaped interpolated values, so expiry dates showed as
26/09/2026 and token names containing & or quotes were mangled in
the created and revoke dialogs; React already escapes on render, so those
strings opt out like utils/streamingStatusUtils does. The scopes fieldset
drops the browser's default padding so it lines up with the other fields, and
the native checkboxes follow the dark colour scheme.
Settings > Access Tokens lists a user's personal access tokens and lets them
create and revoke tokens: scopes grouped by family, optional restriction to
specific resources, and an expiry bounded by the server's policy. The secret
is shown once after creation and never stored client side. Removes the
unused legacy api-key endpoints, wrappers, type and locale block.
Exposes the three per-source graph options in the source's retrieval settings,
shown only when the retriever is graphrag: where the walk starts (entities or
relationships), whether passages join the walk, and whether vector hits are
blended in. Defaults match the backend's measured-best configuration and are
filled in for sources saved before the options existed. A note points at the
"search tool" exposure, which is what offers the graph to an agent.
- Ship tiktoken's cl100k_base inside the package and build the encoding
from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
temp dir, and read tokenizer.json and repo metadata from that cache, so
a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
the endpoints return 404, audio files fail to ingest with a clear
message, /api/config reports tts_available/stt_available, and the UI
hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
The connector OAuth popup posted the session token to window.opener with
a '*' target origin, so any page that opened the popup received it. An
attacker with an account on a multi-user deployment could start a flow for
their own pending session, get a victim to finish the provider consent, and
receive a token backed by the victim's Drive/SharePoint/Confluence tokens.
- Post popup results only to allowed frontend origins: the callback origin,
OIDC_FRONTEND_URL, the new CONNECTOR_ALLOWED_ORIGINS, and localhost:5173
when the callback runs on a loopback host.
- Render the success page from the callback itself so the token never
appears in a URL; callback-status ignores session_token/user_email params.
- ConnectorAuth accepts messages only from the popup it opened, on the
callback origin reported by /api/connectors/auth.
- /api/connectors/disconnect requires auth and only deletes the caller's
session.
- /api/connectors/sync and /api/remote reject session tokens the caller
does not own.
Fixes#2766
The button is disabled once isCopied is true, but copy() is now async, so
isCopied and the disabled prop only catch up after it resolves. Clicks
landing inside that window both passed the guard and started a write.
Nothing leaked, since the existing clearTimeout already handled the
duplicate timer, but the window did not exist before copy-to-clipboard 4
made the call async. A ref cleared in finally closes it.
Raised by CodeRabbit on #2762. 598 tests pass, build and lint unchanged.
eslint.config.js imports @eslint/js and globals but neither was declared;
both resolved transitively through eslint. Declaring them fixes a latent
break and is a prerequisite for ESLint 10, which stops providing them.
ESLint 10 itself is not taken. eslint-plugin-react 7.37.5 is the latest
release and peers at eslint "^3 || ... || ^9.7", and under 10 every rule
load throws "contextOrFilename.getFilename is not a function" because
ESLint 10 removed context.getFilename(). The frontend stays on eslint 9
until that ships; note npm already marks 9.x as no longer supported.
eslint-plugin-n goes to 18, though the flat config does not reference it
(nor eslint-plugin-import, eslint-plugin-promise or eslint-config-prettier
-- all four are unused devDependencies).
Lint output unchanged from baseline.
mermaid 12 needs no source change; mermaidSecurity.ts and its strict-mode
init-directive test pass unchanged.
Adds a lodash-es ^4.18.1 override. mermaid 12 depends on chevrotain ~11.1.2
(mermaid 11 resolved chevrotain 12, which has no lodash-es), and chevrotain
11 pulls lodash-es <= 4.17.23, carrying GHSA-r5fr-rjxr-66jc and
GHSA-f23m-r3pf-42rh, both high. 4.18.1 is the fixed line. Without the
override the frontend gains 5 high-severity advisories.
Verified rendering in Chromium rather than the unit suite, since happy-dom
has no SVG getBBox and cannot render any mermaid version. Flowchart,
sequence, pie, class, state, gantt, ER and mindmap all produce SVG on 12,
matching 11 to the byte except classDiagram (+17 bytes). A strict-mode
`click ... "javascript:..."` target is still stripped.
598 tests pass, build clean, 0 advisories.
react-markdown 10 removed the `className` prop and no longer renders a
wrapper element; passing className now throws at render time. Each of the
three call sites gets an explicit wrapping div carrying the same classes,
which reproduces the v9 DOM exactly.
The large diff is re-indentation from the added wrapper. `git diff -w`
shows the only substantive change per file is the moved className.
Fixes the 19 tests that failed on the upgrade. 598 tests pass, build clean,
lint unchanged.
copy-to-clipboard 4 returns Promise<boolean> instead of boolean, so
CopyButton's handleCopy awaits it. Without the await the success branch
always ran, since a Promise is always truthy. Awaiting also routes a
rejected copy into the existing catch.
vite-plugin-svgr 5 and @shadcn/react 0.3.1 needed no source changes.
598 tests pass, build clean, lint unchanged.
The chat selection was inferred by diffing the refreshed source list
against the snapshot taken before the upload, so any source that appeared
meanwhile -- another upload finishing, a new team share -- could be picked
instead. The upload response already carries the id; match on it, and
select nothing when it is absent rather than guessing.
Drops the now-unused sourceDocs read and its stale dependency.
selectedSourceIdsFromAgent preferred the extra sources over the legacy
single source instead of merging them. Since every save writes both
fields, opening an agent that held both and saving -- even just a rename
-- silently detached the primary source. Retrieval reads both, so merge
them.
The edit form also reported unsaved changes the moment it loaded: the
snapshot taken after the fetch defaulted the retriever to "classic" and
left the default model empty, while the effects that run straight after
rewrite both. Normalise the snapshot through the same serializer and
rules the effects apply.
A source attached to a team-shared agent that the caller cannot list --
the owner's private source -- rendered no picker row, so it showed as
attached with nothing to switch it off. Append rows for selected ids
missing from the list, and memoise the label lookup the trigger and the
picker now share.
Uploading from the agent builder repointed the open conversation onto the
new source and persisted it, because the upload modal always dispatched
the chat selection. Gate that behind a prop the builder opts out of. The
wiki source is created synchronously and never refreshed the source list,
so the builder selected an id the picker could not render and the trigger
read "External KB"; refresh on every completion path that skips ingest
tracking.
Removing the synthetic "Default" source made an empty source list
reachable, which the stored-selection reconciler skipped, leaving a
deleted source checked forever with no way to clear it. Treat an empty
list as a real answer and only skip when the list has not loaded.
Finally, finish the picker extraction: the chat input still hand-rolled
the item mapping, so team users saw the same sources grouped in the agent
builder and flat in chat. Point it at the shared helper, with its own
locale keys, and make the sources link a router navigation instead of a
full page reload.
Review follow-ups:
- The agents list and the pinned list dropped agents that had neither a
source nor a retriever. Publishing such an agent is now allowed, so it
vanished from the list right after publish. Both filters are gone; a
source-less agent skips retrieval and is still runnable. Tests cover
both listings.
- The agent form selected an uploaded source by diffing the source
catalog before and after the upload, which would also pick up any
source created or shared in the meantime. Upload now passes the id of
the source it created to onSuccessfulUpload, and the form selects only
that id.
The sources list used to start with a fake "Default" entry that had no id
and, at run time, meant "no source, skip retrieval". The agent form
pre-selected it, snapped back to it when the last source was deselected,
and refused to publish without it, so a new agent always looked like it
had a knowledge base when it had none.
Backend
- /api/sources returns only ingested sources; no placeholder row.
- Publishing an agent no longer requires a source on create or update.
The legacy "default" value is still accepted and maps to NULL.
Frontend
- The agent source picker starts empty, can be cleared, and shows a hint
that a source-less agent answers from the model and its tools only.
- The picker groups sources into "Your sources" and "Shared with team"
when any team-shared source exists, shows "N sources selected" for a
multi-selection, and gets the same "Go to Sources" / "Upload new"
footer as the chat picker. A source uploaded from the form is selected
when it lands.
- Source selection serialisation and the picker id live in one helper
shared with the chat picker; the four copies in the form are gone.
- The client no longer seeds a placeholder source in the store, and the
dead auto-select of a "default" document is removed.
- The frontend image ran the Vite dev server in development mode, so
.env.development supplied its defaults (notification banner, Google client
id, local API host). The static build only loads .env.production, so the
build stage now copies .env.development in as the baseline and the compose
files pass every VITE_* the app reads through from .env; the runtime script
skips empty values so a blank passthrough keeps the build-time default.
.dockerignore kept only the .local variants out.
- VITE_DISABLE_SOURCE_FE disables sources only when it is the string true.
- DoclingParser: find_spec raises when docling itself is absent; the install
hint now covers that path, with a regression test.
- verify_offline: direct tests for verify(); the PR image check builds and
verifies the -docling variant as well as slim.
- Workflows this branch adds or rewrites pin actions by commit, pass the
release tag through env instead of template expansion, and do not persist
checkout credentials.
- OCR guide no longer claims pre-built images never include docling.
Backend (arc53/docsgpt): 4.5 GB compressed -> 0.9 GB with both embedding
models and tiktoken baked in.
- torch/transformers gone from the default install (docling extra only).
- Ubuntu 24.04 ships python3.12: no deadsnakes PPA, no software-properties-
common; every pin is a wheel, so no gcc/g++/rust in the builder.
- COPY --chown and a prefetch that runs as the process user replace the
trailing chown -R, which duplicated the 600 MB model layer.
- .dockerignore keeps __pycache__, .coverage, local indexes and .env out.
- EXTRAS build arg (INSTALL_DOCLING kept as an alias); the docling variant
also bakes docling's layout/table/RapidOCR models (DOCLING_ARTIFACTS_PATH)
and tesseract, and drops only the discovery documents of Google APIs the
app never builds.
- FLASK_DEBUG env removed (unused); OCI labels added.
Frontend (arc53/docsgpt-fe): 302 MB Vite dev server -> 25 MB static build
behind nginx. VITE_* variables are injected at container start into
/config.js and read through src/env.ts, so the image no longer needs a
rebuild per deployment; docker-compose.yaml keeps hot reload via the dev
target.
Publishing: every release and develop build now pushes a slim tag and a
-docling tag (docling engine + models + tesseract). docker-compose-hub.yaml
takes DOCSGPT_IMAGE_TAG / DOCSGPT_IMAGE_VARIANT; docker-compose-standalone.yaml
runs the stack from pre-built images without a checkout and is attached to
each release. setup.sh selects the -docling variant for OCR instead of
requiring a local build. A new workflow builds the image on PRs that touch
it and runs verify_offline under --network none; lint checks the exported
requirements match uv.lock.
Conflicts, and how each was taken:
- application/core/settings.py — ours. The renamed OCR_ENABLED /
OCR_ATTACHMENTS_ENABLED / OCR_MIN_CHARS_PER_PAGE accept main's
DOCLING_OCR_* spellings as AliasChoices, so nothing is dropped.
- application/Dockerfile — both. Main's install layers plus the
INSTALL_DOCLING build arg.
- application/parser/file/constants.py — both imports.
- deployment/docker-compose.yaml — both. The INSTALL_DOCLING /
INSTALL_TESSERACT build args on backend and worker, and main's
-Q docsgpt,parsing,embeddings, which query embedding needs.
- tests/conftest.py — theirs. Both sides fixed the same pytest-postgresql
9.0.0 autocommit= breakage; main's spelling is the one already on main.
- application/requirements.txt — the comments claimed different reasons
torch is in core. Main's is the true one now: it removed
sentence-transformers, so docling is torch's only remaining consumer.
Two things the merge broke without conflicting:
- onnxruntime. This branch moved it out of core into the docling extra;
main meanwhile made it the runtime local embeddings execute on
(fastembed). Git took the deletion, leaving fastembed with no pinned
runtime in a repo that pins everything. Restored to core, and no longer
pinned twice from the extra.
- The frontend copy of ATTACHMENT_PARSER_EXTENSIONS. The backend list is
derived and picked up the anydoc suffixes; the hand-kept frontend mirror
did not, so the composer would refuse files the API accepts.
tests/parser/file/test_constants.py is what caught it.
ruff, pytest (9897 passed), frontend build and docs build all pass. The
image build is unverified: no Docker daemon on this machine.
Review follow-ups.
A BOM told the sniff which encoding to read, but was also taken as the
verdict: three prepended bytes let any binary through, including as
notes.txt. A BOM now only selects the test — UTF-8 falls through to the
byte rules on the remainder, UTF-16/32 decode and judge the characters
(NUL, unprintable, or replacement chars from bytes the decoder could not
read). Real Notepad-Unicode text still passes, mp4-behind-a-BOM does not,
in either language.
The gate treated the full parser table as a given, but without docling the
fallback extractor has no .tif/.tiff/.bmp/.webp/.vtt/.xml handler, so those
suffixes skipped the content check and reached the plain-text fallthrough —
the original bug, one install away. The worker now passes the keys of the
extractor it actually built, making the second gate stricter than the
route's static one rather than a copy of it.
Cache bins: extraction coerced a missing value to 0 and only non-zero bins
were recorded, so a provider reporting cached_tokens=0 persisted as NULL —
indistinguishable from "not reported", and OpenAI reports exactly that on
every uncached request. Bins are now carried as Optional and recorded when
not None, which is what the nullable columns and the NULL-means-unknown
comment already assumed. Anthropic's cache_read/cache_creation bins had the
same shape and are fixed alongside; the int-or-None coercion is shared in
llm/base.py.
Review follow-ups on the attachment gate.
.txt was listed as parser-backed, but it has no parser — it *is* the
plain-text fallthrough. That let it skip the content check, so renaming a
video to notes.txt walked straight back into the bug the gate exists for
(verified: 5132 chars of binary "extracted" and stored). The list is now
exactly the file extractor's keys, .txt included in the content check like
any other unparsed suffix, and the drift test asserts equality rather than
containment. The sniff now recognises a UTF-16/32 BOM as text, so a
Notepad "Unicode" .txt is not caught by the NUL-byte rule.
The picker's accept filter listed parser-backed suffixes only, hiding .txt,
.py and .log — files the gate reads happily — behind "All files". It now
carries text/* as well, so it can never be narrower than what the upload
accepts.
A rejected batch carries one errors entry per file, but the non-200 branch
applied the top-level message to every chip, so two files failing for two
reasons both reported the first one. Reasons are now matched by
upload_index, with the top-level message as fallback.
_get_store_attachment_user_error no longer reads str(exc): the
unsupported-type message is rebuilt from the filename, so no exception
state can reach a response body (CodeQL py/stack-trace-exposure).
A chat attachment with no parser fell through to SimpleDirectoryReader's
plain-text open(), so a phone-uploaded video was "extracted" into megabytes
of binary garbage, truncated, and stored with extraction.status == "ok".
Gate attachments in two tiers instead. A suffix with a dedicated parser is
admitted on its name — a PDF is binary and parses fine. Anything else has to
read as text: the first 8KB are sampled and refused on a NUL byte or too many
other control bytes. That keeps source, config and log files working through
the plain-text fallthrough, and keeps out videos, archives and renamed
binaries alike. The route checks the staged spool before anything is stored
or queued; the worker repeats the check where the local file exists, raising
the non-retryable AttachmentRejectedError.
SUPPORTED_ATTACHMENT_EXTENSIONS gains the parser-backed suffixes it was
missing (.tiff, .tif, .bmp, .webp, .vtt, .xml) and is now exactly the file
extractor's keys plus .txt, with a test asserting the two agree. The composer
applies the same rule client-side, so an unsupported file is named before it
costs an upload, and a test pins the frontend list to the backend one.
Attachment failures now show their reason inline under the chips rather than
only in a hover tooltip, which a touch user can never see, and only after a
send was attempted. Dropped `accept` from the dropzone: it discarded rejected
drops with no feedback and disagreed with the server about text files.