* fix(build): embed commit SHA for release provenance (#1571 part 2)
Docker builds exclude .git via .dockerignore, so buildvcs cannot read
VCS metadata and published images carry no commit information. A running
image cannot be lined up with the source commit it was built from.
Embed cmd.CommitSHA at link time (maintainer-endorsed option 2) and
surface it in goclaw version output:
- cmd: add CommitSHA var, print commit in version cmd when injected
- Makefile: pass git rev-parse HEAD via LDFLAGS
- Dockerfile: accept COMMIT_SHA build arg (default unknown)
- docker-compose.yml: pass GOCLAW_COMMIT_SHA through as build arg
- release workflows: inject github.sha into release and dev-beta builds
Backward compatible: binaries built without the flag keep the existing
version output format.
* fix(build): pass COMMIT_SHA to docker image builds, surface it in doctor/upgrade
Review follow-up for PR #1590:
- Add COMMIT_SHA=${{ github.sha }} to the build-args of every
docker/build-push-action step (release.yaml, dev-beta-release.yaml,
release-beta.yaml, fork-image.yaml) so published images — the
artifact issue #1571 is about — carry the commit, not just release
tarball binaries.
- Surface the commit in doctor and upgrade output (App version line)
via a shared commitSuffix() helper, so operators can read provenance
from logs without exec'ing goclaw version (review suggestion #2).
- version cmd refactored onto the same helper; output unchanged.
Runs on models without a registered tokenizer (e.g. 9router brand models)
ended with the generic "Agent couldn't generate a response" fallback even
though the real request used about 55% of the context window.
PruneStage counted history with TokenCounter, which falls back to a
chars/2 heuristic for unregistered models and overcounted about 1.8x.
Once over budget it ran memory flush (~35s, invisible in traces), then
mid-loop compaction, which cannot summarize a history made only of tool
call/result pairs. The callback reported the untouched history as
compacted, PruneStage still saw it over budget and returned AbortRun
before any LLM call, and FinalizeStage replaced the empty reply with the
fallback.
- PruneStage and ContextStage overhead count with the request guard's
BudgetCounter. PruneStage no longer controls loop flow; the final
request guard in ThinkStage decides.
- CompactMessages returns ErrNotCompacted when history is unchanged.
Callers stop counting it as a compaction and do not retry it in the
same run, while post-run summarization still sees the pressure.
- When the guard exhausts every reduction step, ThinkStage stops the run
with a localized chat.context_budget_exceeded notice instead of an
error, so the run's tool results are still persisted. The stop reason
marks the trace and agent span as error; team tasks, cron and
heartbeat treat it as a failure via RunOutcome.Failure().
- Memory flush and mid-loop compaction emit event spans.
- Web and desktop UIs treat an unset context_pruning as enabled (the
backend default since 7639a8c0), keep it unset when untouched, and can
re-enable pruning after it was turned off.
The agent loop signals silent replies (NO_REPLY, cancelled runs, suppressed
errors) by publishing an outbound message with empty content plus the inbound
routing metadata (placeholder_key / local_key). Slack, Telegram and Discord
each implement an empty-content branch in Send() that deletes their streamed
'Thinking...' placeholder — the partial draft left visible in the thread when
delivery is suppressed.
deliverOutbound skipped every text-only empty outbound, so that cleanup
signal never reached channel.Send and the stray partial draft stayed
published (issue #1475). Pass the signal through when placeholder routing
metadata is present; keep skipping bare empty messages so channels without
an empty-content branch never render empty bubbles.
Fixes#1475
* fix(gateway): admit operator.provision keys on tenants.create and tenants.users.add
The CVE #866 fail-closed hardening regressed operator.provision: the
role-only router check maps provision-only API keys to viewer and
rejects tenants.create / tenants.users.add with 'requires admin role',
even though the tenant handlers explicitly admit ScopeProvision on
exactly those two methods (issue #1524).
Restore the intended least-privilege provisioning path without any role
promotion: the router now allows credentials carrying ScopeProvision on
exactly the two tenant-provisioning RPCs (permissions.IsProvisionMethod).
Every other admin/write surface stays denied, and viewers without the
provision scope gain nothing.
Regression tests (internal/gateway/router_test.go) prove:
- provision-only succeeds on tenants.create and tenants.users.add
- provision-only stays denied on tenants.update and agents.create
- plain viewers stay denied on the provisioning methods
- unauthenticated clients stay denied
* refactor(gateway): route provisionScopeAllowed through permissions.HasProvisionScope
Address review feedback on #1584: HasProvisionScope was exported and
tested but had no production caller — the router used Client.HasScope
directly. Wire the router to the permissions helper so the provision
scope check has a single source of truth.
The enrich worker recorded a doc in its dedup map even when writing its
body chunks failed, so the next event with the same content hash was
skipped and the doc stayed without chunks until a manual rescan.
indexBody now returns its error and processChunk leaves those docs out of
the dedup map, so the next event for the same hash tries again. Docs whose
chunks were written are still deduplicated as before.
vault_search only indexed title + path + the auto-summary, and the summary
is written from the first 3000 runes of the file. Anything the summary left
out (room names, prices, codes) could never be found. On top of that the
FTS query used plainto_tsquery, which ANDs every word, so a normal question
like "what is the price of the Deluxe Ocean Suite?" matched nothing even when
the key words were indexed.
- Add vault_document_chunks (migration 98): the file body split into
chunks with their own tsvector and embedding. body_indexed_hash on
vault_documents records which content_hash the chunks came from, so
unchanged files are not re-chunked or re-embedded.
- The enrich worker rebuilds chunks when a file changes. This needs no LLM,
so it runs even when no provider is configured.
- Rescan backfills chunks for docs indexed before this change, since they
already have a summary and never go back through the worker.
- FTS matches any query word; ts_rank still ranks docs with more matching
words first. Both FTS and vector search look at the doc and its chunks
and score each doc by its best hit.
SQLite is unchanged: its vault search is LIKE on title/path only.
ACP providers use api_base to store an executable command/path, not a URL.
A previous SSRF-hardening refactor incorrectly classified ACP as a local-URL
provider type, causing create/update to reject valid ACP configs with
"provider URL must use http or https scheme".
Changes:
- Remove ProviderACP from localURLProviderTypes.
- Add validateACPExecutablePath mirroring the runtime registration logic in
cmd/gateway_providers.go (allows empty, claude/codex/gemini, or absolute path;
rejects URLs and relative paths).
- Route ACP through the new validator inside validateProviderURL.
- Add unit tests for validateACPExecutablePath.
- Update existing local-type tests that assumed ACP was URL-based.
- Add create/update HTTP handler regression tests for ACP binary, empty binary,
and URL rejection.
Bug-first verification: without the fix, ACP api_base values like "gemini" or
absolute paths fail URL validation, while URLs are incorrectly accepted; with
the fix the behavior is reversed to match the intended executable-path model.
Fixesnextlevelbuilder/goclaw#1481.
* feat(tools): add Serply as a web_search provider
Serply is a SERP API returning Google web results on a tenant-supplied
key. It joins the existing chain the same way Exa, Tavily and Brave do:
keyed from config_secrets, skipped silently when no key is stored, and
selectable through the provider parameter for cross-engine corroboration.
Freshness maps onto Google's tbs recency filter, which Serply forwards
upstream. Only the pd/pw/pm/py shortcuts are sent: upstream ignores the
equivalent cdr date range and answers with unfiltered results, so a
range leaves tbs unset rather than naming a filter the results do not
honour.
The settings form gained a fifth sortable provider, so the locked
DuckDuckGo card now derives its position number instead of hardcoding
it, and a stored provider_order is reconciled against the full provider
list rather than against Parallel alone.
* test(tools): use a locally named round trip stub in the serply test
The serply cancellation test borrowed parallelRoundTripFunc from
web_search_parallel_test.go. Define serplyRoundTripFunc alongside the
test that uses it so the file does not depend on another provider's
test helper.
When a document contains multiple [[target]] wikilinks pointing to the same
destination, SyncDocLinks produced duplicate (FromDocID, ToDocID, LinkType)
entries. The batch INSERT ... ON CONFLICT DO UPDATE in CreateLinks then failed
because PostgreSQL rejects a single statement updating the same row twice.
Deduplicate by (FromDocID, ToDocID, LinkType) in SyncDocLinks before calling
CreateLinks, concatenating contexts with " | " separator so no surrounding
text is lost.
Fixes#1473
(cherry picked from commit 549c81fd2406875be34cec3d7274125238e7d4a0)
Register Requesty (https://router.requesty.ai/v1) next to OpenRouter on the
existing OpenAI-compatible transport: config and env vars
(GOCLAW_REQUESTY_API_KEY, GOCLAW_REQUESTY_BASE_URL), secret masking, DB and
in-memory registration, CLI setup, doctor, placeholder provider, OpenAPI enum,
Web/Desktop provider lists and docs.
The models list merges Requesty managed policies (GET /models/managed) with
the key's catalog from GET /models.
Map skills/anysearch to the AnySearch integration review checklist
(basic built-in + extensions). Record 2026-09-24 anonymous live smoke
for search, get_sub_domains, batch_search, extract, and vertical search.
Drop the Sh/PS ports in response to PR triage feedback on the
maintenance/security surface. Keep Python (requests) as primary and
Node as the zero third-party-dep fallback. Docs and offline tests
updated accordingly.
Add skills/anysearch (vendored from anysearch-ai/anysearch-skill v3.1.1)
so agents get general/anonymous search, parallel batch search, vertical
domain search, and full-page extract. Includes MAINTENANCE.md and
docs/anysearch-skill.md. No internal/, ui/, or migration changes.
Disclosure: AnySearch Open Source Bounty Claim goclaw#001.
Creating a predefined agent with a description kicks off summoning in a
background goroutine. It finishes 10-20 seconds after the create response and
writes SOUL.md, IDENTITY.md, CAPABILITIES.md and the agent frontmatter.
For a client that manages agents as code that is data loss. Such a client
brings its own context files and writes them through agents.files.set as soon
as create returns; summoning lands afterwards and replaces them. Nothing
fails: no error from the API, nothing in the logs, and the agent quietly runs
on generated text instead of the committed role. The only way to avoid it
today is to poll the agent status until it leaves "summoning" — which means
every automation has to know that summoning exists at all.
POST /v1/agents now accepts an optional "summon" field. Absent or true keeps
today's behaviour, so existing clients are unaffected. False creates the agent
active with the seeded template files and starts no LLM call; summoning is
still available afterwards via POST /v1/agents/{id}/resummon.
The status fallback also now resets an explicitly requested "summoning" status
when nothing will summon the agent: that status is only ever cleared by the
summoner, so such an agent would otherwise stay stuck in it forever.
Surface parity: CLI and WS RPC (agents.create) create agents without a
description and never summoned, so they need no flag; the web UI create flow
wants summoning and keeps the default.
- docs: device.pair.update and the `permanent` option on approve in
docs/04-gateway-protocol.md, docs/19-websocket-rpc.md and
websocket-protocol.md; the paired-device TTL row in docs/09-security.md
now mentions the admin opt-out.
- store.ErrPairedDeviceNotFound: SetPairingPermanent wraps it in both stores.
device.pair.update maps it to NOT_FOUND and any other store error to
INTERNAL, so a DB failure no longer reads as "not found".
- web UI: approve and make-permanent/set-expiry now toast the server error
and reload the list in `finally`. A partially applied approve (paired, but
the permanent write failed) shows up in the table instead of leaving the
dialog dead-ended.
- SQLite ListPaired: a stored expiry that fails to parse stays 0 (expires,
date unknown) rather than being mistaken for permanent; the UI renders it
as "--" instead of a 1970 date.
Tests: gateway handler error mapping (NOT_FOUND / INTERNAL / OK), sentinel
checks in the PG and SQLite store tests, SQLite unreadable-expiry case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every approved pairing expires 30 days after approval (pairedDeviceTTL),
and nothing lets an operator change that: owners of a personal setup have
to re-pair their own Telegram account and browser every month. The schema
already treats a NULL paired_devices.expires_at as "never expires" (IsPaired
and the prune both check for it), only no code path ever writes NULL.
- store: SetPairingPermanent(senderID, channel, permanent) on PairingStore,
PG and SQLite. permanent=true clears expires_at, false restarts the
default TTL from now. An already expired pairing is not revived.
PairedDeviceData gains expires_at (Unix ms, null = never), returned by
ApprovePairing and ListPaired.
- gateway: device.pair.approve accepts `permanent`; new admin RPC
device.pair.update {senderId, channel, permanent} for existing pairings,
classified in permissions/policy.go next to the other pairing methods.
- mcp: pairing approve tool accepts `permanent`.
- web UI (Nodes): "Never expires" switch in the approve dialog, an
Expires column, and a Make permanent / Set expiry action per device.
ConfirmDialog takes optional children for the switch.
The default stays 30 days; permanence is an explicit operator choice.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
read_video called with a url parameter always failed under an agent budget:
tool:read_video: cannot verify streamed native media against the agent
context budget (no in-memory payload to count); refusing to send
read_audio and read_document appeared to work but only for very small files.
Both symptoms share one root cause.
5ca433b8 introduced the complete-input invariant and, to keep out-of-band
media honest, appended the standard-base64 encoding of the payload as a
synthetic guard-only message. That message is counted as text. Measured against
the bundled BudgetCounter, base64 costs 0.956 tokens per raw byte, so an 11.6 MB
video counted 11,161,058 tokens and exceeded a 200k, a 1M and a 2M window alike.
The practical ceiling was roughly 200 KB at a 200k window while videoMaxBytes is
100 MB. The read_video URL transport streams to the File API without buffering,
so it had no bytes to count at all and was routed to a helper that refused to
send whenever an agent budget was present.
Media is now priced by what a provider actually bills for it. Every number below
is either a published per-unit rate or a published provider limit; none is
derived from a byte size, because a static-image video compresses arbitrarily
small and no bitrate floor exists. An earlier revision of this branch tried one
and a 120-second, 20,627-byte clip priced at 526 tokens against a true 31,560.
internal/mediabudget prices one payload:
- Video and audio: ffprobe measures duration. 263 tokens/second for video,
32 for audio, both published for static processing at 1 FPS.
- PDF: pdfinfo counts pages at 258 tokens/page. When the page count cannot be
read, the charge is the proven ceiling of 1000 pages, which is the most a
provider will accept and therefore the most it can bill.
- Video and audio that cannot be measured are refused. Upstream already
refused unverifiable native media for every URL on the Gemini streamed path;
this narrows that refusal from every URL to only what genuinely cannot be
measured, rather than removing it. PDF differs because its page count is
cheaply measurable and its ceiling is small enough to stay usable.
A remote video is measured without downloading it: two ranged GETs, 512 KB from
the head and 512 KB from the tail, written into a sparse temp file sized to the
declared total and handed to ffprobe. The tail matters because every container
that puts its index at the end keeps it there: head-only probing under-reports
mpeg by 98% and ogg by 79%, and a plain prefix makes ffprobe under-report a
30-second WAV as 0.74 seconds because it clamps to the bytes it can see. Sizing
the temp file to the real total fixes that. Verified end to end on an 11,673,105
byte MP4 served by nginx: 1 MB of ranged reads yielded duration 40.000000,
identical to ffprobe reading the whole URL, for a charge of 10,520 tokens. Those
requests reuse the existing SSRF-safe path, security.WithPinnedIP plus
security.NewSafeClient(0), and no URL is ever handed to an external binary.
Beyond the reported bug, two pre-existing gaps let large media reach a provider
almost unpriced. ExecuteWithChain treated every callProvider error as a provider
failure and advanced to the next entry, and the non-Gemini branches of read_video
and read_document reserved without pricing their payload at all. Measured under a
20,000-token window before this change, a 40 MB video and a 1000-page PDF each
reached a provider charged about 1,600 tokens. Budget refusals are now terminal
in the chain and every media branch prices its payload, so both reach no provider
at all. Genuine provider failures still fail over.
Known limits, stated rather than discovered:
- The /Type /Page scan that guards against a forged /Count is a floor, not a
bound. Pages inside a compressed object stream are invisible to it, and a
9,484-byte PDF built that way is charged 258 tokens for 1000 pages. A real
pdfinfo reads such files correctly; the scan only ever raises a probed count.
- A video URL whose origin does not serve byte ranges is now refused on the
non-Gemini path too, and the refusal is terminal. Upstream forwarded such URLs
unpriced. A HEAD giving only a size is not enough to price one.
- read_video and read_audio require ffprobe. Docker images install it by default
except the base variant; bare binaries and the desktop build do not ship it.
- A hostile origin can craft a container ffprobe reads as about one second.
Reservation.Reconcile overwrites the estimate with the provider's reported
usage, so this weakens the gate rather than defeating it.
Byte ceilings videoMaxBytes, audioMaxBytes and documentMaxBytes are unchanged.
No new module dependency; ffprobe and pdfinfo are optional runtime probes.
OpenAI shipped GPT Image 2.5 on 2026-09-08 as two model IDs rather than one:
flare is tuned for speed, sunburst for editing precision. Both accept text and
image input and run on /images/generations, /images/edits and the Responses
API, which is the path the native image_generation tool already uses.
Flare replaces gpt-image-2 as the default because OpenAI reports higher quality
at roughly half the latency, which suits create_image's one-shot call. Both
gpt-image-2 and gpt-image-1.5 stay on the whitelist, so a saved image_model
keeps working untouched.
Also in this change:
- The rejection message is now built from the whitelist itself instead of a
hand-written string that drifts whenever a model is added.
- codex_build.go uses the DefaultImageModel constant instead of hardcoding
"gpt-image-2", so the inline chat path cannot drift from the shared default.
- Both new models join the isEditModel branch so reference images route to
/images/edits.
- ParamField gains an optional labelKey, letting the dropdown read its label
from i18n instead of showing English inside a translated UI. It passes a
defaultValue, so a missing key degrades to the literal rather than breaking.
- Locale keys added to all five catalogs; ko was missing the whole
mediaChain.imageModel* group.
Surface parity: the desktop UI is unchanged because sortable-provider-card
renders no param form and passes params through as Record<string, unknown>.
The CLI is unchanged because cmd/ exposes no image-model flag. The API contract
is unchanged because image_model lives in the free-form params map; only the
backend whitelist widened.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(teams): scope team tasks to the delegation origin, not the delivery channel
A delegated run is delivered on the internal "delegate" channel while its real
origin is preserved separately — buildAgentLinkRunRequest is explicit about it:
// preserves the origin's authorization-bearing identity while keeping
// delegation on its internal delivery channel.
Channel: "delegate",
WorkspaceChannel: req.Channel,
WorkspaceChatID: req.ChatID,
The team tools did not consult that origin, so a task created by a lead reached
through delegate was stamped with the delivery channel. Nothing is registered
for it, so every notification about that task — completion, failure, blocker
escalation, ask_user — was dropped:
unknown channel for outbound message channel=delegate
The delegatee's first answer still arrived, because it travels back as the
delegation result rather than through a channel; everything the lead said after
the delegation closed was lost. One session produced 16 such drops, including a
blocker escalation the user needed to see.
Add OriginChannelFromCtx / OriginChatIDFromCtx next to the existing workspace
scope propagation helpers, and use them where team scoping and notification
routing are decided: task records, list/search scoping, dispatch fallbacks,
event payloads, escalation tasks, ask_user and leader notifications. The
resolution is the identity when no delegation origin is present, so
non-delegated flows are unchanged.
Deliberately not touched: the channel used in authorization decisions
(checkTeamAccess, requireLead, approve/reject lead bypass). Those ask "how did
this call arrive", not "where should the answer go", and widening them is a
separate question.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(teams): cover delegated-lead completion routing and origin isolation
Triage on #1529 named two gates this PR had not met: regression coverage for
delegated lead task completion, and a guard that unrelated origins cannot
receive the notification. The existing tests only covered what create persists.
TestDelegatedLeadCompletionNotifiesOrigin completes a task raised in a delegated
context and asserts the completion event is addressed to the caller's origin.
Reverting team_event_helpers.go to the delivery channel fails it with
"addressed to delegate/system".
TestCompletionNotificationStaysWithinItsOwnOrigin puts two tasks from different
origins on one board, completes one, and asserts no completion notification
carries the other origin.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(teams): scope team tasks to the delegation origin, not the delivery channel
A delegated run is delivered on the internal "delegate" channel while its real
origin is preserved separately — buildAgentLinkRunRequest is explicit about it:
// preserves the origin's authorization-bearing identity while keeping
// delegation on its internal delivery channel.
Channel: "delegate",
WorkspaceChannel: req.Channel,
WorkspaceChatID: req.ChatID,
The team tools did not consult that origin, so a task created by a lead reached
through delegate was stamped with the delivery channel. Nothing is registered
for it, so every notification about that task — completion, failure, blocker
escalation, ask_user — was dropped:
unknown channel for outbound message channel=delegate
The delegatee's first answer still arrived, because it travels back as the
delegation result rather than through a channel; everything the lead said after
the delegation closed was lost. One session produced 16 such drops, including a
blocker escalation the user needed to see.
Add OriginChannelFromCtx / OriginChatIDFromCtx next to the existing workspace
scope propagation helpers, and use them where team scoping and notification
routing are decided: task records, list/search scoping, dispatch fallbacks,
event payloads, escalation tasks, ask_user and leader notifications. The
resolution is the identity when no delegation origin is present, so
non-delegated flows are unchanged.
Deliberately not touched: the channel used in authorization decisions
(checkTeamAccess, requireLead, approve/reject lead bypass). Those ask "how did
this call arrive", not "where should the answer go", and widening them is a
separate question.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(teams): cover delegated-lead completion routing and origin isolation
Triage on #1529 named two gates this PR had not met: regression coverage for
delegated lead task completion, and a guard that unrelated origins cannot
receive the notification. The existing tests only covered what create persists.
TestDelegatedLeadCompletionNotifiesOrigin completes a task raised in a delegated
context and asserts the completion event is addressed to the caller's origin.
Reverting team_event_helpers.go to the delivery channel fails it with
"addressed to delegate/system".
TestCompletionNotificationStaysWithinItsOwnOrigin puts two tasks from different
origins on one board, completes one, and asserts no completion notification
carries the other origin.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(teams): grant a delegated lead read access to its own team
Fixes#1535. Second attempt: the first one set RunRequest.TeamWorkspace on the
delegated run, which loop_context.go rejects outright and deliberately so — that
field also becomes ToolWorkspace, which would displace the exchange outputs
directory. Every delegation then died at setup with "invalid delegation artifact
workspace". This does not go near that field.
The separation the lead needs already exists in the code. A lead addressed
directly resolves its team at loop_context.go:285-321 and gets
ctx = tools.WithToolTeamWorkspace(ctx, wsDir)
ctx = tools.WithToolTeamRoot(ctx, teamRoot)
with no WithToolWorkspace call — a read allowance, not an override. Only the
!isArtifactDelegation gate keeps a delegated lead out of it. So the grant is
made inside the artifact branch instead, from the same inputs, and the guard on
req.TeamWorkspace stays exactly as strict as it was.
Both paths are needed. The workspace alone covers only the lead's own chat leaf;
the deliverable in the report sits at teams/<id>/system/review-....md, written by
a member under a different chat scope. buildAllowedPrefixes adds the team root
for reads and not for writes (filesystem.go:366-370), which is precisely the
asymmetry this case wants: the lead reads the team's output, per-chat write
isolation is untouched.
Team ID is deliberately not set. It switches on the workspace interceptor's
write validation, file-change broadcast and task attachment (workspace_
interceptor.go:36,114,203) — none of which a delegated run should trigger, and
none of which reading needs.
Ambiguity is not resolved silently, as triage asked: an agent leading more than
one active team gets nothing. store.GetTeamForAgent would answer in one call but
it is ORDER BY (lead_agent_id = $1) DESC LIMIT 1 and would quietly pick a team.
Hermeticity is unchanged: send_file and message still refuse inside an artifact
run, publication still happens through delegation outputs on completion.
Tests: TestInjectContext_DelegatedLeadReadsTeamWithoutDisplacingOutputs drives
injectContext through the guard that broke the first attempt and pins both
halves — the run is not refused, and ToolWorkspace stays the outputs directory.
Resolution rules, the shared/isolated path shapes, read-not-write on the team
root, and send_file staying blocked are covered alongside.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(teams): pin that the delegated lead resolver does not re-scope by tenant
l.dataDir arrives already tenant-scoped from resolver.go (config.TenantDataDir),
which is why resolveDelegatedLeadTeamRead must not apply TenantLayer itself.
That was carried only by a comment, here and in the sibling branch of
injectContext, and nothing failed if someone added the layer back.
The failure it guards against is quiet: the path stays plausible, it just gains
a second tenant segment — teams/<id> under tenants/<slug>/tenants/<slug> — so
the lead is handed a directory its own tasks never write to, and reads come back
empty rather than denied.
Verified the assertion has teeth by reintroducing the layer: the subtest fails
with "tenant segment appears 2 times". Prompted by #1546 in this same area,
where a contract that lived outside the unit under test went unpinned and the
feature shipped green and did nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(teams): let a human cancel and retry stuck team tasks from the dashboard
A task that ended up blocked, stale, failed or cancelled could not be
recovered from the UI: the dashboard only deletes terminal tasks and
approves/rejects in_review ones. CancelTask and ResetTaskStatus existed
in the store but were reachable only through the lead agent's team_tasks
tool, which needs the full task UUID the dashboard never shows (#506).
Two WS RPCs, wired to Retry / Cancel buttons in the task detail dialog:
- teams.tasks.cancel — any task not yet completed/cancelled. Optional
reason is stored as the result and posted as a comment so the lead
sees it on the board. Dependents are unblocked by the store; they are
not dispatched here because the dashboard has no agent turn.
- teams.tasks.retry — stale / failed / cancelled / in_review, plus
blocked tasks that nothing blocks any more. A human comment is
required: it is posted on the task and appended to the assignment
prompt, so the assignee gets the missing answer in the same message.
Optional agentId reassigns; the lead is refused as assignee (same
guard as teams.tasks.assign).
ResetTaskStatus (pg + sqlite) now also accepts blocked, guarded on the
handler side by an empty blocked_by. Handler tests cover both actions
with a stub store: reason/comment persistence, required comment,
status gating, blocked_by guard, lead guard, cross-team IDOR.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(permissions): classify teams.tasks.cancel and teams.tasks.retry as write methods
Every WS method must be classified in internal/permissions/policy.go,
otherwise the router answers UNAUTHORIZED for every role. Caught on a
live gateway (the drift test TestMethodRole_DriftCoverage… flags it, I
had not run that package). Both sit next to approve/reject/assign as
operator-level write methods; the write-methods test now pins them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(teams): let the retry dialog pick the assignee
A cancelled or stale task may have no owner (created unassigned, or the
member was removed). teams.tasks.retry then needs agentId, which the
dialog did not send, so Retry would fail with "task has no assignee".
The retry dialog now shows an assignee select (team members minus the
lead, defaulting to the current owner) and passes agentId only when it
differs from the owner. Retry is offered only when there is someone to
pick.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
TestHandleTeammateMessageSchedulesStreamedRun fails CI intermittently under
-race. Two separate races, both in the test rather than in what it exercises:
1. It shared `gotReq` between the scheduler's RunFunc and the assertions. The
RunFunc runs on a scheduler goroutine, and the announce loop schedules a
second run after the teammate one, so the write could land while the test
was reading — and the second run also closed an already-closed channel.
Requests now arrive over a buffered channel: no shared state, and a second
run cannot clobber the first.
2. The deferred sched.Stop() ran while handleTeammateMessage's background
goroutine was still calling Schedule, so Lane.Submit's wg.Add raced
Lane.Stop's wg.Wait. The gateway drains BgWg before stopping the scheduler
(gateway_consumer.go waits on it; sched.Stop is an outer defer in
gateway.go); the test skipped that step. It now drains too, which also makes
the test match the shutdown order it is meant to represent.
Reproduced before the fix with `-race -count=60 -cpu=1,4` (fails within a few
iterations) and clean afterwards over `-count=200 -cpu=1,2,4`, plus the whole
cmd package under -race.
Nothing in production changes; `Stream: true` and the channel-manager assertion
are untouched.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The short identifier (T-015-cc8e) carries only the last four hex chars
of the UUID, while every agent-facing surface — the team_tasks tool
(get / retry / cancel / comment) and the teams.tasks.* RPCs — takes the
full UUID. A human looking at the dashboard therefore had no way to name
a task to the lead in chat. Show the UUID next to the identifier and
status badges, copy it to the clipboard on click, with a short "copied"
confirmation.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
OpenAI rejects non-default temperature on gpt-5 / gpt-5-chat with
HTTP 400 ("Unsupported value: 'temperature' does not support 0.7
with this model. Only the default (1) value is supported.").
skipTemp in buildRequestBody only covered gpt-5-mini / gpt-5-nano and
the o-series, so base gpt-5 and gpt-5-chat fell through and received
config.DefaultTemperature (0.7). On channel integrations the inbound
error is suppressed, leaving the agent silently dead.
The gpt-5.X flagship variants (gpt-5.1, gpt-5.4, gpt-5.5) keep their
temperature — only the base gpt-5 and gpt-5-chat aliases are locked
to the default.
Regression cases added for gpt-5, gpt-5-chat, gpt-5-chat-latest and
provider-prefixed forms; gpt-5 moved to the locked list in
TestBuildRequestBody_TemperatureKeptForNonReasoningModels per the
maintainer triage on #1460.
Fixes#1460
* feat(delegate): add action=list, scoped to the originating chat
Fixes#1545. A delegation result was addressable only by the UUID returned once
in a tool result, which the calling model had to carry forward by hand. One
mistyped character orphaned a completed, durably stored result with no way back:
`get` answers "delegation result not found", and there was nothing else to ask.
Observed in production with a 31B-class caller — one flipped character, and
separately a splice of the previous delegation's tail onto the next one's prefix.
`spawn`, the sibling async mechanism over the same table, has had list/wait/cancel
all along; `delegate` had delegate/get.
Scope is tenant and calling agent, as get already resolves, plus the origin chat.
The chat rather than the session, for three reasons:
- It survives a session reset. Deferring long work, clearing the context and
coming back to ask for status is ordinary use; a SessionKey predicate would
return nothing exactly then — when the handle is most likely already lost.
- It keeps chats apart, which is the enumeration boundary #1525 is about: there
spawn's list filters on the parent agent key alone and ignores the session,
so one chat reads another chat's task text.
- It does not carry a conversation between chats. A delegation raised in a team
chat stays visible in that team chat and does not surface in someone's DM
with the same agent. History stays where it began.
In a direct chat that separates users as well, since the chat ID is per person.
Group chats deliberately show the group what the group started.
get is left as it was, deliberately. #1525 is an enumeration defect — no prior
knowledge needed and task text is disclosed. get is access through an unguessable
handle, and adding a predicate there would break fetching a result by an ID kept
across a reset, which is the very failure this fixes.
No schema change: the origin fields are already persisted by
createDelegateCompletion. ListByParent is filtered in Go behind a cap of 20,
which suits handle recovery; a dedicated predicate would be the next step if this
ever needs to page.
Tests pin the chat boundary, the session-reset case, refusal when there is no
chat to scope to (without querying the store), and the cap. The fake store leaves
ListBySession embedded and nil, so a refactor back to session scoping panics
rather than passing quietly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(delegate): list needs its own store query — ListByParent excludes delegations
The action shipped in 506ecba4 always returned an empty list. It read through
SubagentTaskStore.ListByParent, whose SQL carries
AND COALESCE(metadata->>'completion_kind', 'subagent') <> 'delegate'
ListBySession carries the same clause. Both serve spawn and filter delegations
out on purpose, so no listing in the store could return one — only Get by ID
reaches a delegation. The feature was a no-op in production while its unit tests
were green, because the fake store returned whatever rows the fixture supplied
and never reproduced the predicate that does the damage.
Found by running it against a live cluster: an async delegation was created,
`get` returned it completed with its result, and `list` reported zero.
Adds ListDelegationsByChat to the interface and to both implementations, with
the inverse predicate plus origin_chat_id, and points the tool at it. Chat scope
and delegate-only selection now live in the query rather than in a Go filter over
whatever the store happened to return; an empty chat yields no rows instead of
falling back to everything.
Tests are where the fix matters most:
- internal/store/sqlitestore exercises the real SQL. It pins that
ListDelegationsByChat returns the chat's delegation, that a spawn in the same
chat is not one, that another tenant's identically named chat stays invisible,
and — the part that would have caught this — that ListByParent and
ListBySession still do not return delegations, so the complementarity is
documented rather than assumed.
- the tool's fake now panics if ListByParent or ListBySession is called, so a
regression to either fails loudly instead of quietly listing nothing, and it
applies the chat and kind predicates itself so fixtures behave like the store.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
ChatStream parsed the tool-call accumulator map with a positional loop
(for i := 0; i < len(accumulators); i++ { acc := accumulators[i] }),
assuming OpenAI-compatible providers always emit contiguous zero-based
tc.Index values in streamed delta.tool_calls.
Some OpenAI-compatible gateways (observed with a self-hosted point-p1
proxy) emit non-contiguous or non-zero-based indices. When that
happens accumulators[i] misses a populated key and returns a nil
*toolCallAccumulator, and the next line dereferences acc.rawArgs on
that nil pointer, causing a SIGSEGV panic. The panic kills the
in-flight chat run's goroutine; the gateway's tracing safety-net logs
'finalizing orphan trace' and the user's message goes unanswered.
Fix: range over the accumulators map directly instead of indexing by
position, and skip nil entries defensively.
Verified: built a full backend image (make build-full equivalent,
CGO_ENABLED=0 go build -tags embedui) from this branch, deployed it
against a real point-p1-backed agent, and confirmed multiple
multi-tool-call turns (use_skill, read_file, exec x3) that previously
triggered the panic now complete via v3.run.completed with no panic
and no safety-net trace.
pip >= 23.0 added the PEP 668 --break-system-packages flag. On older pip
builds (e.g. macOS Command Line Tools Python) and on the pip `list`
subcommand of every version, the flag is rejected with
"no such option: --break-system-packages", breaking skill dependency
install/check paths entirely.
- dep_installer: route pip installs through pipRunInstall, which retries
once without the flag when pip rejects it (issue #956)
- pip_update_executor: same retry for the upgrade path
- pip_update_checker: drop the flag from `pip3 list --outdated` - list is
read-only, PEP 668 never applies, and no pip accepts the flag there
- add pip_flags.go helpers + unit tests, and legacy-pip integration tests
reproducing the original failure
Fixes#956