Runs on models without a registered tokenizer (e.g. 9router brand models)
ended with the generic "Agent couldn't generate a response" fallback even
though the real request used about 55% of the context window.
PruneStage counted history with TokenCounter, which falls back to a
chars/2 heuristic for unregistered models and overcounted about 1.8x.
Once over budget it ran memory flush (~35s, invisible in traces), then
mid-loop compaction, which cannot summarize a history made only of tool
call/result pairs. The callback reported the untouched history as
compacted, PruneStage still saw it over budget and returned AbortRun
before any LLM call, and FinalizeStage replaced the empty reply with the
fallback.
- PruneStage and ContextStage overhead count with the request guard's
BudgetCounter. PruneStage no longer controls loop flow; the final
request guard in ThinkStage decides.
- CompactMessages returns ErrNotCompacted when history is unchanged.
Callers stop counting it as a compaction and do not retry it in the
same run, while post-run summarization still sees the pressure.
- When the guard exhausts every reduction step, ThinkStage stops the run
with a localized chat.context_budget_exceeded notice instead of an
error, so the run's tool results are still persisted. The stop reason
marks the trace and agent span as error; team tasks, cron and
heartbeat treat it as a failure via RunOutcome.Failure().
- Memory flush and mid-loop compaction emit event spans.
- Web and desktop UIs treat an unset context_pruning as enabled (the
backend default since 7639a8c0), keep it unset when untouched, and can
re-enable pruning after it was turned off.
* feat(tools): add Serply as a web_search provider
Serply is a SERP API returning Google web results on a tenant-supplied
key. It joins the existing chain the same way Exa, Tavily and Brave do:
keyed from config_secrets, skipped silently when no key is stored, and
selectable through the provider parameter for cross-engine corroboration.
Freshness maps onto Google's tbs recency filter, which Serply forwards
upstream. Only the pd/pw/pm/py shortcuts are sent: upstream ignores the
equivalent cdr date range and answers with unfiltered results, so a
range leaves tbs unset rather than naming a filter the results do not
honour.
The settings form gained a fifth sortable provider, so the locked
DuckDuckGo card now derives its position number instead of hardcoding
it, and a stored provider_order is reconciled against the full provider
list rather than against Parallel alone.
* test(tools): use a locally named round trip stub in the serply test
The serply cancellation test borrowed parallelRoundTripFunc from
web_search_parallel_test.go. Define serplyRoundTripFunc alongside the
test that uses it so the file does not depend on another provider's
test helper.
Register Requesty (https://router.requesty.ai/v1) next to OpenRouter on the
existing OpenAI-compatible transport: config and env vars
(GOCLAW_REQUESTY_API_KEY, GOCLAW_REQUESTY_BASE_URL), secret masking, DB and
in-memory registration, CLI setup, doctor, placeholder provider, OpenAPI enum,
Web/Desktop provider lists and docs.
The models list merges Requesty managed policies (GET /models/managed) with
the key's catalog from GET /models.
- docs: device.pair.update and the `permanent` option on approve in
docs/04-gateway-protocol.md, docs/19-websocket-rpc.md and
websocket-protocol.md; the paired-device TTL row in docs/09-security.md
now mentions the admin opt-out.
- store.ErrPairedDeviceNotFound: SetPairingPermanent wraps it in both stores.
device.pair.update maps it to NOT_FOUND and any other store error to
INTERNAL, so a DB failure no longer reads as "not found".
- web UI: approve and make-permanent/set-expiry now toast the server error
and reload the list in `finally`. A partially applied approve (paired, but
the permanent write failed) shows up in the table instead of leaving the
dialog dead-ended.
- SQLite ListPaired: a stored expiry that fails to parse stays 0 (expires,
date unknown) rather than being mistaken for permanent; the UI renders it
as "--" instead of a 1970 date.
Tests: gateway handler error mapping (NOT_FOUND / INTERNAL / OK), sentinel
checks in the PG and SQLite store tests, SQLite unreadable-expiry case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every approved pairing expires 30 days after approval (pairedDeviceTTL),
and nothing lets an operator change that: owners of a personal setup have
to re-pair their own Telegram account and browser every month. The schema
already treats a NULL paired_devices.expires_at as "never expires" (IsPaired
and the prune both check for it), only no code path ever writes NULL.
- store: SetPairingPermanent(senderID, channel, permanent) on PairingStore,
PG and SQLite. permanent=true clears expires_at, false restarts the
default TTL from now. An already expired pairing is not revived.
PairedDeviceData gains expires_at (Unix ms, null = never), returned by
ApprovePairing and ListPaired.
- gateway: device.pair.approve accepts `permanent`; new admin RPC
device.pair.update {senderId, channel, permanent} for existing pairings,
classified in permissions/policy.go next to the other pairing methods.
- mcp: pairing approve tool accepts `permanent`.
- web UI (Nodes): "Never expires" switch in the approve dialog, an
Expires column, and a Make permanent / Set expiry action per device.
ConfirmDialog takes optional children for the switch.
The default stays 30 days; permanence is an explicit operator choice.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
OpenAI shipped GPT Image 2.5 on 2026-09-08 as two model IDs rather than one:
flare is tuned for speed, sunburst for editing precision. Both accept text and
image input and run on /images/generations, /images/edits and the Responses
API, which is the path the native image_generation tool already uses.
Flare replaces gpt-image-2 as the default because OpenAI reports higher quality
at roughly half the latency, which suits create_image's one-shot call. Both
gpt-image-2 and gpt-image-1.5 stay on the whitelist, so a saved image_model
keeps working untouched.
Also in this change:
- The rejection message is now built from the whitelist itself instead of a
hand-written string that drifts whenever a model is added.
- codex_build.go uses the DefaultImageModel constant instead of hardcoding
"gpt-image-2", so the inline chat path cannot drift from the shared default.
- Both new models join the isEditModel branch so reference images route to
/images/edits.
- ParamField gains an optional labelKey, letting the dropdown read its label
from i18n instead of showing English inside a translated UI. It passes a
defaultValue, so a missing key degrades to the literal rather than breaking.
- Locale keys added to all five catalogs; ko was missing the whole
mediaChain.imageModel* group.
Surface parity: the desktop UI is unchanged because sortable-provider-card
renders no param form and passes params through as Record<string, unknown>.
The CLI is unchanged because cmd/ exposes no image-model flag. The API contract
is unchanged because image_model lives in the free-form params map; only the
backend whitelist widened.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(teams): let a human cancel and retry stuck team tasks from the dashboard
A task that ended up blocked, stale, failed or cancelled could not be
recovered from the UI: the dashboard only deletes terminal tasks and
approves/rejects in_review ones. CancelTask and ResetTaskStatus existed
in the store but were reachable only through the lead agent's team_tasks
tool, which needs the full task UUID the dashboard never shows (#506).
Two WS RPCs, wired to Retry / Cancel buttons in the task detail dialog:
- teams.tasks.cancel — any task not yet completed/cancelled. Optional
reason is stored as the result and posted as a comment so the lead
sees it on the board. Dependents are unblocked by the store; they are
not dispatched here because the dashboard has no agent turn.
- teams.tasks.retry — stale / failed / cancelled / in_review, plus
blocked tasks that nothing blocks any more. A human comment is
required: it is posted on the task and appended to the assignment
prompt, so the assignee gets the missing answer in the same message.
Optional agentId reassigns; the lead is refused as assignee (same
guard as teams.tasks.assign).
ResetTaskStatus (pg + sqlite) now also accepts blocked, guarded on the
handler side by an empty blocked_by. Handler tests cover both actions
with a stub store: reason/comment persistence, required comment,
status gating, blocked_by guard, lead guard, cross-team IDOR.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(permissions): classify teams.tasks.cancel and teams.tasks.retry as write methods
Every WS method must be classified in internal/permissions/policy.go,
otherwise the router answers UNAUTHORIZED for every role. Caught on a
live gateway (the drift test TestMethodRole_DriftCoverage… flags it, I
had not run that package). Both sit next to approve/reject/assign as
operator-level write methods; the write-methods test now pins them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(teams): let the retry dialog pick the assignee
A cancelled or stale task may have no owner (created unassigned, or the
member was removed). teams.tasks.retry then needs agentId, which the
dialog did not send, so Retry would fail with "task has no assignee".
The retry dialog now shows an assignee select (team members minus the
lead, defaulting to the current owner) and passes agentId only when it
differs from the owner. Retry is offered only when there is someone to
pick.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
The short identifier (T-015-cc8e) carries only the last four hex chars
of the UUID, while every agent-facing surface — the team_tasks tool
(get / retry / cancel / comment) and the teams.tasks.* RPCs — takes the
full UUID. A human looking at the dashboard therefore had no way to name
a task to the lead in chat. Show the UUID next to the identifier and
status badges, copy it to the clipboard on click, with a short "copied"
confirmation.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
- Give Name/Status/Tokens/Spans/Time explicit widths so Name cannot shove others off-screen
- Show status label next to the status icon
Fixes#762
Signed-off-by: Vedant Madane <6527493+VedantMadane@users.noreply.github.com>
Editing the description in onboarding step 3 deselected the preset, and the
display name, agent key and emoji were derived from that selection. Losing it
dropped them to a hardcoded "Fox Spirit" / "fox-spirit" / fox emoji, so picking
the Writer preset and rewriting its prompt created an agent named Fox Spirit.
The step also offered no way to name the agent at all.
Identity is now editable state that a preset only prefills:
- Add display name, emoji and agent key inputs to both onboarding steps, with
the key auto-derived from the name until it is edited by hand.
- Editing the prompt only clears the preset highlight; the identity stays.
- Block submit on an empty name or an invalid slug instead of falling back.
- Give the web presets stable English agent keys, so the slug no longer depends
on the UI locale, and strip the emoji prefix off the preset label before it
becomes a display name.
- While a preset is active and its field is untouched, the name and prompt still
follow a locale switch.
Reuses the existing agents:create.* keys, present in en/vi/zh on both surfaces.
The workstations page was written against snake_case field names the
gateway never served. Go's json decoder drops unknown members silently,
so create requests arrived with an empty WorkstationKey and were rejected
with "workstationKey is required", while list rows read undefined and
rendered a blank key, "backend.undefined" and "Invalid Date".
Align the TypeScript contract on the camelCase shape that both the WS
handler and store model already use:
- Rename Workstation, CreateWorkstationParams and every field read on the
workstations page to workstationKey/backendType/createdAt/updatedAt.
- Nest update params under `updates`; a flattened body decoded to an empty
map and was rejected with "no updates provided".
- Replace the Identity File input, which mapped to a field SSHMetadata does
not accept, with a private key or password credential path.
- Add the image field required by DockerMetadata and map the container name
onto host and the daemon endpoint onto socketPath, so Docker workstations
can be created at all.
- Extract the payload builder so the wire contract is unit-testable, and
move validation strings into the en/vi/zh catalogs.
Covered by tests asserting the create payload shape and list rendering
against a verbatim gateway response body.
Co-authored-by: Eddy Lockwood <doakythanh@gmail.com>
The v3 pipeline compacts session history mid-loop (prune_stage +
final-request guard) but only mutates the run's message buffer, never the
session store. Each turn reloads full history and re-compacts from scratch:
message_tokens climb 129k->156k across turns while every turn compacts back
down to ~60k. The lossy compaction differs per run, degrading the agent.
The same missing persistence stalls episodic memory: the cumulative
compaction count never advances, so the episodic worker's idempotency key
(sessionKey:count) is pinned and every cycle after the first is skipped.
Observed on live traffic: 8 run.completed since deploy, 0 new episodic.
Fixes, all reusing existing machinery (no new store methods, no migrations):
- Bug A: emitSessionCompleted reads cumulative GetCompactionCount (matching
the legacy v2 path) instead of the per-run counter that resets to 0.
- Bug B/anti-loop: finalize passes state.Prune.MidLoopCompacted into
maybeSummarize; under pressure it lowers the trigger to a unit-aligned
threshold (compactionInputCap - overhead, same MaxRequestShare the guard
uses) so the compaction is PERSISTED via the existing TruncateHistory +
IncrementCompaction path. Defensive floor prevents over-compaction on
pathological config; tool-result-only bloat still skips (history-only).
- Bug C: SourceID embeds the count (sessionKey:count) so the eventbus dedup
key advances per compaction cycle instead of swallowing rapid same-session
turns within the 5m TTL.
Tests: episodic compaction, maybe_summarize pressure, request budget.
go build (PG + sqliteonly), go vet, go test -race all green.
* feat(hooks/observe): add structured output validation hook (ObserveHook)
- hooks/types.go: Add ObserveHook interface + BuiltInHookType enum
- hooks/dispatcher.go: HookDispatcher emits ObserveHook with PhaseResult payload
- pipeline/observe_stage.go: ObserveStage emits hook after ObserveResult built
- pipeline/substates.go: ObserveStageResult carries hook results + validation errors
- hooks/config.go: BuiltInHookTypeObserve added to BuiltInHookType enum
PhaseResult carries structured output, token usage, tool calls + validation errors
emitted post-ObserveStage. ObserveHook implementations can validate structured
output against schemas, detect tool-call loops, enforce token budgets, etc.
Hook fires after ObserveStage produces ObserveResult, before results propagate
to next stage. ValidationError returned by hook halts pipeline and propagates
error to caller without further stage execution.
Co-Authored-By: Claude <noreply@anthropic.com>
* ui(hooks): add post_model_response event to web UI
- Add event to Zod schema, filter dropdown, and form dialog
- Implement conditional test panel UI for model response payload
- Add translations (en/zh/vi) for new test panel fields
- Updated beta description to reference the new event
Co-Authored-By: Claude <noreply@anthropic.com>
* feat(mcp): add MCP CRUD server exposing goclaw resources at /api/mcp/ with Bearer token auth
and X-GoClaw-Tenant-Id header, default to master tenant
* feat(mcp): add goclaw_skills_write_file tool to edit skill files on disk
The CRUD MCP server's goclaw_skills_update only touched skill DB metadata,
with no way to edit a skill's SKILL.md/file content on the filesystem. Extract
the versioned write logic from the web UI's skill file editor
(SkillsHandler.handleWriteFile) into skills.WriteVersionedFile so both
surfaces share identical validation and versioning, and expose it as a new
MCP tool.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Show an ephemeral "agent is working" indicator (thinking/searching/
generating/analyzing…) in Bitrix24 chat while the agent processes, so
users on this non-streaming channel aren't left staring at silence
until the final reply. Send behavior is unchanged; no LLM call, no DB.
Generic layer (internal/channels):
- New optional ActivityIndicatorChannel interface.
- HandleAgentEvent routes run.started->THINKING, tool.call->mapped
status, tool.result->ANALYZING; static tool->status mapping.
- Conditional heartbeat ticker fills LLM-inference gaps (re-sends only
when idle), 5s per-run throttle caps call rate; ticker stopped on
terminal events AND in UnregisterRun (safety net for missed terminals).
Bitrix24 (internal/channels/bitrix24):
- OnActivityEvent calls imbot.v2.Chat.InputAction.notify best-effort,
drop-on-limit (raw Client.Call, no retry) so cosmetic notifies never
steal leaky-bucket capacity from real message sends.
- Per-channel activity_indicator toggle (default on).
Web UI + docs: dashboard toggle, i18n en/vi/zh, channels doc section.
Tests: 26 unit tests incl. UnregisterRun-stops-ticker regression.
The delete affordance next to each pending portal in the channel-create
form was wired to onClick, but Radix Select closes the popover on the
SelectItem's own pointerdown handler. By the time the click event would
fire on the nested button, the popover — and the button — are gone.
onPointerDown on the button additionally called e.preventDefault(),
which further blocked the browser from synthesising the click at all.
Result: clicking the trash icon looked completely dead. Keyboard users
were already covered by the Delete/Backspace onKeyDown handler on the
SelectItem itself; only mouse users were locked out.
Fire setPendingDeleteName from onPointerDown (with stopPropagation +
preventDefault so the parent SelectItem doesn't also select the pending
portal), and keep the onClick handler so the button stays activatable by
Enter/Space when it takes focus. This mirrors the pattern Radix uses for
its own nested interactive elements (e.g. DropdownMenu.CheckboxItem).
Surface parity:
- Gateway server: N/A because the change is UI-only.
- API contract: N/A because bitrix.portals.delete already exists.
- Web UI: the single component that owned the broken handler.
- CLI/runtime: N/A because no CLI reads this UI state.
Co-authored-by: DangTinh311 <dangtinh31193@gmail.com>
* docs(i18n): design spec for Russian (ru) language support
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(i18n): register ru locale constant and IsSupported
* feat(i18n): scaffold ru catalog (en placeholders)
* feat(i18n): translate ru backend catalog to Russian
* fix(i18n): polish ru catalog grammar per review
* feat(i18n): wire ru into system messages, config enum, parity test
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(i18n): add ru to web language constants and system-message locales
* feat(i18n): scaffold ru web locale dir (en placeholders)
* feat(i18n): wire ru resources into web i18next config
* feat(i18n): translate ru web locale to Russian (41 namespaces)
* fix(i18n): normalize ru web terminology (арендатор, привязка)
* fix(i18n): restore dropped 'named' in ru setup OAuth hint
* feat(i18n): wire ru into desktop i18next and language pickers
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(i18n): add ru desktop locale (Russian, 16 namespaces)
* test(i18n): extend traces/hooks locale guards to cover ru
* chore(deploy): add repeatable fork-dev local build+deploy tooling
Adds a one-command, idempotent way to run goclaw from a locally built
image that merges fresh upstream dev with un-merged fork features
(Telegram trigger-words PR #1383 + Russian locale), independent of the
published :latest image.
- scripts/deploy-forkdev.sh: fetch origin/dev -> merge feature branches
into local/fork-dev worktree -> push fork/dev -> build -> up -d.
- docker-compose.forkdev.yml: overlay pinning image to goclaw:forkdev
with pull_policy:never so :latest never silently replaces the build.
- .gitignore: track the deploy helper (past deploy-*.sh), ignore the
../.goclaw-forkdev integration worktree.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(telegram): agent-callable message reaction (message action=react)
Adds a "react" action to the message tool so an agent can set an emoji
reaction on an existing message (e.g. mark its own status post 👍 once
fully paid). Mirrors the ChannelEditor wiring: ReactionSetter capability
-> Manager.ReactToMessage -> telegram Channel.ReactToMessage (validates
against Telegram's supported reaction set; ✅/❌ are rejected).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(config): remove ignored session scoping controls (#1405)
Co-authored-by: Collective Developer <man@collective.dev>
* feat(telegram): agent-callable message reaction (message action=react) (#1407)
Adds a "react" action to the message tool so an agent can set an emoji
reaction on an existing message (e.g. mark its own status post 👍 once
fully paid). Mirrors the ChannelEditor wiring: ReactionSetter capability
-> Manager.ReactToMessage -> telegram Channel.ReactToMessage (validates
against Telegram's supported reaction set; ✅/❌ are rejected).
Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci: build & publish fork image to GHCR (temporary, pre-upstream)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(memory): use 'russian' FTS config for memory_chunks recall
Memory keyword search (memory_search.go FTS + the generated tsv column)
used the 'simple' text-search config, which does no stemming. Russian
grammatical cases therefore never matched: a genitive query token
("Дианы") could not hit the stored nominative token ("Диана"), so
memory_search returned nothing for natural queries even though the fact
was saved. This is the sole recall path on deployments without an
embedding provider (ChatGPT-OAuth/codex cannot embed), so recall was
effectively dead there.
Switch both sides to the Snowball 'russian' config so index and query
lexemes stem to the same root ("диан"), restoring case-insensitive
recall. Migration 000094 recreates the generated tsv column (a GENERATED
expression cannot be ALTERed in place; ADD COLUMN backfills all rows).
Surface parity: desktop/SQLite N/A — SQLite memory search uses LIKE, a
separate mechanism unaffected by the PG regconfig.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: ナム <namnhfreelancer@gmail.com>
Co-authored-by: Collective Developer <man@collective.dev>
The provider verify call and the provider models-list call used hardcoded
timeouts (30s and 15s) that were too low for slow models. Introduce a
tenant-scoped, UI-editable setting `providers.request_timeout_sec`
(default 30) stored in system_configs, consumed by both handlers, and
overlaid from config.json like tts.timeout_ms. Adds a field to the System
Settings modal (all 4 locales) so operators can raise it without a
redeploy.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
* feat(bitrix24): replace two MCP text inputs with a filtered dropdown [B24:2794]
Bitrix24 channel creation used to demand two hand-typed strings —
mcp_server_name and mcp_base_url — plus zero indication of which MCP
servers can actually auto-onboard. Typos silently disabled provisioning
and admin had to know which servers implement /api/auto-onboard.
Ship a single dropdown backed by mcp_servers.require_user_credentials,
plus the machinery to make it work end-to-end.
Phase 1 — DB & store
* Promote require_user_credentials from settings JSONB to a top-level
column on mcp_servers (PG migration 000089, SQLite migration v54).
* Backfill from existing settings blobs so no admin needs to re-tick.
* Add MCPServerData.RequireUserCredentials to the Go store layer, plumb
through Create / Get / GetByName / List / Update on both stores,
extend the export DTO, and add require_user_credentials to the HTTP
allowlist.
* Bump RequiredSchemaVersion 87 -> 89 (jumping 88, which was on disk
but not wired) and SchemaVersion 53 -> 54 with an
idempotentColumnMigration guard.
Phase 2 — Bitrix24 channel factory
* Add MCPServerID (UUID string) to bitrix24 InstanceConfig, keep
MCPServerName + MCPBaseURL as legacy fallback with a "deprecated"
doc comment.
* Factory validation accepts either mcp_server_id alone or the legacy
pair; half-config still fails fast.
* initMCPProvisioner prefers GetServer(id) and sources the base URL
from mcp_servers.url when the id path is used. Legacy name path
unchanged so pre-migration configs keep working.
* Log line now carries mcp_server_id + require_user_credentials so
operators can eyeball the wiring.
Phase 3 — Frontend types & MCP form
* Add optional top-level require_user_credentials to MCPServerData /
MCPServerInput in both ui/web and ui/desktop/frontend types.
* mcp-form-dialog reads the top-level flag first and falls back to
settings.require_user_credentials so cached responses from
pre-upgrade backends still render correctly.
* On submit send both the top-level flag AND the legacy settings
entry so mid-rollout backends stay consistent.
Phase 4 — Bitrix24 channel form dropdown
* New mcp-select field type + MCPServerSelect component. Uses the
shared useMCP() react-query cache and filters client-side to
servers whose require_user_credentials is true (OR settings
JSONB during the migration window).
* Explicit "None (disable MCP provisioning)" option so admins can
clear the binding without editing config JSON.
* Legacy mcp_server_name / mcp_base_url text inputs kept in the
Advanced panel, relabelled "(legacy)" with pointer help text.
Phase 5 — channel_instances.config backfill
* PG migration 000090 and SQLite migration v55 rewrite existing
bitrix24 channel_instances.config to add mcp_server_id by
resolving mcp_server_name against mcp_servers, tenant-scoped
via agents.tenant_id (channel_instances doesn't carry tenant_id
directly).
* Idempotent — only touches rows already carrying
mcp_server_name that lack mcp_server_id. Legacy keys are left
in place so provisioner.go can still fall back for unmigrated
or future-created legacy configs.
* down.sql drops the mcp_server_id key. Provisioner immediately
reverts to the legacy fallback path.
Tests
* provisioner_test.go: three new cases exercise the mcp_server_id
path (invalid UUID string, valid UUID with missing row, valid
UUID with a per-user row). fakeMCPStore gains a serversByID
map and a real GetServer implementation.
* Existing legacy-config tests unchanged and still green.
Verification
* go build ./... && go build -tags sqliteonly ./...
* go vet ./internal/mcp/... ./internal/channels/bitrix24/...
./internal/store/... ./internal/http/...
* go test ./internal/mcp/... ./internal/channels/bitrix24/... -> ok
* Live-tested against a local docker image on the goclaw-deploy
postgres. Migrations 89 + 90 applied cleanly. Three existing
bitrix24 channels (bitrix-sales / nguyen-dao-openline / tieu-vi)
had their configs backfilled with the b24-syn-mcp UUID and the
provisioner boots with require_user_credentials=true. UI dropdown
correctly shows only b24-syn-mcp (the only server with the flag
ticked) alongside a "None" clearer option.
Surface parity
* Gateway server: store + factory + provisioner + HTTP allowlist.
* API contract: adds require_user_credentials + mcp_server_id
as optional fields on existing routes. No new endpoints.
* Web UI: MCP form + Bitrix24 channel form + shared types.
* CLI/runtime package: N/A because no CLI subcommand reads the
mcp_server_id field.
* fix(bitrix24): derive auto-onboard base URL from mcp_servers.url origin [B24:2794]
The Phase 2 refactor swapped provisioner base-URL sourcing from the
legacy per-channel MCPBaseURL config field (which historically stored the
MCP server's ORIGIN, e.g. https://mcp.example.com) to mcp_servers.url,
which stores the JSON-RPC ENDPOINT the agent loop dials (e.g.
https://mcp.example.com/mcp). The two are semantically different but
share a single column.
mcp_client.newMCPClient then appends "/api/auto-onboard" to whatever
baseURL it receives, so the id-path started POSTing to
".../mcp/api/auto-onboard" — 404 for every per-user credential mint and
refresh. Existing users kept working only until their cached access
tokens expired.
Fix: derive the origin (scheme://host[:port]) from server.URL before
handing it to the auto-onboard client. The legacy path is untouched
because MCPBaseURL from channel config is already the origin.
https://b24-mcp-dev.synity.so/mcp -> https://b24-mcp-dev.synity.sohttps://mcp.example.com/mcp/ -> https://mcp.example.comhttps://mcp.example.com -> https://mcp.example.com
Table-driven test covers six shapes plus four error cases (empty,
whitespace-only, no scheme, no host). Updated the existing
TestInitMCPProvisioner_MCPServerID fixture to seed a URL with the /mcp
subpath so it regression-guards the same code path.
Verified live: user 614 sent a message that triggered the expired-cred
refresh branch; goclaw logged "self-refreshed user credentials
created=false" and the agent immediately reported
"mcp.user_tools_loaded user=614 tools=2". Before this fix the same event
logged 'auto-onboard failed: mcp auto-onboard: 404 Not Found'.
Surface parity:
- Gateway server: provisioner + one new helper (deriveAutoOnboardBaseURL).
- API contract: N/A because the wire shape hasn't changed.
- Web UI: N/A because the UI still writes mcp_server_id verbatim.
- CLI/runtime: N/A because no CLI reads the derived base URL.
---------
Co-authored-by: DangTinh311 <dangtinh31193@gmail.com>
slugify's docstring promises "no leading/trailing dashes" and the sibling
isValidSlug rejects a trailing dash, but slugify only stripped leading
dashes. Any input ending in a non-[a-z0-9] character (e.g. "hello world!",
a trailing space, "trailing---") produced a slug ending in "-" that then
failed isValidSlug, causing invalid persisted keys and spurious validation
errors on agent/channel/provider/cron/mcp name fields. Add a symmetric
trailing-dash strip, matching the Go helpers (strings.Trim(s, "-")).
* feat(skills): allow editing SKILL.md content for database-managed skills
Adds the ability to edit a skill's actual SKILL.md content directly
from the web UI, restricted to database/tenant-managed skills only --
bundled/system skills remain read-only.
Backend: new PUT /v1/skills/{id}/files/{path...} endpoint
(handleWriteFile in skills_versions.go), writing to the current
version directory of a non-system skill only (403 for is_system
skills), with the same path-traversal/symlink/ownership guards as the
existing read endpoint. Bumps skill store version and emits
cache-invalidate + audit event on success.
Frontend: Content tab in the skill detail dialog gets an "Edit
content" button, disabled with a tooltip for system skills. Editing
fetches the RAW file content via the existing read endpoint (which
does not strip YAML frontmatter, unlike the stripped preview shown in
the Content tab by default) to avoid silently destroying frontmatter
on save. Save/cancel follow the established catch-and-toast error
pattern (no uncaught promise rejections, matching PR #1346's config
save button fix).
i18n keys added to all 4 locales (en, ko, vi, zh).
* fix(skills): bump version on content edit, fix scroll regression in content view
handleWriteFile previously overwrote the current skill version in place
instead of creating a new immutable version, inconsistent with the
skill_manage tool's patch action. Now creates a new version directory
and updates the DB pointer, matching that convention.
Also fixes a scroll regression in the skill detail dialog's Content
tab introduced by the edit-mode textarea -- both the read-only preview
and edit textarea are now properly scrollable within the dialog.
* fix(skills): refresh version display after save, fix dialog height constraint for scrolling
Version bump now correctly reflects in the UI immediately after save
without requiring a page reload. Fixed the actual root cause of the
scroll issue: the dialog's height wasn't bounded, so overflow-y-auto
on inner content had no effect since nothing constrained the dialog's
total height in the first place.
---------
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Ollama has two separate request-building code paths: OllamaProvider
(native /api/chat, used for num_ctx control) and OpenAIProvider
(OpenAI-compat /v1/chat/completions). An earlier fix disabled thinking
mode by hardcoding think=false, but only in the OpenAI-compat path --
OllamaProvider.buildRequest() never set the think field at all, so
reasoning-capable models (qwq, deepseek-r1) defaulted to visible
chain-of-thought reasoning regardless of that fix. Confirmed live via
a docker-engineer agent streaming full reasoning traces despite the
existing disable.
Replaced the hardcoded always-off behavior with a provider-level
tri-state setting (llm_providers.settings.thinking_enabled: unset =
default off, explicit true/false overrides), configurable via the
provider's Advanced settings dialog. Both OllamaProvider.buildRequest()
and OpenAIProvider.buildRequestBody() now read and respect this same
setting, so the toggle works regardless of which Ollama code path a
given deployment routes through.
Added tests for setting parsing (unset/true/false/malformed) and both
provider request-builders' handling of the override.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
* fix(prompt): remove duplicate Team Members section from system prompt
The TEAM.md context file already provides the Members section with better
formatting. The system prompt section was redundant and inconsistently
formatted.
- Remove buildTeamMembersSection() call from system prompt building
- Remove unused buildTeamMembersSection() function
* fix(mcp): wire store to manager for prompt preview tool visibility
MCP manager needs database access to query configured servers for tool visibility
in prompt preview.
- Add SetStore() method to MCP manager
- Wire pgStores.MCP to manager after initialization
- Add debug logging for MCP initialization flow
* fix(systemprompt): hide Tooling section when agent has no tools
The Tooling section header and boilerplate were displayed even when the
agent had no tools available. Add early return in buildToolingSection()
to skip the section entirely when toolNames is empty.
This reduces prompt noise for agents with no tool access.
* feat(mcp): cache tool descriptions for prompt preview visibility
MCP tools now show descriptions in prompt preview without requiring a live
server connection.
- Add CacheToolDescriptions() method to MCPServerStore (PG + SQLite)
- Cache tool descriptions in settings['tool_cache'] when server connects
- Use cached descriptions in ListToolsForAgent (fallback: hints → cache → global)
- Descriptions are auto-populated from live server manifest on connection
- Admins can still override via tool_hints in server settings
* test(mcp): add CacheToolDescriptions to MCPServerStore test fakes
Commit 5ce410b6 added CacheToolDescriptions() to the MCPServerStore
interface but missed updating test mock implementations, breaking
go vet across internal/mcp, internal/agent, internal/http, and
internal/channels/bitrix24.
Add no-op implementations matching each fake's existing style.
* fix(providers): enforce per-agent tool policy for Claude CLI provider
The Claude CLI provider (stdio+MCP bridge) was not enforcing per-agent
tool policy, unlike other providers where the policy-filtered tool list
already drives the system prompt's Tooling section.
Two gaps closed:
1. --disallowedTools was previously skipped entirely when no MCP config
path was resolved, letting the CLI subprocess run with its full
native toolset (Bash, Edit, Read, Write, Glob, Grep, WebFetch,
WebSearch) regardless of agent policy. It's now unconditional and
derived from the agent's actual allowed-tools list (state.Tool.AllowedTools),
mapped to Claude CLI's native tool names.
2. The MCP bridge server executed any tool call without checking the
calling agent's policy. It now resolves the agent's policy from
context (via the existing HMAC-verified agent lookup) and denies
calls to tools outside that agent's allowed set, logging
security.mcp_bridge_denied on denial.
Both gaps were closed using existing plumbing (PolicyEngine.WouldAllow,
AgentData.ParseToolsConfig, the bridge context middleware) — no new
cross-cutting mechanism was introduced.
* feat(web): show MCP/tool schemas in system prompt preview dialog
The prompt-preview API response includes a separate `tools` field
(the actual JSON schemas sent to the LLM as the tools API parameter)
alongside `prompt` (the system prompt text), but the web UI only
rendered `prompt`, silently dropping the tools list.
Add a collapsible Tools section to both the full-screen System Prompt
dialog and the inline agent-detail preview, showing tool count, name,
description, and expandable parameter schema per tool. i18n keys added
to en/vi/zh locales.
* fix(prompt): render pinned skills on bootstrap turns
Pinned skills are documented (web UI copy) as "always inlined in the
system prompt", but the entire Skills section was gated behind
!cfg.IsBootstrap, so pinned skill XML never appeared on bootstrap
turns (first message of a session) despite the promise.
Separate pinned-skill rendering from bootstrap-suppressed guidance:
- Bootstrap + pinned skills present: render pinned XML only, no
search/manage guidance (which stays suppressed as before)
- Non-bootstrap: unchanged behavior
- Minimal/none modes: pinned skills always render regardless of
bootstrap state
Add regression tests covering all four prompt modes on bootstrap
turns, plus a non-bootstrap guard confirming existing behavior is
preserved.
* fix(skills): resolve managed skills directory per-tenant, not master-only
skills.Loader was wired at startup to scan a single fixed directory
(the master tenant's managed-skills dir), making any skill belonging
to a non-master tenant invisible to both pinned-skills prompt
resolution and skill_search/use_skill, regardless of DB visibility
settings.
- Loader now resolves the calling tenant's managed-skills directory
per-call via context (store.TenantIDFromContext), never enumerating
other tenants' directories
- Skill cache is now tenant-keyed to prevent slug collisions and
cross-tenant cache leaks across tenants using the same skill slug
- gateway_setup.go passes the root data dir instead of a pre-resolved
master-tenant path
Write-side tooling (skill_manage, publish_skill) was already correctly
tenant-scoped per-operation — no changes needed there.
Added TestLoader_ManagedSkills_TenantIsolation proving two tenants
with same-slug/different-content skills never see each other's
content, including after cache population from a different tenant's
lookup.
Known follow-up (not in this commit): skill_search's BM25 index is
still a single process-global index shared across tenants, which is
a related but separate cross-tenant search-result leak requiring its
own scoped fix (per-tenant index maps + threading tenant context
through ensureIndex/rebuildIndex).
* fix(tools): scope skill_search BM25 index per-tenant
SkillSearchTool held a single process-global BM25 index built once
from whichever tenant's context first triggered ensureIndex, then
reused for all subsequent Execute() calls regardless of caller —
leaking one tenant's skill search results into another's, the
search-path counterpart to the managed-directory bug fixed in
7b4668ad.
- index/lastVersion are now keyed per-tenant (map[uuid.UUID]*tenantIndexState)
- ensureIndex resolves the calling tenant from context and only
builds/reads that tenant's index entry, never touching another
tenant's cached state
- Builtin/bundled skills remain visible in every tenant's index
(Loader already merges those tiers correctly per 7b4668ad)
Loader.Version() remains a single global counter — a version bump in
one tenant causes unnecessary rebuilds in others but does not cause
cross-tenant leakage, an acceptable tradeoff to avoid scope creep.
Added TestSkillSearchTool_TenantIsolation proving two tenants with
same-slug/different-content skills never see each other's search
results, including after cache population from a different tenant.
* fix(tools): fix group-spec expansion in tool policy engine
PolicyEngine.registry was only ever set via SetRegistry(), which was
never called in production (only in one test) — so pe.registry was
permanently nil in production. Every group-expansion helper
(applyProfile, intersectWithSpec, unionWithSpec, subtractSpec,
expandSpec, matchDenySpec, filterByCapability) silently dropped any
"group:*" spec entry instead of expanding it when registry was nil.
Concretely: "group:mcp" (auto-injected into agentToolPolicy.AlsoAllow
for any agent with MCP tools) never resolved to real tool names, so
MCP tools connected successfully and appeared in prompt text (which
reads the registry directly, bypassing PolicyEngine) but were never
included in the actual ChatRequest.Tools payload sent to the LLM —
confirmed live via mcp.agent.tools_loaded tools=6 immediately followed
by mcp.filtered_tools mcp_defs_count=0 in the same request. This
affects any agent relying on group-based grants, not just MCP.
PolicyEngine is a shared/global singleton used concurrently across
all agents (constructed once at gateway startup), so mutating a
registry field per-call would be a data race. Fix instead threads the
registry as an explicit parameter from FilterTools down through all
internal group-expansion helpers, and adds IsDenied/WouldAllow
registry parameters, removing the dead SetRegistry() mechanism
entirely.
Also fixes group expansion for the per-user-MCP-tools path: FilterTools
is sometimes called with a userToolOverlay wrapping a *Registry rather
than a *Registry directly; added Unwrap() to userToolOverlay so the
new registry-resolution logic works for both cases.
Added 4 tests proving group expansion works via the threaded parameter
alone (no SetRegistry): plain registry allow, userToolOverlay allow,
deny-side group expansion, and the WouldAllow bridge-server path.
Blast radius note: this restores intended access for every agent
configured with group:* specs (group:mcp, group:vault, group:goclaw,
group:coding, etc.) that were silently inert before. Existing agent
configs relying on group grants will gain the tool access they were
nominally already configured for.
* fix(mcp): cache tool descriptions from pool-connected servers too
5ce410b6 added tool-description caching (for prompt-preview visibility)
only inside connectServer. connectViaPool — the separate connect path
used when MCP connections go through the shared pool — never got the
same caching hook, even though it shares the same underlying
connectAndDiscover wire handshake.
Confirmed live: cloudflare/docker connected via connectViaPool and
received real descriptions over the wire, but prompt preview still
showed blank descriptions because this path never wrote to the cache.
Also removes the temporary mcp.connect.raw_tool debug log added
earlier this session for diagnosing the same issue — no longer needed
now that the root cause is fixed.
* chore(skills): remove temporary pinned-skills diagnostic logging
Confirmed live: pinned skills (caveman, infra-ansible-knowledge) now
resolve correctly end-to-end for tenant-scoped agents. Debug logging
added to trace the resolution chain is no longer needed.
* fix(tools): deny always wins over AlsoAllow group grants
AlsoAllow's unionWithSpec could reintroduce a tool explicitly listed
in Deny, since it added tools back from allTools without re-checking
deny specs. Previously masked because AlsoAllow's group-expansion was
also broken (fixed in f7af95de this session) — group specs silently
expanded to nothing, so this ordering bug never manifested. Now that
group expansion works, an admin-denied tool that's also reachable via
a group:* AlsoAllow entry (e.g. group:mcp) would silently reappear.
Re-apply deny-spec subtraction as a final step after AlsoAllow union,
for both global and per-agent policy, so deny always wins regardless
of which allow mechanism tries to add a tool back.
Added tests proving global and per-agent Deny correctly override an
overlapping AlsoAllow group grant, while sibling non-denied tools in
the same group remain allowed.
* fix(mcp): enumerate cached tools instead of wildcard placeholder in prompt preview
ListToolsForAgent (the prompt-preview path) collapsed any server with
an empty ToolAllow (unrestricted grant — the common case) into a
single "server__*" placeholder entry, even when tool_cache already
had every real tool name and description from connect time (5ce410b6,
8ffd67b9). This meant agents with unrestricted MCP server access never
saw individual tool names or descriptions in prompt preview, forcing
trial-and-error tool usage.
When ToolAllow is empty and tool_cache is populated, enumerate every
cached tool (skipping any explicitly denied) and emit one
MCPToolPreviewInfo per tool, matching the construction logic already
used for the ToolAllow-non-empty case. Falls back to the single
placeholder only when tool_cache is also empty (server never
connected).
The live (non-preview) conversation path, buildMCPToolDescs, does not
have this bug — it resolves tool identity from the live connected
registry, never from ToolAllow, so no placeholder shortcut exists
there.
Added tests covering both the cache-populated enumeration case and
the no-cache placeholder fallback.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(prompt): filter alias re-injection and add MCP schemas to preview ToolDefs
Prompt-preview's tools: schema array (PreviewResult.ToolDefs) had two
bugs, both preview-only — confirmed live conversations already use
correctly-filtered tool payloads via buildToolsPayload/PolicyEngine.FilterTools,
unaffected by either bug:
1. Alias re-injection iterated ALL registry aliases globally with no
check against the deny-filtered toolNames list, letting denied
tools reappear via their alias name (e.g. a denied canonical tool
still showing up as its Claude-Code-compat alias like Bash/Edit/Write/Read).
Now skips any alias whose canonical tool isn't in the filtered set.
2. MCP tool descriptions (from the store-based, connection-free
ListToolsForAgent path) only ever fed the prompt TEXT section,
never got converted into ToolDefinition schema objects — so the
tools: array never showed MCP tools at all, even when the MCP
section of the prompt text correctly listed them. Now appends a
ToolDefinition per MCP tool with name+description populated and a
placeholder {"type":"object"} Parameters schema, documented as
preview-only (real parameter schemas require a live MCP connection,
only available during an actual conversation turn).
Added tests proving denied-tool-alias exclusion and MCP tool inclusion
in preview ToolDefs.
* feat(skills): inline full content for pinned skills instead of pointer-only
Web UI documented "Pinned skills are always inlined in the system
prompt", but BuildPinnedSummary was just BuildSummary with an
allowlist filter — same as the general searchable skill list:
name/description(truncated)/location pointer only, requiring
use_skill+read_file round trips to get actual content. No code path
inlined real SKILL.md content for pinned skills specifically.
BuildPinnedSummary now reads and inlines full SKILL.md content
(frontmatter stripped) per pinned skill inside <skill_instructions>
tags. Per-skill (10000 bytes) and total (30000 bytes) size caps fall
back to the original pointer-only format with a note when a skill is
too large to inline, so oversized skills degrade gracefully instead
of blowing the prompt budget.
Added tests: full-content inlining, size-cap fallback, and tenant
isolation for the new inline path (mirroring the existing managed-
skills tenant isolation test).
* fix(mcp): real parameter schemas and full policy enforcement in prompt preview
Three interconnected fixes to prompt preview, none affecting the live
conversation path (which was already correct):
1. Real MCP parameter schemas instead of a useless empty placeholder.
tool_cache previously stored only name+description; extended to
also capture the real JSON Schema from the MCP server's tools/list
response (CachedToolInfo{Description, Parameters}) at connect time,
for both direct-connect and pool-connect paths. Preview now shows
complete, real input schemas instead of {"type":"object"} with no
properties — verified against actual wire data from a live
cloudflare MCP server showing genuinely rich schemas (zone_id,
type, name, content, ttl, proxied, all typed with descriptions and
correct required arrays) that were previously being discarded.
Backward-compatible: old-shape cache entries degrade gracefully to
description-only rather than crashing, self-healing on next
connect.
2. Global tool deny now enforced in preview. BuildPreviewPrompt
previously hand-rolled a partial policy reimplementation
(per-agent deny only, explicitly skipping the full PolicyEngine
"because runtime state isn't available in preview") — but
PolicyEngine.WouldAllow already handles this per-tool-name without
needing channel context. A tool denied via the global config (not
per-agent) would appear in preview despite being correctly denied
in every real conversation. Preview now calls WouldAllow per
candidate tool, with a graceful per-agent-deny-only fallback when
no PolicyEngine is wired (e.g. in tests).
3. MCP tools now also subject to the same policy check — previously
the MCP tool supplement (store-based, connection-free tool listing)
added MCP tool names unconditionally, bypassing WouldAllow entirely,
so a denied MCP tool could still appear in preview.
Added tests for all three: real-schema presence, global-deny exclusion
for both core and MCP tools, and backward-compat cache handling.
* fix(prompt): remove redundant per-tool MCP enumeration from prompt text
Now that MCP tool schemas in the tools: API parameter are real and
complete (1290d4f1), the ## MCP Tools prompt-text section's per-tool
"- mcp_x__y: description" enumeration is pure duplication with zero
added value — the model already gets each tool's real schema
(including description) via tools:.
buildMCPToolsInlineSection now keeps only the behavioral instructions
that aren't expressible via JSON schema and thus aren't duplicated:
prefer-MCP-over-core-tools guidance, and the optional-parameter
guidance (don't guess/fill optional fields). The per-tool name+
description enumeration loop is removed. Section still only appears
when the agent has MCP tools (len(cfg.MCPToolDescs) > 0, unchanged
gate).
Updated tests to assert the enumeration is gone while the behavioral
instructions remain; ToolDefs assertions are now the authoritative
check for MCP tool allow/deny filtering behavior (prompt text no
longer enumerates names at all).
* test(agent): update TeamContextInjection test for removed Team Members section
TestBuildSystemPrompt_TeamContextInjection asserted the presence of a
'Team Members' prompt-text section that was intentionally removed in
71d33180 (duplicate of the canonical TEAM.md-context-file Members
section, which has better formatting). The test was never updated to
match, causing it to fail on every run since. Moved the assertion
from wantIn to wantNotIn for the 3 affected subtests -- BuildSystemPrompt
correctly no longer renders team-member roster info directly; that
info now comes exclusively from the TEAM.md context-file mechanism,
outside this test's isolated scope.
* fix(prompt): resolve real registry for WouldAllow calls in preview
BuildPreviewPrompt's two WouldAllow calls hardcoded reg=nil, silently
breaking group:* expansion (e.g. group:mcp) needed to resolve the
AlsoAllow grant production actually uses to grant MCP tool access
(resolver_helpers.go's agentToolPolicyWithMCP injects
AlsoAllow: ["group:mcp"]). With reg=nil, WouldAllow could match
literal tool names fine (the 18 core/static tools) but could never
resolve group-based grants, so every MCP tool silently failed
WouldAllow and was excluded from preview -- confirmed live via curl:
18 tools returned, zero mcp_* ones, for an agent with genuinely
working MCP access in real conversations.
Live conversations were never affected -- internal/mcp/bridge_server.go's
WouldAllow call already correctly passes a real registry.
Fix resolves a real *tools.Registry from deps.ToolLister via
tools.ResolveConcreteRegistry (the same helper used at the live call
site), passing it to both WouldAllow calls instead of nil. Falls back
to nil gracefully for test mocks that don't implement the full
ToolExecutor interface, preserving existing test behavior.
Added a test proving an MCP tool granted via the exact production
AlsoAllow: ["group:mcp"] pattern is now correctly included in preview
ToolDefs, where the old reg=nil bug would have silently excluded it.
* fix(prompt): use literal deny check for MCP tools in preview, not group expansion
The MCP-tools policy gate added in 1290d4f1 called WouldAllow with a
real registry (per 9b5fd2eb), which requires group:mcp expansion
against that registry to grant access via the production
AlsoAllow: ["group:mcp"] pattern. But MCP tools are only ever
registered into ephemeral per-agent registry clones at live connection
time (manager_connect.go) -- never into the shared/global registry
preview uses. group:mcp always resolved empty in preview's
connection-free context, so WouldAllow denied every MCP tool --
confirmed live via tool-name diff: live conversations correctly
included all 6 MCP tools, preview included zero.
MCP access-granting is already correctly handled by
ListToolsForAgent's own per-server tool_allow/tool_deny grant logic
(confirmed working correctly earlier this session). The preview gate
only needs to catch the narrower case of a literally-denied tool name
via global/per-agent policy config -- it never needed group
expansion. Replaced WouldAllow with IsDenied(nil, name, agentPolicy),
which forces a pure literal-name match with zero registry dependency,
matching the existing usage pattern already established elsewhere in
policy.go. This class of bug cannot recur: there's no registry-passing
code path left in this check to silently reintroduce group-expansion
dependence.
Added a test proving MCP tool inclusion in preview is independent of
group-expansion outcome (no AlsoAllow: group:mcp needed for a
non-denied tool to appear).
Also reverts the temporary loop.filtered_tool_names/
preview_prompt.filtered_tool_names diagnostic logging used to capture
the live-vs-preview tool-name comparison that diagnosed this bug.
* fix(http): preserve real MCP parameter schemas through HTTP preview adapter
mcpPreviewAdapter.ListToolsForAgent (the HTTP-layer glue converting
mcp.MCPToolPreviewInfo to agent.MCPToolPreviewInfo for BuildPreviewPrompt)
only copied RegisteredName and Description, silently dropping
Parameters -- a bug present since this adapter was introduced
(2499d0be/7e250244), unrelated to today's other MCP preview fixes.
This was masked until d5fc6344 fixed MCP tools being excluded from
preview entirely (a separate bug) -- once MCP tools started appearing
again, this pre-existing adapter gap became visible: tools showed up
correctly, but always with the bare {"type":"object"} placeholder
instead of their real cached schema (confirmed live: update_dns_record
missing its 7 real properties).
One-line fix: copy Parameters through in the adapter's struct literal.
Added a regression test constructing a real *mcp.Manager with
populated tool_cache, asserting the adapter's output preserves
specific real schema properties (not just non-nil Parameters) --
verified this test fails without the fix and passes with it.
---------
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
fix(usage): repair cost analytics and display precision
- Automatic OpenRouter pricing sync with cost backfill for traces/snapshots/events
- Live usage data merging for current-hour dashboard accuracy
- 2-decimal API cost formatting across usage/overview pages
- Comprehensive test coverage (PG, SQLite, HTTP, UI)
Merged by github-maintain automation.