ThinkStage folded each turn's usage into state.Think.TotalUsage but summed
only prompt/completion/total/thinking tokens, dropping CacheReadTokens,
CacheCreationTokens and the PromptTokensIncludeCachedSegments flag. The
aggregated RunResult.Usage is the source for webhook usage responses and
usage_events analytics, so async multi-turn agent runs reported zero cache
tokens even when traces showed heavy prompt-cache use (cache is read
per-response for spans, but never summed into the aggregate).
Sum cache read/creation across turns and OR the include-cached flag (a
provider-level property, consistent across a run). Restores correct cache
accounting for webhook usage and usage_events without any envelope change.
The internal sandbox.Config already supports a custom container Workdir
that doubles as the workspace mount target, but config.SandboxConfig
(the JSONB-backed per-agent config) did not expose it. Without this,
agents must use the hardcoded /workspace mount point even when their
scripts expect a different in-container path.
Adds an optional "workdir" JSON field that maps through to
sandbox.Config.Workdir. Backward compatible — existing agents without
workdir set fall back to the /workspace default.
Co-authored-by: tuannt23065 <tuannt23065@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Reparent tool_call spans under their producing llm_call span via
RunState.CurrentLLMSpanID so a model turn's trace visibly contains the
tool calls it triggered.
- Wire internal/hooks EmitHookSpan into the dispatcher writeExec so every
hook execution (pre/post tool use, etc.) produces a trace span with input,
console output, decision, error, and duration for agent troubleshooting.
- Fix usage_events_span_id_fkey violations: usage events were inserted
synchronously referencing a span flushed ~5s later, silently dropping
token/cost data. Route them through the collector so they flush after
spans in the same cycle. No schema migration; FK preserved.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Adds a "react" action to the message tool so an agent can set an emoji
reaction on an existing message (e.g. mark its own status post 👍 once
fully paid). Mirrors the ChannelEditor wiring: ReactionSetter capability
-> Manager.ReactToMessage -> telegram Channel.ReactToMessage (validates
against Telegram's supported reaction set; ✅/❌ are rejected).
Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Fix image_generation nil-pointer crash on codex agents: the sentinel now
carries a name-only Function so the many sites reading td.Function.Name never
nil-deref; codex_build still branches on Type.
- Agent-declared trigger words in IDENTITY.md wake the bot in groups without an
@mention (whole-word, Cyrillic-aware; text + caption), cached per-agent 60s.
- channel_post support with a synthetic sender + a recover() guard so a
malformed update can't crash the gateway.
- message tool action=edit (editMessageText + editMessageCaption fallback),
targeting the replied-to message via reply_to_message_id.
- message tool topic=<name> posts into a named forum topic; topics learned from
forum_topic_created into channel_contacts and resolved by name.
Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
bitrix.portals.create only persisted the DB row — it never called
Router.RegisterPortal, so a portal created while the gateway was
already running stayed invisible to PortalByDomain/PortalByKey until
the next full restart. The only other RegisterPortal call site is
BootstrapPortals, which runs once at boot. Result: the OAuth install
callback 404'd with "unknown portal" for any portal created through
the self-service UI mid-uptime — hit live on web1trang.bitrix24.com.
Fix: inject registerPortal/unregisterPortal func fields into
BitrixPortalsMethods (mirrors the existing gatewayPublicURL func()
pattern), defaulting to bitrix24.WebhookRouter() + NewPortal/
RegisterPortal (same construction path as BootstrapPortals) and
Router.UnregisterPortal respectively. handleCreate calls
registerPortal after a successful Create; handleDelete calls
unregisterPortal after a successful Delete for symmetry. Both are
best-effort — a registration failure is logged, not returned as an
RPC error, since the DB row is already valid either way.
Injected via func fields (not calling bitrix24.WebhookRouter()
inline) so tests can stub live-router registration without touching
that process-wide singleton, whose test-reset helper isn't exported
outside the bitrix24 package.
Tests: 4 new cases covering register-on-create, best-effort failure
handling, unregister-on-delete, and no-unregister-when-delete-blocked.
All 20 tests in bitrix_portals_test.go pass; go build (both PG and
sqliteonly tags) and go vet clean.
Co-authored-by: DangTinh311 <dangtinh31193@gmail.com>
* fix: warm Zalo group-approval cache so forwards from non-group origins route correctly
send.go picks ThreadTypeGroup vs ThreadTypeUser for a target chat ID by
checking BaseChannel's approvedGroups cache (or an explicit "group_id"
metadata key). That cache is only ever populated from INBOUND group
traffic under the "pairing" group policy — under "allowlist" (this
integration's actual default) it stays empty for the whole process
lifetime unless something else populates it, and it's wiped on every
restart regardless (in-memory sync.Map, no persistence).
message.go's cross-target forward only attaches "group_id" metadata when
isGroupContext(ctx) is true, which reflects the ORIGIN session's peer
kind, not the destination target's. Forwarding from a DM (or any non-group
session) into a group therefore sends with no group_id metadata and an
empty approval cache, so Send() defaults to ThreadTypeUser — the message
goes out addressed as if to a user account, not the group, and is never
seen there. No error is returned anywhere in this path, so nothing in the
existing (or previously fixed) error-reporting surfaces it.
ListGroups now marks every returned group ID as approved via
MarkGroupApproved, and Channel.Start kicks off a best-effort ListGroups
call right after connecting so the cache is warm from process start, not
just after an explicit zalo_list_groups tool call.
* fix: expose a media tag for quote-forwarded images, not just vision access
extractQuoteMedia downloads the image attached to a quoted message so the
model can see it via vision, but the composed content only ever carried
the literal "[Quoted image]" placeholder text (from TQuote.Text()) — no
<media:image> tag was ever inserted for it. The later agent-side
enrichment (enrichImageIDs/enrichImagePaths) only fills id/path attributes
into a bare tag it finds already present in content; with no tag to find,
the downloaded file had no path reference the model could hand back to
message(MEDIA:<path>) to forward it elsewhere, even though the file was
sitting on disk the whole time. Symptom: asked to forward a quoted image
to another chat, the agent reported "the image is only a quote, no direct
file to attach" despite genuinely having already downloaded it.
buildQuoteMediaTag renders the same bare <media:image> tag a direct
(non-quoted) attachment gets, appended after the quote wrapper text in all
three call sites (handleDM, handleGroupMessage,
extractContentAndMediaWithQuote) — ordered after any of the current
message's own attachment tags to match the media slice's append order, so
positional id/path enrichment lines up correctly.
The provider verify call and the provider models-list call used hardcoded
timeouts (30s and 15s) that were too low for slow models. Introduce a
tenant-scoped, UI-editable setting `providers.request_timeout_sec`
(default 30) stored in system_configs, consumed by both handlers, and
overlaid from config.json like tts.timeout_ms. Adds a field to the System
Settings modal (all 4 locales) so operators can raise it without a
redeploy.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Eleven built-in tools were registered in the runtime tool registry (and
thus appeared in agent system prompts and were subject to allow/deny) but
were absent from builtinToolSeedData(), so operators could not see or
toggle them in the agent allow/deny UI. Seed datetime, heartbeat,
memory_expand, list_group_members, zalo_list_groups, vault_search,
vault_read, delegate, workstation_exec, claude_remote, and mcp_tool_search
so the UI catalog reflects the actual registered tool set. telegram_manager
is intentionally excluded per existing test guard.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
When a tool_call_prefix template has both a prefix and a suffix
(e.g. "pre_{tool_name}_suf"), StripToolPrefix stripped the prefix and
then sliced off the suffix length without checking the two ends don't
overlap. A model-emitted name that satisfies both HasPrefix and
HasSuffix but is shorter than prefix+suffix (e.g. "pre_suf") produced a
negative slice length -> "slice bounds out of range [:-1]" panic. This
runs in the agent turn loop on LLM-controlled names
(resolveToolCallName -> StripToolPrefix), so a crafted/degenerate name
crashes the turn.
Guard with len(prefix)+len(suffix) <= len(name); overlapping names now
fall through to returning the name unchanged, consistent with the other
no-match fallbacks. Adds a regression row to the prefix test table.
When a predefined agent has an operator-authored USER_PREDEFINED.md, the
built-in per-user USER.md template is no longer seeded or injected into the
system prompt. This gives operators full control over user-context in the
prompt and stops the per-turn name/timezone profile nag. Open agents and
predefined agents without USER_PREDEFINED.md are unaffected.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
The system prompt's Tooling section previously listed tools even when
they were disabled for a tenant and already stripped from the API
tools parameter, confusing the LLM into thinking it could use them.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
* fix: resolve Zalo group name to real chat ID, notify origin on failed forward
Two related root causes behind "forward message to group by name" silently
failing while the agent reports success:
1. No tool let the agent resolve a group's display name (e.g. "Ban Dieu
Hanh") to the chat ID the message tool actually requires. sessions_list
only exposes session keys (numeric group IDs), never human-readable
names, so the agent had no reliable way to turn a name into a real
target and ended up passing the display name itself as `target`.
Adds zalo_list_groups, wrapping the already-used (dashboard picker)
protocol.FetchGroups behind the same optional-interface pattern as
list_group_members/GroupMemberProvider (GroupListProvider on
channels.Manager, gated to zalo_personal via RequiredChannelTypes).
2. When the resulting send fails downstream (e.g. Zalo rejects a bad
chat_id), dispatchOutbound only ever retried/notified media failures on
the same (already-broken) destination, and dropped text-only failures
entirely — even though message.go's own postCrossTargetNotice comment
states forwards must never announce a fake delivery. Because the bus
publish is fire-and-forget, the tool had already returned "sent" and
announced success to the origin chat before the real send was even
attempted.
message.go now tags cross-target forwards with origin channel/chat in
OutboundMessage.Metadata; dispatchOutbound uses it to notify the ORIGIN
chat with the real failure instead of silently dropping it or retrying
against the same invalid target.
* test: cover forward-origin metadata tagging and dispatch failure notice
Extracts dispatchOutbound's error branch into handleSendFailure so it can
be unit tested without driving the consumer loop/goroutine, and adds
coverage for: forward failures notifying the origin chat (not the broken
destination), pre-existing non-forward media/text-only behavior staying
unchanged, message.go tagging cross-target group forwards with origin
metadata alongside group_id, and the new zalo_list_groups tool/Manager
delegator.
* feat(bitrix24): replace two MCP text inputs with a filtered dropdown [B24:2794]
Bitrix24 channel creation used to demand two hand-typed strings —
mcp_server_name and mcp_base_url — plus zero indication of which MCP
servers can actually auto-onboard. Typos silently disabled provisioning
and admin had to know which servers implement /api/auto-onboard.
Ship a single dropdown backed by mcp_servers.require_user_credentials,
plus the machinery to make it work end-to-end.
Phase 1 — DB & store
* Promote require_user_credentials from settings JSONB to a top-level
column on mcp_servers (PG migration 000089, SQLite migration v54).
* Backfill from existing settings blobs so no admin needs to re-tick.
* Add MCPServerData.RequireUserCredentials to the Go store layer, plumb
through Create / Get / GetByName / List / Update on both stores,
extend the export DTO, and add require_user_credentials to the HTTP
allowlist.
* Bump RequiredSchemaVersion 87 -> 89 (jumping 88, which was on disk
but not wired) and SchemaVersion 53 -> 54 with an
idempotentColumnMigration guard.
Phase 2 — Bitrix24 channel factory
* Add MCPServerID (UUID string) to bitrix24 InstanceConfig, keep
MCPServerName + MCPBaseURL as legacy fallback with a "deprecated"
doc comment.
* Factory validation accepts either mcp_server_id alone or the legacy
pair; half-config still fails fast.
* initMCPProvisioner prefers GetServer(id) and sources the base URL
from mcp_servers.url when the id path is used. Legacy name path
unchanged so pre-migration configs keep working.
* Log line now carries mcp_server_id + require_user_credentials so
operators can eyeball the wiring.
Phase 3 — Frontend types & MCP form
* Add optional top-level require_user_credentials to MCPServerData /
MCPServerInput in both ui/web and ui/desktop/frontend types.
* mcp-form-dialog reads the top-level flag first and falls back to
settings.require_user_credentials so cached responses from
pre-upgrade backends still render correctly.
* On submit send both the top-level flag AND the legacy settings
entry so mid-rollout backends stay consistent.
Phase 4 — Bitrix24 channel form dropdown
* New mcp-select field type + MCPServerSelect component. Uses the
shared useMCP() react-query cache and filters client-side to
servers whose require_user_credentials is true (OR settings
JSONB during the migration window).
* Explicit "None (disable MCP provisioning)" option so admins can
clear the binding without editing config JSON.
* Legacy mcp_server_name / mcp_base_url text inputs kept in the
Advanced panel, relabelled "(legacy)" with pointer help text.
Phase 5 — channel_instances.config backfill
* PG migration 000090 and SQLite migration v55 rewrite existing
bitrix24 channel_instances.config to add mcp_server_id by
resolving mcp_server_name against mcp_servers, tenant-scoped
via agents.tenant_id (channel_instances doesn't carry tenant_id
directly).
* Idempotent — only touches rows already carrying
mcp_server_name that lack mcp_server_id. Legacy keys are left
in place so provisioner.go can still fall back for unmigrated
or future-created legacy configs.
* down.sql drops the mcp_server_id key. Provisioner immediately
reverts to the legacy fallback path.
Tests
* provisioner_test.go: three new cases exercise the mcp_server_id
path (invalid UUID string, valid UUID with missing row, valid
UUID with a per-user row). fakeMCPStore gains a serversByID
map and a real GetServer implementation.
* Existing legacy-config tests unchanged and still green.
Verification
* go build ./... && go build -tags sqliteonly ./...
* go vet ./internal/mcp/... ./internal/channels/bitrix24/...
./internal/store/... ./internal/http/...
* go test ./internal/mcp/... ./internal/channels/bitrix24/... -> ok
* Live-tested against a local docker image on the goclaw-deploy
postgres. Migrations 89 + 90 applied cleanly. Three existing
bitrix24 channels (bitrix-sales / nguyen-dao-openline / tieu-vi)
had their configs backfilled with the b24-syn-mcp UUID and the
provisioner boots with require_user_credentials=true. UI dropdown
correctly shows only b24-syn-mcp (the only server with the flag
ticked) alongside a "None" clearer option.
Surface parity
* Gateway server: store + factory + provisioner + HTTP allowlist.
* API contract: adds require_user_credentials + mcp_server_id
as optional fields on existing routes. No new endpoints.
* Web UI: MCP form + Bitrix24 channel form + shared types.
* CLI/runtime package: N/A because no CLI subcommand reads the
mcp_server_id field.
* fix(bitrix24): derive auto-onboard base URL from mcp_servers.url origin [B24:2794]
The Phase 2 refactor swapped provisioner base-URL sourcing from the
legacy per-channel MCPBaseURL config field (which historically stored the
MCP server's ORIGIN, e.g. https://mcp.example.com) to mcp_servers.url,
which stores the JSON-RPC ENDPOINT the agent loop dials (e.g.
https://mcp.example.com/mcp). The two are semantically different but
share a single column.
mcp_client.newMCPClient then appends "/api/auto-onboard" to whatever
baseURL it receives, so the id-path started POSTing to
".../mcp/api/auto-onboard" — 404 for every per-user credential mint and
refresh. Existing users kept working only until their cached access
tokens expired.
Fix: derive the origin (scheme://host[:port]) from server.URL before
handing it to the auto-onboard client. The legacy path is untouched
because MCPBaseURL from channel config is already the origin.
https://b24-mcp-dev.synity.so/mcp -> https://b24-mcp-dev.synity.sohttps://mcp.example.com/mcp/ -> https://mcp.example.comhttps://mcp.example.com -> https://mcp.example.com
Table-driven test covers six shapes plus four error cases (empty,
whitespace-only, no scheme, no host). Updated the existing
TestInitMCPProvisioner_MCPServerID fixture to seed a URL with the /mcp
subpath so it regression-guards the same code path.
Verified live: user 614 sent a message that triggered the expired-cred
refresh branch; goclaw logged "self-refreshed user credentials
created=false" and the agent immediately reported
"mcp.user_tools_loaded user=614 tools=2". Before this fix the same event
logged 'auto-onboard failed: mcp auto-onboard: 404 Not Found'.
Surface parity:
- Gateway server: provisioner + one new helper (deriveAutoOnboardBaseURL).
- API contract: N/A because the wire shape hasn't changed.
- Web UI: N/A because the UI still writes mcp_server_id verbatim.
- CLI/runtime: N/A because no CLI reads the derived base URL.
---------
Co-authored-by: DangTinh311 <dangtinh31193@gmail.com>
Go's regexp (RE2) treats \b as an ASCII-only word boundary, so the
privilege_escalation deny pattern \bsu\b matched the ASCII "su" inside
non-ASCII words such as the Vietnamese "suất" — the accented 'ấ' is not
an ASCII word char, so \b sees a spurious word boundary right after "su".
Any exec command containing such a word (for example a Vietnamese
document embedded in a python heredoc to render a PDF) was wrongly
blocked as a privilege-escalation attempt. Replace the pattern with a
Unicode-aware boundary (\pL, \pN) that still matches the real su command.
slugify's docstring promises "no leading/trailing dashes" and the sibling
isValidSlug rejects a trailing dash, but slugify only stripped leading
dashes. Any input ending in a non-[a-z0-9] character (e.g. "hello world!",
a trailing space, "trailing---") produced a slug ending in "-" that then
failed isValidSlug, causing invalid persisted keys and spurious validation
errors on agent/channel/provider/cron/mcp name fields. Add a symmetric
trailing-dash strip, matching the Go helpers (strings.Trim(s, "-")).
The tenants.* methods were the only RPC group without MethodXxx
constants in pkg/protocol/methods.go. Their wire names were hardcoded
as raw string literals across three call sites — router registration,
and the read/admin allowlists in the permission policy — which is
inconsistent with the other 147 methods and risks silent drift: a
typo in any allowlist string is not caught by the compiler and would
fail-closed (deny) at runtime, which is hard to diagnose.
Add MethodTenants{List,Get,Create,Update,UsersList,UsersAdd,
UsersRemove,Mine} and replace the method-identifier literals in
internal/permissions/policy.go and internal/gateway/methods/tenants.go
with the constants.
No behavior change — string values are identical. Human-facing error
labels and log prefixes in internal/http/tenants.go are intentionally
left as free-form text, matching the sibling HTTP handler convention.
Python skill scripts using zoneinfo.ZoneInfo (and Go time.LoadLocation with
named zones) fail on the published images because alpine:3.23 ships without
tzdata: ZoneInfoNotFoundError: 'No time zone found with key Europe/Paris'.
Add tzdata to the base runtime apk install so every image variant (base,
latest, full, claude-cli) has the system zoneinfo database.
Current Claude CLI releases (>= 2.x) no longer register TodoRead and
NotebookRead. Passing them via --disallowedTools makes the CLI print
'Permission deny rule "..." matches no known tool' on every invocation,
including background episodic summarization, burying real errors.
Keep TodoWrite/NotebookEdit blocked (still registered, no GoClaw
equivalent). Add a regression test guarding both directions.
Fixes#1374
The bridge exposed a static BridgeToolNames subset that drifted from the
tool registry: use_skill, datetime, knowledge_graph_search and skill_manage
were never added, while the system prompt's skill-loading protocol requires
agents to call use_skill. claude_cli agents following the protocol hit a
nonexistent tool and could fabricate results.
Implement the structural fix recommended in #1373 triage:
- register the full bridge-capable surface (registry minus hard exclusions
spawn/create_forum_topic) instead of the static list
- gate BOTH tools/list (new WithToolFilter) and tools/call through one
shared predicate bridgeToolAllowed:
* callers WITH a verified agent policy get exactly the policy-filtered
surface (same WouldAllow check the call path always enforced)
* callers WITHOUT one (anonymous, or agent without tools_config) keep the
legacy conservative BridgeToolNames set - no exposure widening
- downgrade the per-call denial log Warn->Info; list filtering makes probes
of denied tools rare and the call gate is the intended enforcement point
Fixes#1373
Cron job execution set the tenant ID on its context but not the tenant
slug. Tenant-scoped filesystem paths (skills-store, workspace, media via
config.TenantScopedDir) key off the slug and fall back to an id-based
path when it is absent — a different directory than where HTTP/WS skill
upload materialized the files (which sets the slug). As a result a cron
agent turn in a non-master tenant saw NONE of its tenant's managed
skills: skill_search returned 0 results and the agent, unable to run the
skill, produced an ungrounded answer.
Add cronTenantContext() which resolves the tenant slug via TenantStore
and sets both WithTenantID and WithTenantSlug. Master tenant and
nil-store/lookup-failure paths fall back to id-only (prior behavior).
Thread TenantStore into makeCronJobHandler and runCommandCronJob.
Tested: added unit tests for cronTenantContext (slug injected for
non-master; master skips lookup; nil store and lookup error fall back to
id-only). Verified end-to-end on a live tenant: before, a daily-agenda
cron guessed an empty day; after, it read the real event from the DB.
Note: other background executors that build a context from a tenant ID
(e.g. heartbeat) likely share this gap and are worth an audit.
* fix: unify NO_REPLY detection (#1233)
Co-authored-by: GoClaw Operator <operator@goclaw>
* feat(acp): surface session/update tool_call notifications at Info level (#1141)
session/update notifications carrying a ToolCall (or inline tool_call /
tool_call_update kind) were only dumped at Debug level inside the params
blob. Operators could see `security.tool_granted` (permission granted) but
had no way to tell whether the tool actually executed successfully — both
"permission granted then failed silently" and "permission granted and
succeeded" looked identical in journalctl.
Add a structured Info log emitting toolCallId, name/title, status, and a
content preview (truncated at 400 chars) whenever the notification
contains tool-call state. This is what made it possible to diagnose the
recent .goclaw/-path-deny regression — `status=failed` immediately after
`security.tool_granted` revealed the gap that the granted-only log hid.
No behavior change beyond logging volume; preview is bounded.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(exec): exempt venv python interpreter from .goclaw/ path deny (#1140)
The ExecTool path-deny rule blocks any token containing `.goclaw/` unless
it matches one of the AllowPathExemptions prefixes (skills-store, tenants).
This silently rejected legitimate commands invoking the goclaw-managed
Python interpreter via its absolute path:
/home/user/.goclaw/venv/bin/python3 .../script.py
The first token `/home/user/.goclaw/venv/bin/python3` matched the deny
pattern but no exemption, so the entire command was denied.
Naive exemption (".goclaw/venv/bin/") does not work: matchesAnyPathExemption
resolves both tokens and exemption candidates via EvalSymlinks, and the
venv's python3 is a symlink into the host's python cellar (e.g. linuxbrew).
The token canonicalizes to /home/linuxbrew/.../python3.14 while a literal
".goclaw/venv/bin/" prefix never gets touched.
Fix: resolve venv/bin/python3 once at startup and exempt the dirname of
the resolved target. Failure to resolve (no venv present) silently falls
through.
Without this, ACP-driven agents either fail outright or work only via
fragile heuristics (cwd-local symlinks generated on the fly by the LLM).
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(cron): invert stateless gate to match UI label (#1139)
The cron job handler reset the session when stateless=false and skipped
reset when stateless=true — the opposite of what every UI locale labels
the field (en/ko/zh/vi all describe stateless as "each run starts fresh
without loading previous messages").
The buggy gate caused stateless=true crons to silently accumulate session
history across every execution, leading to context bloat and increasing
the chance of LLMs short-circuiting tool calls in favor of replaying
prior assistant turns. One affected daily ETL cron grew to 38 messages
over 18 days before the regression was noticed.
Fix: gate the Reset on `if job.Stateless` so the runtime matches the UI
contract. No DB migration is required — existing values were set by users
based on the UI label, so they already encode the intended behavior.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(providers): retry transient Codex response failures (#1332)
* fix(providers): retry transient Codex response failures
* test(cron): tolerate nil session store in handler tests
* fix(background): clear stale provider alerts (#1362)
Co-authored-by: Collective Developer <man@collective.dev>
---------
Co-authored-by: Duy /zuey/ <duy@wearetopgroup.com>
Co-authored-by: nguyenha935 <nguyenthanhha935@gmail.com>
Co-authored-by: GoClaw Operator <operator@goclaw>
Co-authored-by: codebit0 <34156842+codebit0@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Zezae Oh <zezaeoh@gmail.com>
Co-authored-by: Collective Developer <man@collective.dev>
* fix(mcp): return bare tool names with descriptions for already-connected MCP servers, fixing empty descriptions/schemas in agent prompts
* fix(mcp): resolve live tool descriptions/schemas in ListToolsForAgent for legacy prefixed tool_allow entries
The previous fix (d612da8e) only corrected the admin "browse/select tools"
endpoint. The actual runtime path building the LLM system prompt's tool list,
Manager.ListToolsForAgent, still looked up tool_cache/registry entries keyed
by the raw stored tool_allow name. Agents with grants captured before that
fix have tool_allow persisted as registered (prefixed) names instead of bare
MCP tool names, so every lookup missed and every tool in the prompt preview
came back with an empty description and schema.
Add bareMCPToolName to normalize both legacy-prefixed and current bare
tool_allow/tool_deny entries before matching, and resolveMCPToolInfo to
prefer the live registry's current description/schema (falling back to the
settings tool_cache) instead of trusting whatever was persisted verbatim.
Both the explicit-allow and cache-enumeration branches now share this single
resolution path.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Adds an operator opt-in env var (comma-separated hostnames/IPs) to
whitelist additional local hosts for ollama/acp provider-type URL
validation, on top of the existing hardcoded localhost allowlist.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
* feat(skills): allow editing SKILL.md content for database-managed skills
Adds the ability to edit a skill's actual SKILL.md content directly
from the web UI, restricted to database/tenant-managed skills only --
bundled/system skills remain read-only.
Backend: new PUT /v1/skills/{id}/files/{path...} endpoint
(handleWriteFile in skills_versions.go), writing to the current
version directory of a non-system skill only (403 for is_system
skills), with the same path-traversal/symlink/ownership guards as the
existing read endpoint. Bumps skill store version and emits
cache-invalidate + audit event on success.
Frontend: Content tab in the skill detail dialog gets an "Edit
content" button, disabled with a tooltip for system skills. Editing
fetches the RAW file content via the existing read endpoint (which
does not strip YAML frontmatter, unlike the stripped preview shown in
the Content tab by default) to avoid silently destroying frontmatter
on save. Save/cancel follow the established catch-and-toast error
pattern (no uncaught promise rejections, matching PR #1346's config
save button fix).
i18n keys added to all 4 locales (en, ko, vi, zh).
* fix(skills): bump version on content edit, fix scroll regression in content view
handleWriteFile previously overwrote the current skill version in place
instead of creating a new immutable version, inconsistent with the
skill_manage tool's patch action. Now creates a new version directory
and updates the DB pointer, matching that convention.
Also fixes a scroll regression in the skill detail dialog's Content
tab introduced by the edit-mode textarea -- both the read-only preview
and edit textarea are now properly scrollable within the dialog.
* fix(skills): refresh version display after save, fix dialog height constraint for scrolling
Version bump now correctly reflects in the UI immediately after save
without requiring a page reload. Fixed the actual root cause of the
scroll issue: the dialog's height wasn't bounded, so overflow-y-auto
on inner content had no effect since nothing constrained the dialog's
total height in the first place.
---------
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Browser pairing authenticates users as RoleOperator, but every provider
and OAuth route was wrapped in requireAuth(RoleAdmin). The /setup flow reads
GET /v1/providers, so paired browsers got 403 and hung on the setup page.
Split provider/OAuth route auth by operation:
- providers.go: GET list/get/models/status/embedding/codex-activity now use a
readAuth wrapper (method-derived min → Viewer); POST/PUT/DELETE and
reconnect/verify stay Admin. Responses already mask API keys and queries
stay tenant-scoped, so no secrets are exposed.
- oauth.go: GET status/quota use readAuth (they return only connection state
and quota, never tokens); start/callback/logout stay Admin (preserving the
#450 admin-only-management decision for mutations).
Updates the two auth tests to assert the new contract: an Operator can read
providers/OAuth status (200) but still cannot mutate them (403).
`goclaw migrate up` (and every other migrate subcommand) failed on Windows
with `create migrator: failed to open source, "file:///D:/.../migrations":
open .: The filename, directory name, or volume label syntax is incorrect.`
golang-migrate's file source driver mis-parses absolute drive-letter file://
URLs; the drive-letter formatting in absoluteToFileURI produced a URL the
driver could not open. Replace the file:// URL with an iofs source over
os.DirFS, which uses native OS path handling and works identically on every
platform. Drop the now-unused absoluteToFileURI/migrationsSourceURL helpers
and their URL-shape tests, and add a DB-free regression test that opens the
real migrations directory via newMigrationSource().
Verified end-to-end on Windows: `migrate version` and `migrate up` now run.
Closes#1076, #1338.
Fix A — tenant backup aborted with SQLSTATE 42703 "column id does not
exist" because exportQuery() hardcoded ORDER BY id. Several tenant-scoped
tables have composite PKs and no id column. Add a TableDef.OrderBy field,
honor it in exportQuery(), and set it for every id-less registry table:
config_secrets, agent_team_members, tenant_hook_budget, system_configs,
builtin_tool_tenant_configs, skill_tenant_configs, user_agent_profiles.
Fix B — hooks, tenant_hook_budget and webhook config were missing from
the backup registry, silently dropping their data on backup/restore. Add
hooks, tenant_hook_budget, webhooks (preserve) plus the hook_agents junction
(via ParentJoin through hooks, composite PK), and mark hook_executions and
webhook_calls as ephemeral in the skipped list.
Fix C — restoring on a fresh server failed with "N active DB connection(s)
detected" because the gateway's own pool connections were counted as active
clients. Tag pool connections with application_name='goclaw' (pg.OpenDB) and
exclude them in CheckActiveConnections; genuine external clients still block.
Adds unit + integration regression tests, including an export-over-every-
registered-table test that surfaced the additional id-less tables.