Commit Graph
1946 Commits
Author SHA1 Message Date
thotam e95851e690 fix(pipeline): aggregate cache tokens across agent turns (#1419)
ThinkStage folded each turn's usage into state.Think.TotalUsage but summed
only prompt/completion/total/thinking tokens, dropping CacheReadTokens,
CacheCreationTokens and the PromptTokensIncludeCachedSegments flag. The
aggregated RunResult.Usage is the source for webhook usage responses and
usage_events analytics, so async multi-turn agent runs reported zero cache
tokens even when traces showed heavy prompt-cache use (cache is read
per-response for spans, but never summed into the aggregate).

Sum cache read/creation across turns and OR the include-cached flag (a
provider-level property, consistent across a run). Restores correct cache
accounting for webhook usage and usage_events without any envelope change.
2026-07-10 20:25:06 +07:00
b459d40531 feat(sandbox): expose Workdir field in SandboxConfig (#853)
The internal sandbox.Config already supports a custom container Workdir
that doubles as the workspace mount target, but config.SandboxConfig
(the JSONB-backed per-agent config) did not expose it. Without this,
agents must use the hardcoded /workspace mount point even when their
scripts expect a different in-container path.

Adds an optional "workdir" JSON field that maps through to
sandbox.Config.Workdir. Backward compatible — existing agents without
workdir set fall back to the /workspace default.

Co-authored-by: tuannt23065 <tuannt23065@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-07-10 15:56:31 +07:00
SYNITY 641d4f8756 feat(bitrix24): ship per-user OAuth re-authorization flow (#1417)
feat(bitrix24): ship per-user OAuth re-authorization flow (#1417)
2026-07-10 10:56:45 +07:00
Zezae Oh 9be468edad feat(providers): add GPT-5.6 Codex models (#1416) 2026-07-10 06:55:30 +07:00
db69ccf882 tracing: nest tool calls under LLM spans, emit hook spans, fix usage_events FK (#1415)
- Reparent tool_call spans under their producing llm_call span via
  RunState.CurrentLLMSpanID so a model turn's trace visibly contains the
  tool calls it triggered.
- Wire internal/hooks EmitHookSpan into the dispatcher writeExec so every
  hook execution (pre/post tool use, etc.) produces a trace span with input,
  console output, decision, error, and duration for agent troubleshooting.
- Fix usage_events_span_id_fkey violations: usage events were inserted
  synchronously referencing a span flushed ~5s later, silently dropping
  token/cost data. Route them through the collector so they flush after
  spans in the same cycle. No schema migration; FK preserved.

Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-07-10 03:54:42 +07:00
ナムandCollective Developer d84576939d fix(kg): count newly inserted dedup candidates (#1413)
Co-authored-by: Collective Developer <man@collective.dev>
2026-07-10 01:54:25 +07:00
ナムandCollective Developer 84fab59226 fix(memory): guide recency recall tool calls (#1412)
* fix(memory): guide recency recall tool calls

* fix(tools): align memory and vault builtin visibility

---------

Co-authored-by: Collective Developer <man@collective.dev>
2026-07-10 01:24:39 +07:00
ナムandCollective Developer e2af8b5f15 feat(memory): add episodic time filters to search (#1411)
Co-authored-by: Collective Developer <man@collective.dev>
2026-07-10 00:55:13 +07:00
ナムandCollective Developer 97e316f460 fix(memory): surface passive channel episodic recall (#1410)
* fix(memory): surface passive channel episodic recall

* fix(memory): expose episodic key topics

---------

Co-authored-by: Collective Developer <man@collective.dev>
2026-07-09 23:55:11 +07:00
ナムandCollective Developer 12d5619cdc fix(graph): preserve knowledge graph node colors (#1409)
Co-authored-by: Collective Developer <man@collective.dev>
2026-07-09 23:24:42 +07:00
ナムandCollective Developer e9898be505 fix(discord): honor disabled group history limit (#1406)
Co-authored-by: Collective Developer <man@collective.dev>
2026-07-09 17:25:37 +07:00
2b7bfec504 feat(telegram): agent-callable message reaction (message action=react) (#1407)
Adds a "react" action to the message tool so an agent can set an emoji
reaction on an existing message (e.g. mark its own status post 👍 once
fully paid). Mirrors the ChannelEditor wiring: ReactionSetter capability
-> Manager.ReactToMessage -> telegram Channel.ReactToMessage (validates
against Telegram's supported reaction set; ✅/❌ are rejected).

Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-09 16:55:54 +07:00
ナムandCollective Developer fb28b72b27 fix(config): remove ignored session scoping controls (#1405)
Co-authored-by: Collective Developer <man@collective.dev>
2026-07-09 16:24:58 +07:00
ba621391c2 feat(telegram): trigger-words, channel posts, message edit & topic posting (#1383)
- Fix image_generation nil-pointer crash on codex agents: the sentinel now
  carries a name-only Function so the many sites reading td.Function.Name never
  nil-deref; codex_build still branches on Type.
- Agent-declared trigger words in IDENTITY.md wake the bot in groups without an
  @mention (whole-word, Cyrillic-aware; text + caption), cached per-agent 60s.
- channel_post support with a synthetic sender + a recover() guard so a
  malformed update can't crash the gateway.
- message tool action=edit (editMessageText + editMessageCaption fallback),
  targeting the replied-to message via reply_to_message_id.
- message tool topic=<name> posts into a named forum topic; topics learned from
  forum_topic_created into channel_contacts and resolved by name.

Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-09 12:28:18 +07:00
SYNITYandDangTinh311 51ed7845e2 fix(bitrix24): register newly-created portal in live router [B24:2794] (#1404)
bitrix.portals.create only persisted the DB row — it never called
Router.RegisterPortal, so a portal created while the gateway was
already running stayed invisible to PortalByDomain/PortalByKey until
the next full restart. The only other RegisterPortal call site is
BootstrapPortals, which runs once at boot. Result: the OAuth install
callback 404'd with "unknown portal" for any portal created through
the self-service UI mid-uptime — hit live on web1trang.bitrix24.com.

Fix: inject registerPortal/unregisterPortal func fields into
BitrixPortalsMethods (mirrors the existing gatewayPublicURL func()
pattern), defaulting to bitrix24.WebhookRouter() + NewPortal/
RegisterPortal (same construction path as BootstrapPortals) and
Router.UnregisterPortal respectively. handleCreate calls
registerPortal after a successful Create; handleDelete calls
unregisterPortal after a successful Delete for symmetry. Both are
best-effort — a registration failure is logged, not returned as an
RPC error, since the DB row is already valid either way.

Injected via func fields (not calling bitrix24.WebhookRouter()
inline) so tests can stub live-router registration without touching
that process-wide singleton, whose test-reset helper isn't exported
outside the bitrix24 package.

Tests: 4 new cases covering register-on-create, best-effort failure
handling, unregister-on-delete, and no-unregister-when-delete-blocked.
All 20 tests in bitrix_portals_test.go pass; go build (both PG and
sqliteonly tags) and go vet clean.

Co-authored-by: DangTinh311 <dangtinh31193@gmail.com>
2026-07-09 10:54:56 +07:00
SYNITYandDangTinh311 47068ad563 feat: add delete affordance for pending portals in channel creation form [B24:2794] (#1403)
- Add delete icon (X) to portal dropdown items for pending (not-yet-installed) Bitrix24 portals
- Support mouse click and keyboard delete key (Del/Backspace) for removing pending portals
- Implement portal removal handler in bitrix-portal-select component
- Add comprehensive tests for delete functionality
- Update i18n strings (en/vi/zh) with delete action labels

Co-authored-by: DangTinh311 <dangtinh31193@gmail.com>
2026-07-09 10:24:43 +07:00
Duc Nguyen ee37017e0e fix: warm Zalo group-approval cache + attach quoted-image content tag (#1402)
* fix: warm Zalo group-approval cache so forwards from non-group origins route correctly

send.go picks ThreadTypeGroup vs ThreadTypeUser for a target chat ID by
checking BaseChannel's approvedGroups cache (or an explicit "group_id"
metadata key). That cache is only ever populated from INBOUND group
traffic under the "pairing" group policy — under "allowlist" (this
integration's actual default) it stays empty for the whole process
lifetime unless something else populates it, and it's wiped on every
restart regardless (in-memory sync.Map, no persistence).

message.go's cross-target forward only attaches "group_id" metadata when
isGroupContext(ctx) is true, which reflects the ORIGIN session's peer
kind, not the destination target's. Forwarding from a DM (or any non-group
session) into a group therefore sends with no group_id metadata and an
empty approval cache, so Send() defaults to ThreadTypeUser — the message
goes out addressed as if to a user account, not the group, and is never
seen there. No error is returned anywhere in this path, so nothing in the
existing (or previously fixed) error-reporting surfaces it.

ListGroups now marks every returned group ID as approved via
MarkGroupApproved, and Channel.Start kicks off a best-effort ListGroups
call right after connecting so the cache is warm from process start, not
just after an explicit zalo_list_groups tool call.

* fix: expose a media tag for quote-forwarded images, not just vision access

extractQuoteMedia downloads the image attached to a quoted message so the
model can see it via vision, but the composed content only ever carried
the literal "[Quoted image]" placeholder text (from TQuote.Text()) — no
<media:image> tag was ever inserted for it. The later agent-side
enrichment (enrichImageIDs/enrichImagePaths) only fills id/path attributes
into a bare tag it finds already present in content; with no tag to find,
the downloaded file had no path reference the model could hand back to
message(MEDIA:<path>) to forward it elsewhere, even though the file was
sitting on disk the whole time. Symptom: asked to forward a quoted image
to another chat, the agent reported "the image is only a quote, no direct
file to attach" despite genuinely having already downloaded it.

buildQuoteMediaTag renders the same bare <media:image> tag a direct
(non-quoted) attachment gets, appended after the quote wrapper text in all
three call sites (handleDM, handleGroupMessage,
extractContentAndMediaWithQuote) — ordered after any of the current
message's own attachment tags to match the media slice's append order, so
positional id/path enrichment lines up correctly.
2026-07-09 07:55:02 +07:00
Bruno ClermontandBruno Clermont d856a58218 feat(providers): make provider request timeout configurable (default 30s) (#1400)
The provider verify call and the provider models-list call used hardcoded
timeouts (30s and 15s) that were too low for slow models. Introduce a
tenant-scoped, UI-editable setting `providers.request_timeout_sec`
(default 30) stored in system_configs, consumed by both handlers, and
overlaid from config.json like tts.timeout_ms. Adds a field to the System
Settings modal (all 4 locales) so operators can raise it without a
redeploy.

Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
2026-07-09 04:54:15 +07:00
Bruno ClermontandBruno Clermont 951acad714 feat(tools): seed missing built-in tools into builtin_tools catalog (#1399)
Eleven built-in tools were registered in the runtime tool registry (and
thus appeared in agent system prompts and were subject to allow/deny) but
were absent from builtinToolSeedData(), so operators could not see or
toggle them in the agent allow/deny UI. Seed datetime, heartbeat,
memory_expand, list_group_members, zalo_list_groups, vault_search,
vault_read, delegate, workstation_exec, claude_remote, and mcp_tool_search
so the UI catalog reflects the actual registered tool set. telegram_manager
is intentionally excluded per existing test guard.

Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
2026-07-09 04:24:19 +07:00
Syed Osama Ali Shah f9b07d0a9c fix(tools): avoid panic in StripToolPrefix on overlapping prefix+suffix (#1398)
When a tool_call_prefix template has both a prefix and a suffix
(e.g. "pre_{tool_name}_suf"), StripToolPrefix stripped the prefix and
then sliced off the suffix length without checking the two ends don't
overlap. A model-emitted name that satisfies both HasPrefix and
HasSuffix but is shorter than prefix+suffix (e.g. "pre_suf") produced a
negative slice length -> "slice bounds out of range [:-1]" panic. This
runs in the agent turn loop on LLM-controlled names
(resolveToolCallName -> StripToolPrefix), so a crafted/degenerate name
crashes the turn.

Guard with len(prefix)+len(suffix) <= len(name); overlapping names now
fall through to returning the name unchanged, consistent with the other
no-match fallbacks. Adds a regression row to the prefix test table.
2026-07-09 03:54:33 +07:00
Bruno ClermontandBruno Clermont d3c04ab3e5 feat(bootstrap): skip built-in USER.md for predefined agents with USER_PREDEFINED.md (#1397)
When a predefined agent has an operator-authored USER_PREDEFINED.md, the
built-in per-user USER.md template is no longer seeded or injected into the
system prompt. This gives operators full control over user-context in the
prompt and stops the per-turn name/timezone profile nag. Open agents and
predefined agents without USER_PREDEFINED.md are unaffected.

Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
2026-07-09 03:24:15 +07:00
Bruno ClermontandBruno Clermont 6bafe54b08 fix(agent): filter per-tenant disabled tools from system prompt (#1396)
The system prompt's Tooling section previously listed tools even when
they were disabled for a tenant and already stripped from the API
tools parameter, confusing the LLM into thinking it could use them.

Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
2026-07-09 02:23:46 +07:00
Duc Nguyen aa15317d5b fix: resolve Zalo group by name + notify origin on failed forward (#1395)
* fix: resolve Zalo group name to real chat ID, notify origin on failed forward

Two related root causes behind "forward message to group by name" silently
failing while the agent reports success:

1. No tool let the agent resolve a group's display name (e.g. "Ban Dieu
   Hanh") to the chat ID the message tool actually requires. sessions_list
   only exposes session keys (numeric group IDs), never human-readable
   names, so the agent had no reliable way to turn a name into a real
   target and ended up passing the display name itself as `target`.

   Adds zalo_list_groups, wrapping the already-used (dashboard picker)
   protocol.FetchGroups behind the same optional-interface pattern as
   list_group_members/GroupMemberProvider (GroupListProvider on
   channels.Manager, gated to zalo_personal via RequiredChannelTypes).

2. When the resulting send fails downstream (e.g. Zalo rejects a bad
   chat_id), dispatchOutbound only ever retried/notified media failures on
   the same (already-broken) destination, and dropped text-only failures
   entirely — even though message.go's own postCrossTargetNotice comment
   states forwards must never announce a fake delivery. Because the bus
   publish is fire-and-forget, the tool had already returned "sent" and
   announced success to the origin chat before the real send was even
   attempted.

   message.go now tags cross-target forwards with origin channel/chat in
   OutboundMessage.Metadata; dispatchOutbound uses it to notify the ORIGIN
   chat with the real failure instead of silently dropping it or retrying
   against the same invalid target.

* test: cover forward-origin metadata tagging and dispatch failure notice

Extracts dispatchOutbound's error branch into handleSendFailure so it can
be unit tested without driving the consumer loop/goroutine, and adds
coverage for: forward failures notifying the origin chat (not the broken
destination), pre-existing non-forward media/text-only behavior staying
unchanged, message.go tagging cross-target group forwards with origin
metadata alongside group_id, and the new zalo_list_groups tool/Manager
delegator.
2026-07-08 23:23:08 +07:00
ナムandCollective Developer fc7d183961 Enhance passive memory extraction tuning (#1393)
* feat(channelmemory): add passive extraction tuning

* feat(ui): expose passive memory tuning controls

* docs(memory): document passive extraction tuning

---------

Co-authored-by: Collective Developer <man@collective.dev>
2026-07-08 20:53:19 +07:00
SYNITYandDangTinh311 3b8bc1bc4d feat(bitrix24): replace two MCP text inputs with a filtered dropdown (#1392)
* feat(bitrix24): replace two MCP text inputs with a filtered dropdown [B24:2794]

Bitrix24 channel creation used to demand two hand-typed strings —
mcp_server_name and mcp_base_url — plus zero indication of which MCP
servers can actually auto-onboard. Typos silently disabled provisioning
and admin had to know which servers implement /api/auto-onboard.

Ship a single dropdown backed by mcp_servers.require_user_credentials,
plus the machinery to make it work end-to-end.

Phase 1 — DB & store
* Promote require_user_credentials from settings JSONB to a top-level
  column on mcp_servers (PG migration 000089, SQLite migration v54).
* Backfill from existing settings blobs so no admin needs to re-tick.
* Add MCPServerData.RequireUserCredentials to the Go store layer, plumb
  through Create / Get / GetByName / List / Update on both stores,
  extend the export DTO, and add require_user_credentials to the HTTP
  allowlist.
* Bump RequiredSchemaVersion 87 -> 89 (jumping 88, which was on disk
  but not wired) and SchemaVersion 53 -> 54 with an
  idempotentColumnMigration guard.

Phase 2 — Bitrix24 channel factory
* Add MCPServerID (UUID string) to bitrix24 InstanceConfig, keep
  MCPServerName + MCPBaseURL as legacy fallback with a "deprecated"
  doc comment.
* Factory validation accepts either mcp_server_id alone or the legacy
  pair; half-config still fails fast.
* initMCPProvisioner prefers GetServer(id) and sources the base URL
  from mcp_servers.url when the id path is used. Legacy name path
  unchanged so pre-migration configs keep working.
* Log line now carries mcp_server_id + require_user_credentials so
  operators can eyeball the wiring.

Phase 3 — Frontend types & MCP form
* Add optional top-level require_user_credentials to MCPServerData /
  MCPServerInput in both ui/web and ui/desktop/frontend types.
* mcp-form-dialog reads the top-level flag first and falls back to
  settings.require_user_credentials so cached responses from
  pre-upgrade backends still render correctly.
* On submit send both the top-level flag AND the legacy settings
  entry so mid-rollout backends stay consistent.

Phase 4 — Bitrix24 channel form dropdown
* New mcp-select field type + MCPServerSelect component. Uses the
  shared useMCP() react-query cache and filters client-side to
  servers whose require_user_credentials is true (OR settings
  JSONB during the migration window).
* Explicit "None (disable MCP provisioning)" option so admins can
  clear the binding without editing config JSON.
* Legacy mcp_server_name / mcp_base_url text inputs kept in the
  Advanced panel, relabelled "(legacy)" with pointer help text.

Phase 5 — channel_instances.config backfill
* PG migration 000090 and SQLite migration v55 rewrite existing
  bitrix24 channel_instances.config to add mcp_server_id by
  resolving mcp_server_name against mcp_servers, tenant-scoped
  via agents.tenant_id (channel_instances doesn't carry tenant_id
  directly).
* Idempotent — only touches rows already carrying
  mcp_server_name that lack mcp_server_id. Legacy keys are left
  in place so provisioner.go can still fall back for unmigrated
  or future-created legacy configs.
* down.sql drops the mcp_server_id key. Provisioner immediately
  reverts to the legacy fallback path.

Tests
* provisioner_test.go: three new cases exercise the mcp_server_id
  path (invalid UUID string, valid UUID with missing row, valid
  UUID with a per-user row). fakeMCPStore gains a serversByID
  map and a real GetServer implementation.
* Existing legacy-config tests unchanged and still green.

Verification
* go build ./... && go build -tags sqliteonly ./...
* go vet ./internal/mcp/... ./internal/channels/bitrix24/...
    ./internal/store/... ./internal/http/...
* go test ./internal/mcp/... ./internal/channels/bitrix24/...  -> ok
* Live-tested against a local docker image on the goclaw-deploy
  postgres. Migrations 89 + 90 applied cleanly. Three existing
  bitrix24 channels (bitrix-sales / nguyen-dao-openline / tieu-vi)
  had their configs backfilled with the b24-syn-mcp UUID and the
  provisioner boots with require_user_credentials=true. UI dropdown
  correctly shows only b24-syn-mcp (the only server with the flag
  ticked) alongside a "None" clearer option.

Surface parity
* Gateway server: store + factory + provisioner + HTTP allowlist.
* API contract: adds require_user_credentials + mcp_server_id
  as optional fields on existing routes. No new endpoints.
* Web UI: MCP form + Bitrix24 channel form + shared types.
* CLI/runtime package: N/A because no CLI subcommand reads the
  mcp_server_id field.

* fix(bitrix24): derive auto-onboard base URL from mcp_servers.url origin [B24:2794]

The Phase 2 refactor swapped provisioner base-URL sourcing from the
legacy per-channel MCPBaseURL config field (which historically stored the
MCP server's ORIGIN, e.g. https://mcp.example.com) to mcp_servers.url,
which stores the JSON-RPC ENDPOINT the agent loop dials (e.g.
https://mcp.example.com/mcp). The two are semantically different but
share a single column.

mcp_client.newMCPClient then appends "/api/auto-onboard" to whatever
baseURL it receives, so the id-path started POSTing to
".../mcp/api/auto-onboard" — 404 for every per-user credential mint and
refresh. Existing users kept working only until their cached access
tokens expired.

Fix: derive the origin (scheme://host[:port]) from server.URL before
handing it to the auto-onboard client. The legacy path is untouched
because MCPBaseURL from channel config is already the origin.

  https://b24-mcp-dev.synity.so/mcp   ->  https://b24-mcp-dev.synity.so
  https://mcp.example.com/mcp/        ->  https://mcp.example.com
  https://mcp.example.com             ->  https://mcp.example.com

Table-driven test covers six shapes plus four error cases (empty,
whitespace-only, no scheme, no host). Updated the existing
TestInitMCPProvisioner_MCPServerID fixture to seed a URL with the /mcp
subpath so it regression-guards the same code path.

Verified live: user 614 sent a message that triggered the expired-cred
refresh branch; goclaw logged "self-refreshed user credentials
created=false" and the agent immediately reported
"mcp.user_tools_loaded user=614 tools=2". Before this fix the same event
logged 'auto-onboard failed: mcp auto-onboard: 404 Not Found'.

Surface parity:
  - Gateway server: provisioner + one new helper (deriveAutoOnboardBaseURL).
  - API contract: N/A because the wire shape hasn't changed.
  - Web UI: N/A because the UI still writes mcp_server_id verbatim.
  - CLI/runtime: N/A because no CLI reads the derived base URL.

---------

Co-authored-by: DangTinh311 <dangtinh31193@gmail.com>
2026-07-08 19:23:21 +07:00
ナムandCollective Developer 7263f771c8 fix(channelmemory): honor Discord parent channel excludes (#1390)
* fix(channelmemory): harden passive extraction

* fix(channelmemory): honor discord parent channel excludes

---------

Co-authored-by: Collective Developer <man@collective.dev>
2026-07-08 13:23:16 +07:00
Duc Nguyen 8367a9bd0c fix(exec): stop privilege-escalation su deny matching non-ASCII text (#1389)
Go's regexp (RE2) treats \b as an ASCII-only word boundary, so the
privilege_escalation deny pattern \bsu\b matched the ASCII "su" inside
non-ASCII words such as the Vietnamese "suất" — the accented 'ấ' is not
an ASCII word char, so \b sees a spurious word boundary right after "su".

Any exec command containing such a word (for example a Vietnamese
document embedded in a python heredoc to render a PDF) was wrongly
blocked as a privilege-escalation attempt. Replace the pattern with a
Unicode-aware boundary (\pL, \pN) that still matches the real su command.
2026-07-08 07:23:05 +07:00
Syed Osama Ali Shah e65f70e2ef fix(ui): strip trailing dashes in slugify (#1387)
slugify's docstring promises "no leading/trailing dashes" and the sibling
isValidSlug rejects a trailing dash, but slugify only stripped leading
dashes. Any input ending in a non-[a-z0-9] character (e.g. "hello world!",
a trailing space, "trailing---") produced a slug ending in "-" that then
failed isValidSlug, causing invalid persisted keys and spurious validation
errors on agent/channel/provider/cron/mcp name fields. Add a symmetric
trailing-dash strip, matching the Go helpers (strings.Trim(s, "-")).
2026-07-08 05:53:31 +07:00
ナムandCollective Developer a9c7feb99c fix(channelmemory): harden passive extraction reliability (#1386)
Co-authored-by: Collective Developer <man@collective.dev>
2026-07-08 01:25:08 +07:00
nguyenha935andnguyenha935 086532484d feat: add configurable appearance settings (#1384)
Co-authored-by: nguyenha935 <208228297+nguyenha935@users.noreply.github.com>
2026-07-07 19:54:37 +07:00
Zezae Oh 4d3f0a4e1d fix(providers): tolerate unknown Claude CLI deny rules (#1381) 2026-07-07 17:52:59 +07:00
fastopencn 93eef5feac refactor(protocol): centralize tenants.* RPC method names as constants (#1371)
The tenants.* methods were the only RPC group without MethodXxx
constants in pkg/protocol/methods.go. Their wire names were hardcoded
as raw string literals across three call sites — router registration,
and the read/admin allowlists in the permission policy — which is
inconsistent with the other 147 methods and risks silent drift: a
typo in any allowlist string is not caught by the compiler and would
fail-closed (deny) at runtime, which is hard to diagnose.

Add MethodTenants{List,Get,Create,Update,UsersList,UsersAdd,
UsersRemove,Mine} and replace the method-identifier literals in
internal/permissions/policy.go and internal/gateway/methods/tenants.go
with the constants.

No behavior change — string values are identical. Human-facing error
labels and log prefixes in internal/http/tenants.go are intentionally
left as free-form text, matching the sibling HTTP handler convention.
2026-07-07 12:54:08 +07:00
MToan b5895c0a05 fix(docker): add tzdata to runtime image for zoneinfo support (#1372)
Python skill scripts using zoneinfo.ZoneInfo (and Go time.LoadLocation with
named zones) fail on the published images because alpine:3.23 ships without
tzdata: ZoneInfoNotFoundError: 'No time zone found with key Europe/Paris'.

Add tzdata to the base runtime apk install so every image variant (base,
latest, full, claude-cli) has the system zoneinfo database.
2026-07-07 12:53:48 +07:00
MToan afd0bd89e6 fix(claude-cli): drop TodoRead/NotebookRead from always-blocked deny rules (#1376)
Current Claude CLI releases (>= 2.x) no longer register TodoRead and
NotebookRead. Passing them via --disallowedTools makes the CLI print
'Permission deny rule "..." matches no known tool' on every invocation,
including background episodic summarization, burying real errors.

Keep TodoWrite/NotebookEdit blocked (still registered, no GoClaw
equivalent). Add a regression test guarding both directions.

Fixes #1374
2026-07-07 11:54:23 +07:00
MToan 09d8660f3b fix(mcp): derive claude_cli bridge tool surface from agent policy (#1377)
The bridge exposed a static BridgeToolNames subset that drifted from the
tool registry: use_skill, datetime, knowledge_graph_search and skill_manage
were never added, while the system prompt's skill-loading protocol requires
agents to call use_skill. claude_cli agents following the protocol hit a
nonexistent tool and could fabricate results.

Implement the structural fix recommended in #1373 triage:
- register the full bridge-capable surface (registry minus hard exclusions
  spawn/create_forum_topic) instead of the static list
- gate BOTH tools/list (new WithToolFilter) and tools/call through one
  shared predicate bridgeToolAllowed:
  * callers WITH a verified agent policy get exactly the policy-filtered
    surface (same WouldAllow check the call path always enforced)
  * callers WITHOUT one (anonymous, or agent without tools_config) keep the
    legacy conservative BridgeToolNames set - no exposure widening
- downgrade the per-call denial log Warn->Info; list filtering makes probes
  of denied tools rare and the call gate is the intended enforcement point

Fixes #1373
2026-07-07 11:23:11 +07:00
MToan 7844df74f5 fix(cron): scope cron context to tenant slug so managed skills resolve (#1378)
Cron job execution set the tenant ID on its context but not the tenant
slug. Tenant-scoped filesystem paths (skills-store, workspace, media via
config.TenantScopedDir) key off the slug and fall back to an id-based
path when it is absent — a different directory than where HTTP/WS skill
upload materialized the files (which sets the slug). As a result a cron
agent turn in a non-master tenant saw NONE of its tenant's managed
skills: skill_search returned 0 results and the agent, unable to run the
skill, produced an ungrounded answer.

Add cronTenantContext() which resolves the tenant slug via TenantStore
and sets both WithTenantID and WithTenantSlug. Master tenant and
nil-store/lookup-failure paths fall back to id-only (prior behavior).
Thread TenantStore into makeCronJobHandler and runCommandCronJob.

Tested: added unit tests for cronTenantContext (slug injected for
non-master; master skips lookup; nil store and lookup error fall back to
id-only). Verified end-to-end on a live tenant: before, a daily-agenda
cron guessed an empty day; after, it read the real event from the DB.

Note: other background executors that build a context from a tenant ID
(e.g. heartbeat) likely share this gap and are worth an audit.
2026-07-07 10:52:17 +07:00
ナムandCollective Developer 2a082f4edf fix(discord): preserve channel agent routing (#1380)
Co-authored-by: Collective Developer <man@collective.dev>
2026-07-07 09:53:02 +07:00
nguyenha935 f826738ee6 feat: add team work classification routing (#1379)
Approved by github-maintain bot. Clean team work classification feature with comprehensive tests and i18n.
2026-07-07 01:21:12 +07:00
ナムandCollective Developer f31214f5ce feat(channel-memory): enhance passive extraction for Discord (#1361)
* fix(channels): persist pending history under instance names

* feat(channel-memory): add run-all extraction workflow

* feat(discord): resolve passive memory channel titles

* feat(web): enhance passive memory controls

* docs(channel-memory): document extraction endpoints

---------

Co-authored-by: Collective Developer <man@collective.dev>
2026-07-06 11:51:07 +07:00
c1b89e96e1 chore: merge main into dev (#1370)
* fix: unify NO_REPLY detection (#1233)

Co-authored-by: GoClaw Operator <operator@goclaw>

* feat(acp): surface session/update tool_call notifications at Info level (#1141)

session/update notifications carrying a ToolCall (or inline tool_call /
tool_call_update kind) were only dumped at Debug level inside the params
blob. Operators could see `security.tool_granted` (permission granted) but
had no way to tell whether the tool actually executed successfully — both
"permission granted then failed silently" and "permission granted and
succeeded" looked identical in journalctl.

Add a structured Info log emitting toolCallId, name/title, status, and a
content preview (truncated at 400 chars) whenever the notification
contains tool-call state. This is what made it possible to diagnose the
recent .goclaw/-path-deny regression — `status=failed` immediately after
`security.tool_granted` revealed the gap that the granted-only log hid.

No behavior change beyond logging volume; preview is bounded.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(exec): exempt venv python interpreter from .goclaw/ path deny (#1140)

The ExecTool path-deny rule blocks any token containing `.goclaw/` unless
it matches one of the AllowPathExemptions prefixes (skills-store, tenants).
This silently rejected legitimate commands invoking the goclaw-managed
Python interpreter via its absolute path:

    /home/user/.goclaw/venv/bin/python3 .../script.py

The first token `/home/user/.goclaw/venv/bin/python3` matched the deny
pattern but no exemption, so the entire command was denied.

Naive exemption (".goclaw/venv/bin/") does not work: matchesAnyPathExemption
resolves both tokens and exemption candidates via EvalSymlinks, and the
venv's python3 is a symlink into the host's python cellar (e.g. linuxbrew).
The token canonicalizes to /home/linuxbrew/.../python3.14 while a literal
".goclaw/venv/bin/" prefix never gets touched.

Fix: resolve venv/bin/python3 once at startup and exempt the dirname of
the resolved target. Failure to resolve (no venv present) silently falls
through.

Without this, ACP-driven agents either fail outright or work only via
fragile heuristics (cwd-local symlinks generated on the fly by the LLM).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(cron): invert stateless gate to match UI label (#1139)

The cron job handler reset the session when stateless=false and skipped
reset when stateless=true — the opposite of what every UI locale labels
the field (en/ko/zh/vi all describe stateless as "each run starts fresh
without loading previous messages").

The buggy gate caused stateless=true crons to silently accumulate session
history across every execution, leading to context bloat and increasing
the chance of LLMs short-circuiting tool calls in favor of replaying
prior assistant turns. One affected daily ETL cron grew to 38 messages
over 18 days before the regression was noticed.

Fix: gate the Reset on `if job.Stateless` so the runtime matches the UI
contract. No DB migration is required — existing values were set by users
based on the UI label, so they already encode the intended behavior.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(providers): retry transient Codex response failures (#1332)

* fix(providers): retry transient Codex response failures

* test(cron): tolerate nil session store in handler tests

* fix(background): clear stale provider alerts (#1362)

Co-authored-by: Collective Developer <man@collective.dev>

---------

Co-authored-by: Duy /zuey/ <duy@wearetopgroup.com>
Co-authored-by: nguyenha935 <nguyenthanhha935@gmail.com>
Co-authored-by: GoClaw Operator <operator@goclaw>
Co-authored-by: codebit0 <34156842+codebit0@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Zezae Oh <zezaeoh@gmail.com>
Co-authored-by: Collective Developer <man@collective.dev>
2026-07-06 10:49:54 +07:00
a3fe5f7858 fix(mcp): resolve tool descriptions/schemas consistently across all MCP tool-listing paths (#1369)
* fix(mcp): return bare tool names with descriptions for already-connected MCP servers, fixing empty descriptions/schemas in agent prompts

* fix(mcp): resolve live tool descriptions/schemas in ListToolsForAgent for legacy prefixed tool_allow entries

The previous fix (d612da8e) only corrected the admin "browse/select tools"
endpoint. The actual runtime path building the LLM system prompt's tool list,
Manager.ListToolsForAgent, still looked up tool_cache/registry entries keyed
by the raw stored tool_allow name. Agents with grants captured before that
fix have tool_allow persisted as registered (prefixed) names instead of bare
MCP tool names, so every lookup missed and every tool in the prompt preview
came back with an empty description and schema.

Add bareMCPToolName to normalize both legacy-prefixed and current bare
tool_allow/tool_deny entries before matching, and resolveMCPToolInfo to
prefer the live registry's current description/schema (falling back to the
settings tool_cache) instead of trusting whatever was persisted verbatim.
Both the explicit-allow and cache-enumeration branches now share this single
resolution path.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-06 04:39:07 +07:00
Bruno ClermontandBruno Clermont ca097052b6 fix(skills): allow archived skills to be found by slug/name lookup, matching UUID lookup and ListSkills behavior (#1368)
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
2026-07-06 03:09:08 +07:00
Bruno ClermontandBruno Clermont db8ac12b16 feat(providers): allow LAN Ollama hosts via GOCLAW_OLLAMA_ALLOWED_HOSTS (#1367)
Adds an operator opt-in env var (comma-separated hostnames/IPs) to
whitelist additional local hosts for ollama/acp provider-type URL
validation, on top of the existing hardcoded localhost allowlist.

Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
2026-07-06 00:38:43 +07:00
Bruno ClermontandBruno Clermont 75d89da5e7 fix(tracing): dedupe tool_call span rows in hourly snapshot aggregation to prevent ON CONFLICT DO UPDATE double-affect error (#1366)
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
2026-07-05 23:38:56 +07:00
ナムandCollective Developer af80890edf fix(channel-memory): retry truncated extraction JSON (#1364)
Co-authored-by: Collective Developer <man@collective.dev>
2026-07-05 22:09:59 +07:00
Bruno ClermontandBruno Clermont 374defe04e fix(skills): allow editing SKILL.md content for database-managed skills (#1363)
* feat(skills): allow editing SKILL.md content for database-managed skills

Adds the ability to edit a skill's actual SKILL.md content directly
from the web UI, restricted to database/tenant-managed skills only --
bundled/system skills remain read-only.

Backend: new PUT /v1/skills/{id}/files/{path...} endpoint
(handleWriteFile in skills_versions.go), writing to the current
version directory of a non-system skill only (403 for is_system
skills), with the same path-traversal/symlink/ownership guards as the
existing read endpoint. Bumps skill store version and emits
cache-invalidate + audit event on success.

Frontend: Content tab in the skill detail dialog gets an "Edit
content" button, disabled with a tooltip for system skills. Editing
fetches the RAW file content via the existing read endpoint (which
does not strip YAML frontmatter, unlike the stripped preview shown in
the Content tab by default) to avoid silently destroying frontmatter
on save. Save/cancel follow the established catch-and-toast error
pattern (no uncaught promise rejections, matching PR #1346's config
save button fix).

i18n keys added to all 4 locales (en, ko, vi, zh).

* fix(skills): bump version on content edit, fix scroll regression in content view

handleWriteFile previously overwrote the current skill version in place
instead of creating a new immutable version, inconsistent with the
skill_manage tool's patch action. Now creates a new version directory
and updates the DB pointer, matching that convention.

Also fixes a scroll regression in the skill detail dialog's Content
tab introduced by the edit-mode textarea -- both the read-only preview
and edit textarea are now properly scrollable within the dialog.

* fix(skills): refresh version display after save, fix dialog height constraint for scrolling

Version bump now correctly reflects in the UI immediately after save
without requiring a page reload. Fixed the actual root cause of the
scroll issue: the dialog's height wasn't bounded, so overflow-y-auto
on inner content had no effect since nothing constrained the dialog's
total height in the first place.

---------

Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
2026-07-05 22:09:27 +07:00
thotam 84f06c0443 feat(web): restore agent Shares tab for per-user access sharing (#1360)
Approved by github-maintain bot. Clean UI restoration of agent Shares tab.
2026-07-05 13:56:24 +07:00
thotam dcf000ab6a fix(http): allow Operator to read providers and OAuth status (#1075) (#1359)
Browser pairing authenticates users as RoleOperator, but every provider
and OAuth route was wrapped in requireAuth(RoleAdmin). The /setup flow reads
GET /v1/providers, so paired browsers got 403 and hung on the setup page.

Split provider/OAuth route auth by operation:
- providers.go: GET list/get/models/status/embedding/codex-activity now use a
  readAuth wrapper (method-derived min → Viewer); POST/PUT/DELETE and
  reconnect/verify stay Admin. Responses already mask API keys and queries
  stay tenant-scoped, so no secrets are exposed.
- oauth.go: GET status/quota use readAuth (they return only connection state
  and quota, never tokens); start/callback/logout stay Admin (preserving the
  #450 admin-only-management decision for mutations).

Updates the two auth tests to assert the new contract: an Operator can read
providers/OAuth status (200) but still cannot mutate them (403).
2026-07-05 12:56:13 +07:00
thotam cab639d15e fix(migrate): use iofs source so migrations load on Windows (#1358)
`goclaw migrate up` (and every other migrate subcommand) failed on Windows
with `create migrator: failed to open source, "file:///D:/.../migrations":
open .: The filename, directory name, or volume label syntax is incorrect.`

golang-migrate's file source driver mis-parses absolute drive-letter file://
URLs; the drive-letter formatting in absoluteToFileURI produced a URL the
driver could not open. Replace the file:// URL with an iofs source over
os.DirFS, which uses native OS path handling and works identically on every
platform. Drop the now-unused absoluteToFileURI/migrationsSourceURL helpers
and their URL-shape tests, and add a DB-free regression test that opens the
real migrations directory via newMigrationSource().

Verified end-to-end on Windows: `migrate version` and `migrate up` now run.
2026-07-05 12:26:42 +07:00
thotam 25eaa0166f fix(backup): repair tenant backup/restore (config_secrets order, hooks registry, restore conn check) (#1357)
Closes #1076, #1338.

Fix A — tenant backup aborted with SQLSTATE 42703 "column id does not
exist" because exportQuery() hardcoded ORDER BY id. Several tenant-scoped
tables have composite PKs and no id column. Add a TableDef.OrderBy field,
honor it in exportQuery(), and set it for every id-less registry table:
config_secrets, agent_team_members, tenant_hook_budget, system_configs,
builtin_tool_tenant_configs, skill_tenant_configs, user_agent_profiles.

Fix B — hooks, tenant_hook_budget and webhook config were missing from
the backup registry, silently dropping their data on backup/restore. Add
hooks, tenant_hook_budget, webhooks (preserve) plus the hook_agents junction
(via ParentJoin through hooks, composite PK), and mark hook_executions and
webhook_calls as ephemeral in the skipped list.

Fix C — restoring on a fresh server failed with "N active DB connection(s)
detected" because the gateway's own pool connections were counted as active
clients. Tag pool connections with application_name='goclaw' (pg.OpenDB) and
exclude them in CheckActiveConnections; genuine external clients still block.

Adds unit + integration regression tests, including an export-over-every-
registered-table test that surfaced the additional id-less tables.
2026-07-05 11:56:36 +07:00