mirror of
https://github.com/tiennm99/goclaw.git
synced 2026-10-11 12:18:59 +00:00
e3dad3c60ceb68030e50397a68dfb658314f85d9
470
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e3dad3c60c |
fix(tts): unify voice resolution — dashboard settings affect tts tool (#1175)
The dashboard /tts page writes to system_configs[tts.<provider>.voice], but the LLM-invoked tts tool was only checking args > agent OtherConfig > builtin_tool_tenant_configs[tts].default_voice_id. Two different storage locations → the user's chosen voice was ignored, Edge defaulted to en-US-AriaNeural even when "HoaiMy" (vi-VN-HoaiMyNeural) was set correctly in the dashboard. Add system_configs as a 4th-level fallback so the dashboard becomes the single source of truth. - TtsTool gains SetSystemConfigStore(s) setter - resolveVoiceAndModel takes providerName + looks up tts.<provider>.voice/model when no higher-precedence source set - effectiveProvider resolved BEFORE voice/model so the right key is hit - Wired at gateway boot in cmd/gateway.go (after pgStores ready) - 4 unit tests covering: fallback, arg precedence, no-store, empty provider Verified against production trace 019e6036-44bb-703b-85fa-dee34f7ab2c0 where the tts tool was called with provider=edge, no voice arg, and defaulted to English instead of the configured Vietnamese voice. |
||
|
|
cc9956c1f9 |
feat(audio): openai_compat TTS/STT provider for self-hosted endpoints (#1447)
* feat(audio): add openai_compat TTS/STT provider for self-hosted endpoints
Adds a provider pair speaking the OpenAI audio wire format against any
compatible endpoint (gpu-manager, Speaches, vLLM, LocalAI, llama.cpp,
Ollama), so a self-hosted engine no longer needs a bespoke HTTP shim
implementing goclaw's proprietary /transcribe_audio contract.
Kept distinct from the openai package on purpose. Both speak the same
wire format, but "openai" carries api.openai.com semantics, and
audio.IsVoiceCompatible applies OpenAI's voice allowlist to any provider
by that name — FilterVoiceForProvider then silently substitutes "alloy"
for anything outside it. A self-hosted voice ID such as Piper's
"fr_FR-gilles-low" would be replaced with no error and no log line, and
the caller would hear the wrong language. Providers without validation
rules pass through untouched, which is the correct behaviour for an
engine whose voice namespace is its own.
Other deliberate departures from the openai package:
- api_base is required. There is no public endpoint to fall back to, and
defaulting to api.openai.com would send self-hosted traffic to a vendor.
- An empty api_key omits the Authorization header rather than sending an
empty Bearer token; self-hosted engines commonly have no auth.
- Empty voice/model are omitted from the request instead of sent blank,
since some engines resolve the model from the requested voice.
- The multipart filename is never blank: several engines type the audio
by extension rather than Content-Type and reject what they cannot type.
Endpoint error bodies are preserved in the returned error so a format
mismatch (e.g. an engine that cannot encode mp3) is diagnosable.
* feat(audio): wire openai_compat TTS/STT provider into gateway config
Adds tts.openai_compat config block and registers both providers at
startup. api_base is the enable switch rather than an API key, since
self-hosted endpoints commonly have no auth and there is no vendor
default to fall back to.
The STT chain is now built from the providers that actually registered
instead of being hardcoded to {elevenlabs, proxy}. Transcribe skips
unregistered names with a warning on every call, so a static chain
naming absent providers is log noise on the hot path. openai_compat goes
first when present: it is an explicit operator choice and keeps audio on
the local network. "proxy" stays last, as BridgeLegacySTT registers it
later per channel.
Behaviour is unchanged when tts.openai_compat.api_base is unset: the
chain resolves to {elevenlabs, proxy} exactly as before.
---------
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
|
||
|
|
dc605b911e |
feat(mcp): expand CRUD server to near-complete CLI parity (#1440)
* feat(mcp): expand CRUD MCP server to near-complete CLI parity Closes the CLI-vs-MCP tool coverage gap identified by auditing every `goclaw` CLI command against the existing goclaw_* MCP tool set. Adds skill grant/revoke, agent skill pin/unpin, and full CRUD/inspection surfaces for memory, knowledge graph (including dedup/merge/prune), tenants, providers, LLM traces, channel contacts, pending messages, audit activity, system config, tenant storage (list/size/delete/move), scoped agent config export/import, secure-CLI binary registry, and a DB-backed health check. Deliberately out of scope, documented inline where relevant: - `goclaw credentials`: confirmed CLI-local (~/.goclaw/config.yaml + keychain), no server resource to wrap. goclaw_secure_cli_binaries_* covers the closest real, previously-uncovered server resource instead. - Full tar-archive agent export/import (KG + workspace files): the CLI's version streams a multi-section archive with progress events, a shape that doesn't map to a single MCP tool call. Config + context files (the portable "brain") is covered. - `kg extract` (LLM-driven text extraction): goclaw_kg_ingest accepts the same Entity/Relation shapes the extractor produces, so a caller can run extraction itself and hand off the result. Wires 8 new store dependencies (Memory, KnowledgeGraph, Tracing, Contacts, PendingMessages, Activity, SystemConfigs, SecureCLI) through gateway.Server setters -> cmd/gateway.go -> CRUDDeps, following the existing Providers/Tenants pattern. Storage and secure-CLI-binary handlers duplicate internal/http's path-escape/symlink-hiding validation logic (documented inline) since internal/http already imports internal/mcp and the reverse would cycle. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(mcp): return sessionKey in goclaw_chat_send response for persistent conversations The goclaw_chat_send tool creates a new session internally when sessionKey is empty, but never returned the key to the caller. This prevented using goclaw_chat_history to fetch previous messages in the session. Add SessionKey field to ChatSendResult so callers can: 1. Start a new agent chat without providing sessionKey 2. Receive the sessionKey back in the response 3. Use that sessionKey for follow-up messages and history queries Fixes the training loop pattern: start chat → get sessionKey → call goclaw_chat_history with that key → iterate skill based on actual failures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(mcp): add timing diagnostics to goclaw_chat_send for timeout troubleshooting Log request arrival, processing duration, and errors with millisecond precision. Helps identify whether timeouts occur at MCP client→goclaw, goclaw→ollama, or during agent execution. Critical for production debugging. Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com> * fix(gateway): add http.Server timeouts for defensive timeout handling Set explicit timeouts in http.Server: - ReadTimeout: 1h (allow large uploads, long-running agent operations) - WriteTimeout: 1h (allow streaming responses to slow clients) - IdleTimeout: 30s (close idle keep-alive connections quickly) Provides defense-in-depth when Nginx/Traefik timeouts are misconfigured. Matches Nginx timeout (3600s) to prevent race conditions. Timeout chain: traefik 3600s = nginx 3600s = goclaw 3600s Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com> --------- Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
293aa2a5e1 |
(feat) MCP Server (#1365)
* feat(hooks/observe): add structured output validation hook (ObserveHook) - hooks/types.go: Add ObserveHook interface + BuiltInHookType enum - hooks/dispatcher.go: HookDispatcher emits ObserveHook with PhaseResult payload - pipeline/observe_stage.go: ObserveStage emits hook after ObserveResult built - pipeline/substates.go: ObserveStageResult carries hook results + validation errors - hooks/config.go: BuiltInHookTypeObserve added to BuiltInHookType enum PhaseResult carries structured output, token usage, tool calls + validation errors emitted post-ObserveStage. ObserveHook implementations can validate structured output against schemas, detect tool-call loops, enforce token budgets, etc. Hook fires after ObserveStage produces ObserveResult, before results propagate to next stage. ValidationError returned by hook halts pipeline and propagates error to caller without further stage execution. Co-Authored-By: Claude <noreply@anthropic.com> * ui(hooks): add post_model_response event to web UI - Add event to Zod schema, filter dropdown, and form dialog - Implement conditional test panel UI for model response payload - Add translations (en/zh/vi) for new test panel fields - Updated beta description to reference the new event Co-Authored-By: Claude <noreply@anthropic.com> * feat(mcp): add MCP CRUD server exposing goclaw resources at /api/mcp/ with Bearer token auth and X-GoClaw-Tenant-Id header, default to master tenant * feat(mcp): add goclaw_skills_write_file tool to edit skill files on disk The CRUD MCP server's goclaw_skills_update only touched skill DB metadata, with no way to edit a skill's SKILL.md/file content on the filesystem. Extract the versioned write logic from the web UI's skill file editor (SkillsHandler.handleWriteFile) into skills.WriteVersionedFile so both surfaces share identical validation and versioning, and expose it as a new MCP tool. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com> Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
4f558e2c3b |
fix(ollama): resolve context window per model via POST /api/show (#1437)
The Ollama integration never applied a correct context window, so agents
with real prompts (20k-100k tokens) were rejected with HTTP 400
exceed_context_size_error against a 4096-token default.
Three root causes fixed:
1. FetchOllamaModelContext issued a GET to /api/show, which Ollama answers
with 405 (the endpoint is POST-only). Now POSTs {"model": "<name>"}.
2. The response parser expected a flat model_info.context_length, but a real
Ollama server namespaces the key by architecture (gemma4.context_length,
qwen3.context_length, ...). extractContextLength now matches "context_length"
or any "*.context_length" key.
3. num_ctx was resolved once at startup for a hardcoded "llama3.3" model and
never for the model an agent actually uses. Resolution now happens per
request for the real model inside OllamaProvider.resolveNumCtx, cached under
an RWMutex, with an explicit settings override winning and the fetched value
bounded by OllamaDefaultNumCtx so an enormous advertised window (Qwen3.5
reports 262144) cannot balloon the KV cache beyond VRAM.
Also classify Ollama's "exceed_context_size" 400 as a context-overflow error so
the pipeline's emergency-compaction+retry path (Issue 958) engages gracefully
instead of surfacing a raw error.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
27f4b7743b |
fix(skills): preserve custom skills during bundled seeding (#1435)
Co-authored-by: ntduc <ntduc@cpp.ai.vn> |
||
|
|
6167a53bbe |
fix(bitrix24): mount shared webhook route independent of bot readiness [B24:2794] (#1432)
/bitrix24/* only mounted when a bitrix24 Channel loaded successfully via WebhookHandlers(), which requires portal + bot_code + bot_name already set. But completing portal OAuth requires /bitrix24/install to be reachable, and the UI won't let you pick an uninstalled portal when creating a bot -- deadlock on any fresh deployment: no bot can load until OAuth completes, OAuth can't complete without the route mounted. Claim and mount the shared bitrix24.WebhookRouter() singleton directly at boot, independent of whether any Channel loaded. ClaimWebhookRoute is idempotent (first-claim-wins via CompareAndSwap), so this is a no-op once a real Channel has already claimed the route. Co-authored-by: DangTinh311 <dangtinh31193@gmail.com> |
||
|
|
6b738924d5 |
fix(discord): qualify group titles across surfaces (#1420)
Co-authored-by: Collective Developer <man@collective.dev> |
||
|
|
db69ccf882 |
tracing: nest tool calls under LLM spans, emit hook spans, fix usage_events FK (#1415)
- Reparent tool_call spans under their producing llm_call span via RunState.CurrentLLMSpanID so a model turn's trace visibly contains the tool calls it triggered. - Wire internal/hooks EmitHookSpan into the dispatcher writeExec so every hook execution (pre/post tool use, etc.) produces a trace span with input, console output, decision, error, and duration for agent troubleshooting. - Fix usage_events_span_id_fkey violations: usage events were inserted synchronously referencing a span flushed ~5s later, silently dropping token/cost data. Route them through the collector so they flush after spans in the same cycle. No schema migration; FK preserved. Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com> Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
84fab59226 |
fix(memory): guide recency recall tool calls (#1412)
* fix(memory): guide recency recall tool calls * fix(tools): align memory and vault builtin visibility --------- Co-authored-by: Collective Developer <man@collective.dev> |
||
|
|
2b7bfec504 |
feat(telegram): agent-callable message reaction (message action=react) (#1407)
Adds a "react" action to the message tool so an agent can set an emoji reaction on an existing message (e.g. mark its own status post 👍 once fully paid). Mirrors the ChannelEditor wiring: ReactionSetter capability -> Manager.ReactToMessage -> telegram Channel.ReactToMessage (validates against Telegram's supported reaction set; ✅/❌ are rejected). Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ba621391c2 |
feat(telegram): trigger-words, channel posts, message edit & topic posting (#1383)
- Fix image_generation nil-pointer crash on codex agents: the sentinel now carries a name-only Function so the many sites reading td.Function.Name never nil-deref; codex_build still branches on Type. - Agent-declared trigger words in IDENTITY.md wake the bot in groups without an @mention (whole-word, Cyrillic-aware; text + caption), cached per-agent 60s. - channel_post support with a synthetic sender + a recover() guard so a malformed update can't crash the gateway. - message tool action=edit (editMessageText + editMessageCaption fallback), targeting the replied-to message via reply_to_message_id. - message tool topic=<name> posts into a named forum topic; topics learned from forum_topic_created into channel_contacts and resolved by name. Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
51ed7845e2 |
fix(bitrix24): register newly-created portal in live router [B24:2794] (#1404)
bitrix.portals.create only persisted the DB row — it never called Router.RegisterPortal, so a portal created while the gateway was already running stayed invisible to PortalByDomain/PortalByKey until the next full restart. The only other RegisterPortal call site is BootstrapPortals, which runs once at boot. Result: the OAuth install callback 404'd with "unknown portal" for any portal created through the self-service UI mid-uptime — hit live on web1trang.bitrix24.com. Fix: inject registerPortal/unregisterPortal func fields into BitrixPortalsMethods (mirrors the existing gatewayPublicURL func() pattern), defaulting to bitrix24.WebhookRouter() + NewPortal/ RegisterPortal (same construction path as BootstrapPortals) and Router.UnregisterPortal respectively. handleCreate calls registerPortal after a successful Create; handleDelete calls unregisterPortal after a successful Delete for symmetry. Both are best-effort — a registration failure is logged, not returned as an RPC error, since the DB row is already valid either way. Injected via func fields (not calling bitrix24.WebhookRouter() inline) so tests can stub live-router registration without touching that process-wide singleton, whose test-reset helper isn't exported outside the bitrix24 package. Tests: 4 new cases covering register-on-create, best-effort failure handling, unregister-on-delete, and no-unregister-when-delete-blocked. All 20 tests in bitrix_portals_test.go pass; go build (both PG and sqliteonly tags) and go vet clean. Co-authored-by: DangTinh311 <dangtinh31193@gmail.com> |
||
|
|
d856a58218 |
feat(providers): make provider request timeout configurable (default 30s) (#1400)
The provider verify call and the provider models-list call used hardcoded timeouts (30s and 15s) that were too low for slow models. Introduce a tenant-scoped, UI-editable setting `providers.request_timeout_sec` (default 30) stored in system_configs, consumed by both handlers, and overlaid from config.json like tts.timeout_ms. Adds a field to the System Settings modal (all 4 locales) so operators can raise it without a redeploy. Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com> |
||
|
|
951acad714 |
feat(tools): seed missing built-in tools into builtin_tools catalog (#1399)
Eleven built-in tools were registered in the runtime tool registry (and thus appeared in agent system prompts and were subject to allow/deny) but were absent from builtinToolSeedData(), so operators could not see or toggle them in the agent allow/deny UI. Seed datetime, heartbeat, memory_expand, list_group_members, zalo_list_groups, vault_search, vault_read, delegate, workstation_exec, claude_remote, and mcp_tool_search so the UI catalog reflects the actual registered tool set. telegram_manager is intentionally excluded per existing test guard. Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com> |
||
|
|
6bafe54b08 |
fix(agent): filter per-tenant disabled tools from system prompt (#1396)
The system prompt's Tooling section previously listed tools even when they were disabled for a tenant and already stripped from the API tools parameter, confusing the LLM into thinking it could use them. Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com> |
||
|
|
aa15317d5b |
fix: resolve Zalo group by name + notify origin on failed forward (#1395)
* fix: resolve Zalo group name to real chat ID, notify origin on failed forward Two related root causes behind "forward message to group by name" silently failing while the agent reports success: 1. No tool let the agent resolve a group's display name (e.g. "Ban Dieu Hanh") to the chat ID the message tool actually requires. sessions_list only exposes session keys (numeric group IDs), never human-readable names, so the agent had no reliable way to turn a name into a real target and ended up passing the display name itself as `target`. Adds zalo_list_groups, wrapping the already-used (dashboard picker) protocol.FetchGroups behind the same optional-interface pattern as list_group_members/GroupMemberProvider (GroupListProvider on channels.Manager, gated to zalo_personal via RequiredChannelTypes). 2. When the resulting send fails downstream (e.g. Zalo rejects a bad chat_id), dispatchOutbound only ever retried/notified media failures on the same (already-broken) destination, and dropped text-only failures entirely — even though message.go's own postCrossTargetNotice comment states forwards must never announce a fake delivery. Because the bus publish is fire-and-forget, the tool had already returned "sent" and announced success to the origin chat before the real send was even attempted. message.go now tags cross-target forwards with origin channel/chat in OutboundMessage.Metadata; dispatchOutbound uses it to notify the ORIGIN chat with the real failure instead of silently dropping it or retrying against the same invalid target. * test: cover forward-origin metadata tagging and dispatch failure notice Extracts dispatchOutbound's error branch into handleSendFailure so it can be unit tested without driving the consumer loop/goroutine, and adds coverage for: forward failures notifying the origin chat (not the broken destination), pre-existing non-forward media/text-only behavior staying unchanged, message.go tagging cross-target group forwards with origin metadata alongside group_id, and the new zalo_list_groups tool/Manager delegator. |
||
|
|
fc7d183961 |
Enhance passive memory extraction tuning (#1393)
* feat(channelmemory): add passive extraction tuning * feat(ui): expose passive memory tuning controls * docs(memory): document passive extraction tuning --------- Co-authored-by: Collective Developer <man@collective.dev> |
||
|
|
3b8bc1bc4d |
feat(bitrix24): replace two MCP text inputs with a filtered dropdown (#1392)
* feat(bitrix24): replace two MCP text inputs with a filtered dropdown [B24:2794]
Bitrix24 channel creation used to demand two hand-typed strings —
mcp_server_name and mcp_base_url — plus zero indication of which MCP
servers can actually auto-onboard. Typos silently disabled provisioning
and admin had to know which servers implement /api/auto-onboard.
Ship a single dropdown backed by mcp_servers.require_user_credentials,
plus the machinery to make it work end-to-end.
Phase 1 — DB & store
* Promote require_user_credentials from settings JSONB to a top-level
column on mcp_servers (PG migration 000089, SQLite migration v54).
* Backfill from existing settings blobs so no admin needs to re-tick.
* Add MCPServerData.RequireUserCredentials to the Go store layer, plumb
through Create / Get / GetByName / List / Update on both stores,
extend the export DTO, and add require_user_credentials to the HTTP
allowlist.
* Bump RequiredSchemaVersion 87 -> 89 (jumping 88, which was on disk
but not wired) and SchemaVersion 53 -> 54 with an
idempotentColumnMigration guard.
Phase 2 — Bitrix24 channel factory
* Add MCPServerID (UUID string) to bitrix24 InstanceConfig, keep
MCPServerName + MCPBaseURL as legacy fallback with a "deprecated"
doc comment.
* Factory validation accepts either mcp_server_id alone or the legacy
pair; half-config still fails fast.
* initMCPProvisioner prefers GetServer(id) and sources the base URL
from mcp_servers.url when the id path is used. Legacy name path
unchanged so pre-migration configs keep working.
* Log line now carries mcp_server_id + require_user_credentials so
operators can eyeball the wiring.
Phase 3 — Frontend types & MCP form
* Add optional top-level require_user_credentials to MCPServerData /
MCPServerInput in both ui/web and ui/desktop/frontend types.
* mcp-form-dialog reads the top-level flag first and falls back to
settings.require_user_credentials so cached responses from
pre-upgrade backends still render correctly.
* On submit send both the top-level flag AND the legacy settings
entry so mid-rollout backends stay consistent.
Phase 4 — Bitrix24 channel form dropdown
* New mcp-select field type + MCPServerSelect component. Uses the
shared useMCP() react-query cache and filters client-side to
servers whose require_user_credentials is true (OR settings
JSONB during the migration window).
* Explicit "None (disable MCP provisioning)" option so admins can
clear the binding without editing config JSON.
* Legacy mcp_server_name / mcp_base_url text inputs kept in the
Advanced panel, relabelled "(legacy)" with pointer help text.
Phase 5 — channel_instances.config backfill
* PG migration 000090 and SQLite migration v55 rewrite existing
bitrix24 channel_instances.config to add mcp_server_id by
resolving mcp_server_name against mcp_servers, tenant-scoped
via agents.tenant_id (channel_instances doesn't carry tenant_id
directly).
* Idempotent — only touches rows already carrying
mcp_server_name that lack mcp_server_id. Legacy keys are left
in place so provisioner.go can still fall back for unmigrated
or future-created legacy configs.
* down.sql drops the mcp_server_id key. Provisioner immediately
reverts to the legacy fallback path.
Tests
* provisioner_test.go: three new cases exercise the mcp_server_id
path (invalid UUID string, valid UUID with missing row, valid
UUID with a per-user row). fakeMCPStore gains a serversByID
map and a real GetServer implementation.
* Existing legacy-config tests unchanged and still green.
Verification
* go build ./... && go build -tags sqliteonly ./...
* go vet ./internal/mcp/... ./internal/channels/bitrix24/...
./internal/store/... ./internal/http/...
* go test ./internal/mcp/... ./internal/channels/bitrix24/... -> ok
* Live-tested against a local docker image on the goclaw-deploy
postgres. Migrations 89 + 90 applied cleanly. Three existing
bitrix24 channels (bitrix-sales / nguyen-dao-openline / tieu-vi)
had their configs backfilled with the b24-syn-mcp UUID and the
provisioner boots with require_user_credentials=true. UI dropdown
correctly shows only b24-syn-mcp (the only server with the flag
ticked) alongside a "None" clearer option.
Surface parity
* Gateway server: store + factory + provisioner + HTTP allowlist.
* API contract: adds require_user_credentials + mcp_server_id
as optional fields on existing routes. No new endpoints.
* Web UI: MCP form + Bitrix24 channel form + shared types.
* CLI/runtime package: N/A because no CLI subcommand reads the
mcp_server_id field.
* fix(bitrix24): derive auto-onboard base URL from mcp_servers.url origin [B24:2794]
The Phase 2 refactor swapped provisioner base-URL sourcing from the
legacy per-channel MCPBaseURL config field (which historically stored the
MCP server's ORIGIN, e.g. https://mcp.example.com) to mcp_servers.url,
which stores the JSON-RPC ENDPOINT the agent loop dials (e.g.
https://mcp.example.com/mcp). The two are semantically different but
share a single column.
mcp_client.newMCPClient then appends "/api/auto-onboard" to whatever
baseURL it receives, so the id-path started POSTing to
".../mcp/api/auto-onboard" — 404 for every per-user credential mint and
refresh. Existing users kept working only until their cached access
tokens expired.
Fix: derive the origin (scheme://host[:port]) from server.URL before
handing it to the auto-onboard client. The legacy path is untouched
because MCPBaseURL from channel config is already the origin.
https://b24-mcp-dev.synity.so/mcp -> https://b24-mcp-dev.synity.so
https://mcp.example.com/mcp/ -> https://mcp.example.com
https://mcp.example.com -> https://mcp.example.com
Table-driven test covers six shapes plus four error cases (empty,
whitespace-only, no scheme, no host). Updated the existing
TestInitMCPProvisioner_MCPServerID fixture to seed a URL with the /mcp
subpath so it regression-guards the same code path.
Verified live: user 614 sent a message that triggered the expired-cred
refresh branch; goclaw logged "self-refreshed user credentials
created=false" and the agent immediately reported
"mcp.user_tools_loaded user=614 tools=2". Before this fix the same event
logged 'auto-onboard failed: mcp auto-onboard: 404 Not Found'.
Surface parity:
- Gateway server: provisioner + one new helper (deriveAutoOnboardBaseURL).
- API contract: N/A because the wire shape hasn't changed.
- Web UI: N/A because the UI still writes mcp_server_id verbatim.
- CLI/runtime: N/A because no CLI reads the derived base URL.
---------
Co-authored-by: DangTinh311 <dangtinh31193@gmail.com>
|
||
|
|
086532484d |
feat: add configurable appearance settings (#1384)
Co-authored-by: nguyenha935 <208228297+nguyenha935@users.noreply.github.com> |
||
|
|
7844df74f5 |
fix(cron): scope cron context to tenant slug so managed skills resolve (#1378)
Cron job execution set the tenant ID on its context but not the tenant slug. Tenant-scoped filesystem paths (skills-store, workspace, media via config.TenantScopedDir) key off the slug and fall back to an id-based path when it is absent — a different directory than where HTTP/WS skill upload materialized the files (which sets the slug). As a result a cron agent turn in a non-master tenant saw NONE of its tenant's managed skills: skill_search returned 0 results and the agent, unable to run the skill, produced an ungrounded answer. Add cronTenantContext() which resolves the tenant slug via TenantStore and sets both WithTenantID and WithTenantSlug. Master tenant and nil-store/lookup-failure paths fall back to id-only (prior behavior). Thread TenantStore into makeCronJobHandler and runCommandCronJob. Tested: added unit tests for cronTenantContext (slug injected for non-master; master skips lookup; nil store and lookup error fall back to id-only). Verified end-to-end on a live tenant: before, a daily-agenda cron guessed an empty day; after, it read the real event from the DB. Note: other background executors that build a context from a tenant ID (e.g. heartbeat) likely share this gap and are worth an audit. |
||
|
|
2a082f4edf |
fix(discord): preserve channel agent routing (#1380)
Co-authored-by: Collective Developer <man@collective.dev> |
||
|
|
f826738ee6 |
feat: add team work classification routing (#1379)
Approved by github-maintain bot. Clean team work classification feature with comprehensive tests and i18n. |
||
|
|
c1b89e96e1 |
chore: merge main into dev (#1370)
* fix: unify NO_REPLY detection (#1233) Co-authored-by: GoClaw Operator <operator@goclaw> * feat(acp): surface session/update tool_call notifications at Info level (#1141) session/update notifications carrying a ToolCall (or inline tool_call / tool_call_update kind) were only dumped at Debug level inside the params blob. Operators could see `security.tool_granted` (permission granted) but had no way to tell whether the tool actually executed successfully — both "permission granted then failed silently" and "permission granted and succeeded" looked identical in journalctl. Add a structured Info log emitting toolCallId, name/title, status, and a content preview (truncated at 400 chars) whenever the notification contains tool-call state. This is what made it possible to diagnose the recent .goclaw/-path-deny regression — `status=failed` immediately after `security.tool_granted` revealed the gap that the granted-only log hid. No behavior change beyond logging volume; preview is bounded. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(exec): exempt venv python interpreter from .goclaw/ path deny (#1140) The ExecTool path-deny rule blocks any token containing `.goclaw/` unless it matches one of the AllowPathExemptions prefixes (skills-store, tenants). This silently rejected legitimate commands invoking the goclaw-managed Python interpreter via its absolute path: /home/user/.goclaw/venv/bin/python3 .../script.py The first token `/home/user/.goclaw/venv/bin/python3` matched the deny pattern but no exemption, so the entire command was denied. Naive exemption (".goclaw/venv/bin/") does not work: matchesAnyPathExemption resolves both tokens and exemption candidates via EvalSymlinks, and the venv's python3 is a symlink into the host's python cellar (e.g. linuxbrew). The token canonicalizes to /home/linuxbrew/.../python3.14 while a literal ".goclaw/venv/bin/" prefix never gets touched. Fix: resolve venv/bin/python3 once at startup and exempt the dirname of the resolved target. Failure to resolve (no venv present) silently falls through. Without this, ACP-driven agents either fail outright or work only via fragile heuristics (cwd-local symlinks generated on the fly by the LLM). Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(cron): invert stateless gate to match UI label (#1139) The cron job handler reset the session when stateless=false and skipped reset when stateless=true — the opposite of what every UI locale labels the field (en/ko/zh/vi all describe stateless as "each run starts fresh without loading previous messages"). The buggy gate caused stateless=true crons to silently accumulate session history across every execution, leading to context bloat and increasing the chance of LLMs short-circuiting tool calls in favor of replaying prior assistant turns. One affected daily ETL cron grew to 38 messages over 18 days before the regression was noticed. Fix: gate the Reset on `if job.Stateless` so the runtime matches the UI contract. No DB migration is required — existing values were set by users based on the UI label, so they already encode the intended behavior. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(providers): retry transient Codex response failures (#1332) * fix(providers): retry transient Codex response failures * test(cron): tolerate nil session store in handler tests * fix(background): clear stale provider alerts (#1362) Co-authored-by: Collective Developer <man@collective.dev> --------- Co-authored-by: Duy /zuey/ <duy@wearetopgroup.com> Co-authored-by: nguyenha935 <nguyenthanhha935@gmail.com> Co-authored-by: GoClaw Operator <operator@goclaw> Co-authored-by: codebit0 <34156842+codebit0@users.noreply.github.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Zezae Oh <zezaeoh@gmail.com> Co-authored-by: Collective Developer <man@collective.dev> |
||
|
|
cab639d15e |
fix(migrate): use iofs source so migrations load on Windows (#1358)
`goclaw migrate up` (and every other migrate subcommand) failed on Windows with `create migrator: failed to open source, "file:///D:/.../migrations": open .: The filename, directory name, or volume label syntax is incorrect.` golang-migrate's file source driver mis-parses absolute drive-letter file:// URLs; the drive-letter formatting in absoluteToFileURI produced a URL the driver could not open. Replace the file:// URL with an iofs source over os.DirFS, which uses native OS path handling and works identically on every platform. Drop the now-unused absoluteToFileURI/migrationsSourceURL helpers and their URL-shape tests, and add a DB-free regression test that opens the real migrations directory via newMigrationSource(). Verified end-to-end on Windows: `migrate version` and `migrate up` now run. |
||
|
|
8000a1d0f5 |
feat(providers): add provider-level thinking/reasoning toggle for Ollama (#1355)
Ollama has two separate request-building code paths: OllamaProvider (native /api/chat, used for num_ctx control) and OpenAIProvider (OpenAI-compat /v1/chat/completions). An earlier fix disabled thinking mode by hardcoding think=false, but only in the OpenAI-compat path -- OllamaProvider.buildRequest() never set the think field at all, so reasoning-capable models (qwq, deepseek-r1) defaulted to visible chain-of-thought reasoning regardless of that fix. Confirmed live via a docker-engineer agent streaming full reasoning traces despite the existing disable. Replaced the hardcoded always-off behavior with a provider-level tri-state setting (llm_providers.settings.thinking_enabled: unset = default off, explicit true/false overrides), configurable via the provider's Advanced settings dialog. Both OllamaProvider.buildRequest() and OpenAIProvider.buildRequestBody() now read and respect this same setting, so the toggle works regardless of which Ollama code path a given deployment routes through. Added tests for setting parsing (unset/true/false/malformed) and both provider request-builders' handling of the override. Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com> |
||
|
|
e1fcccea4f |
feat: add configurable system messages (#1343)
Co-authored-by: GoClaw Operator <operator@goclaw> |
||
|
|
4d3c6bcb05 |
feat: add Telegram manager channel permissions (#1339)
Co-authored-by: GoClaw Operator <operator@goclaw> |
||
|
|
2810ad96f0 |
feat(providers): add AI/ML API preset (#1336)
Approved by github-maintain. Clean implementation of AI/ML API as branded OpenAI-compatible provider. All CI green. Closes #1325. |
||
|
|
8ea166f2e5 |
fix(providers): retry transient Codex response failures (#1334)
Backport of #1332 to dev — retry transient Codex failures |
||
|
|
90e7035100 |
fix(mcp): tenant isolation, tool policy engine, and prompt preview fixes (#1333)
* fix(prompt): remove duplicate Team Members section from system prompt The TEAM.md context file already provides the Members section with better formatting. The system prompt section was redundant and inconsistently formatted. - Remove buildTeamMembersSection() call from system prompt building - Remove unused buildTeamMembersSection() function * fix(mcp): wire store to manager for prompt preview tool visibility MCP manager needs database access to query configured servers for tool visibility in prompt preview. - Add SetStore() method to MCP manager - Wire pgStores.MCP to manager after initialization - Add debug logging for MCP initialization flow * fix(systemprompt): hide Tooling section when agent has no tools The Tooling section header and boilerplate were displayed even when the agent had no tools available. Add early return in buildToolingSection() to skip the section entirely when toolNames is empty. This reduces prompt noise for agents with no tool access. * feat(mcp): cache tool descriptions for prompt preview visibility MCP tools now show descriptions in prompt preview without requiring a live server connection. - Add CacheToolDescriptions() method to MCPServerStore (PG + SQLite) - Cache tool descriptions in settings['tool_cache'] when server connects - Use cached descriptions in ListToolsForAgent (fallback: hints → cache → global) - Descriptions are auto-populated from live server manifest on connection - Admins can still override via tool_hints in server settings * test(mcp): add CacheToolDescriptions to MCPServerStore test fakes Commit 5ce410b6 added CacheToolDescriptions() to the MCPServerStore interface but missed updating test mock implementations, breaking go vet across internal/mcp, internal/agent, internal/http, and internal/channels/bitrix24. Add no-op implementations matching each fake's existing style. * fix(providers): enforce per-agent tool policy for Claude CLI provider The Claude CLI provider (stdio+MCP bridge) was not enforcing per-agent tool policy, unlike other providers where the policy-filtered tool list already drives the system prompt's Tooling section. Two gaps closed: 1. --disallowedTools was previously skipped entirely when no MCP config path was resolved, letting the CLI subprocess run with its full native toolset (Bash, Edit, Read, Write, Glob, Grep, WebFetch, WebSearch) regardless of agent policy. It's now unconditional and derived from the agent's actual allowed-tools list (state.Tool.AllowedTools), mapped to Claude CLI's native tool names. 2. The MCP bridge server executed any tool call without checking the calling agent's policy. It now resolves the agent's policy from context (via the existing HMAC-verified agent lookup) and denies calls to tools outside that agent's allowed set, logging security.mcp_bridge_denied on denial. Both gaps were closed using existing plumbing (PolicyEngine.WouldAllow, AgentData.ParseToolsConfig, the bridge context middleware) — no new cross-cutting mechanism was introduced. * feat(web): show MCP/tool schemas in system prompt preview dialog The prompt-preview API response includes a separate `tools` field (the actual JSON schemas sent to the LLM as the tools API parameter) alongside `prompt` (the system prompt text), but the web UI only rendered `prompt`, silently dropping the tools list. Add a collapsible Tools section to both the full-screen System Prompt dialog and the inline agent-detail preview, showing tool count, name, description, and expandable parameter schema per tool. i18n keys added to en/vi/zh locales. * fix(prompt): render pinned skills on bootstrap turns Pinned skills are documented (web UI copy) as "always inlined in the system prompt", but the entire Skills section was gated behind !cfg.IsBootstrap, so pinned skill XML never appeared on bootstrap turns (first message of a session) despite the promise. Separate pinned-skill rendering from bootstrap-suppressed guidance: - Bootstrap + pinned skills present: render pinned XML only, no search/manage guidance (which stays suppressed as before) - Non-bootstrap: unchanged behavior - Minimal/none modes: pinned skills always render regardless of bootstrap state Add regression tests covering all four prompt modes on bootstrap turns, plus a non-bootstrap guard confirming existing behavior is preserved. * fix(skills): resolve managed skills directory per-tenant, not master-only skills.Loader was wired at startup to scan a single fixed directory (the master tenant's managed-skills dir), making any skill belonging to a non-master tenant invisible to both pinned-skills prompt resolution and skill_search/use_skill, regardless of DB visibility settings. - Loader now resolves the calling tenant's managed-skills directory per-call via context (store.TenantIDFromContext), never enumerating other tenants' directories - Skill cache is now tenant-keyed to prevent slug collisions and cross-tenant cache leaks across tenants using the same skill slug - gateway_setup.go passes the root data dir instead of a pre-resolved master-tenant path Write-side tooling (skill_manage, publish_skill) was already correctly tenant-scoped per-operation — no changes needed there. Added TestLoader_ManagedSkills_TenantIsolation proving two tenants with same-slug/different-content skills never see each other's content, including after cache population from a different tenant's lookup. Known follow-up (not in this commit): skill_search's BM25 index is still a single process-global index shared across tenants, which is a related but separate cross-tenant search-result leak requiring its own scoped fix (per-tenant index maps + threading tenant context through ensureIndex/rebuildIndex). * fix(tools): scope skill_search BM25 index per-tenant SkillSearchTool held a single process-global BM25 index built once from whichever tenant's context first triggered ensureIndex, then reused for all subsequent Execute() calls regardless of caller — leaking one tenant's skill search results into another's, the search-path counterpart to the managed-directory bug fixed in 7b4668ad. - index/lastVersion are now keyed per-tenant (map[uuid.UUID]*tenantIndexState) - ensureIndex resolves the calling tenant from context and only builds/reads that tenant's index entry, never touching another tenant's cached state - Builtin/bundled skills remain visible in every tenant's index (Loader already merges those tiers correctly per 7b4668ad) Loader.Version() remains a single global counter — a version bump in one tenant causes unnecessary rebuilds in others but does not cause cross-tenant leakage, an acceptable tradeoff to avoid scope creep. Added TestSkillSearchTool_TenantIsolation proving two tenants with same-slug/different-content skills never see each other's search results, including after cache population from a different tenant. * fix(tools): fix group-spec expansion in tool policy engine PolicyEngine.registry was only ever set via SetRegistry(), which was never called in production (only in one test) — so pe.registry was permanently nil in production. Every group-expansion helper (applyProfile, intersectWithSpec, unionWithSpec, subtractSpec, expandSpec, matchDenySpec, filterByCapability) silently dropped any "group:*" spec entry instead of expanding it when registry was nil. Concretely: "group:mcp" (auto-injected into agentToolPolicy.AlsoAllow for any agent with MCP tools) never resolved to real tool names, so MCP tools connected successfully and appeared in prompt text (which reads the registry directly, bypassing PolicyEngine) but were never included in the actual ChatRequest.Tools payload sent to the LLM — confirmed live via mcp.agent.tools_loaded tools=6 immediately followed by mcp.filtered_tools mcp_defs_count=0 in the same request. This affects any agent relying on group-based grants, not just MCP. PolicyEngine is a shared/global singleton used concurrently across all agents (constructed once at gateway startup), so mutating a registry field per-call would be a data race. Fix instead threads the registry as an explicit parameter from FilterTools down through all internal group-expansion helpers, and adds IsDenied/WouldAllow registry parameters, removing the dead SetRegistry() mechanism entirely. Also fixes group expansion for the per-user-MCP-tools path: FilterTools is sometimes called with a userToolOverlay wrapping a *Registry rather than a *Registry directly; added Unwrap() to userToolOverlay so the new registry-resolution logic works for both cases. Added 4 tests proving group expansion works via the threaded parameter alone (no SetRegistry): plain registry allow, userToolOverlay allow, deny-side group expansion, and the WouldAllow bridge-server path. Blast radius note: this restores intended access for every agent configured with group:* specs (group:mcp, group:vault, group:goclaw, group:coding, etc.) that were silently inert before. Existing agent configs relying on group grants will gain the tool access they were nominally already configured for. * fix(mcp): cache tool descriptions from pool-connected servers too 5ce410b6 added tool-description caching (for prompt-preview visibility) only inside connectServer. connectViaPool — the separate connect path used when MCP connections go through the shared pool — never got the same caching hook, even though it shares the same underlying connectAndDiscover wire handshake. Confirmed live: cloudflare/docker connected via connectViaPool and received real descriptions over the wire, but prompt preview still showed blank descriptions because this path never wrote to the cache. Also removes the temporary mcp.connect.raw_tool debug log added earlier this session for diagnosing the same issue — no longer needed now that the root cause is fixed. * chore(skills): remove temporary pinned-skills diagnostic logging Confirmed live: pinned skills (caveman, infra-ansible-knowledge) now resolve correctly end-to-end for tenant-scoped agents. Debug logging added to trace the resolution chain is no longer needed. * fix(tools): deny always wins over AlsoAllow group grants AlsoAllow's unionWithSpec could reintroduce a tool explicitly listed in Deny, since it added tools back from allTools without re-checking deny specs. Previously masked because AlsoAllow's group-expansion was also broken (fixed in f7af95de this session) — group specs silently expanded to nothing, so this ordering bug never manifested. Now that group expansion works, an admin-denied tool that's also reachable via a group:* AlsoAllow entry (e.g. group:mcp) would silently reappear. Re-apply deny-spec subtraction as a final step after AlsoAllow union, for both global and per-agent policy, so deny always wins regardless of which allow mechanism tries to add a tool back. Added tests proving global and per-agent Deny correctly override an overlapping AlsoAllow group grant, while sibling non-denied tools in the same group remain allowed. * fix(mcp): enumerate cached tools instead of wildcard placeholder in prompt preview ListToolsForAgent (the prompt-preview path) collapsed any server with an empty ToolAllow (unrestricted grant — the common case) into a single "server__*" placeholder entry, even when tool_cache already had every real tool name and description from connect time (5ce410b6, 8ffd67b9). This meant agents with unrestricted MCP server access never saw individual tool names or descriptions in prompt preview, forcing trial-and-error tool usage. When ToolAllow is empty and tool_cache is populated, enumerate every cached tool (skipping any explicitly denied) and emit one MCPToolPreviewInfo per tool, matching the construction logic already used for the ToolAllow-non-empty case. Falls back to the single placeholder only when tool_cache is also empty (server never connected). The live (non-preview) conversation path, buildMCPToolDescs, does not have this bug — it resolves tool identity from the live connected registry, never from ToolAllow, so no placeholder shortcut exists there. Added tests covering both the cache-populated enumeration case and the no-cache placeholder fallback. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(prompt): filter alias re-injection and add MCP schemas to preview ToolDefs Prompt-preview's tools: schema array (PreviewResult.ToolDefs) had two bugs, both preview-only — confirmed live conversations already use correctly-filtered tool payloads via buildToolsPayload/PolicyEngine.FilterTools, unaffected by either bug: 1. Alias re-injection iterated ALL registry aliases globally with no check against the deny-filtered toolNames list, letting denied tools reappear via their alias name (e.g. a denied canonical tool still showing up as its Claude-Code-compat alias like Bash/Edit/Write/Read). Now skips any alias whose canonical tool isn't in the filtered set. 2. MCP tool descriptions (from the store-based, connection-free ListToolsForAgent path) only ever fed the prompt TEXT section, never got converted into ToolDefinition schema objects — so the tools: array never showed MCP tools at all, even when the MCP section of the prompt text correctly listed them. Now appends a ToolDefinition per MCP tool with name+description populated and a placeholder {"type":"object"} Parameters schema, documented as preview-only (real parameter schemas require a live MCP connection, only available during an actual conversation turn). Added tests proving denied-tool-alias exclusion and MCP tool inclusion in preview ToolDefs. * feat(skills): inline full content for pinned skills instead of pointer-only Web UI documented "Pinned skills are always inlined in the system prompt", but BuildPinnedSummary was just BuildSummary with an allowlist filter — same as the general searchable skill list: name/description(truncated)/location pointer only, requiring use_skill+read_file round trips to get actual content. No code path inlined real SKILL.md content for pinned skills specifically. BuildPinnedSummary now reads and inlines full SKILL.md content (frontmatter stripped) per pinned skill inside <skill_instructions> tags. Per-skill (10000 bytes) and total (30000 bytes) size caps fall back to the original pointer-only format with a note when a skill is too large to inline, so oversized skills degrade gracefully instead of blowing the prompt budget. Added tests: full-content inlining, size-cap fallback, and tenant isolation for the new inline path (mirroring the existing managed- skills tenant isolation test). * fix(mcp): real parameter schemas and full policy enforcement in prompt preview Three interconnected fixes to prompt preview, none affecting the live conversation path (which was already correct): 1. Real MCP parameter schemas instead of a useless empty placeholder. tool_cache previously stored only name+description; extended to also capture the real JSON Schema from the MCP server's tools/list response (CachedToolInfo{Description, Parameters}) at connect time, for both direct-connect and pool-connect paths. Preview now shows complete, real input schemas instead of {"type":"object"} with no properties — verified against actual wire data from a live cloudflare MCP server showing genuinely rich schemas (zone_id, type, name, content, ttl, proxied, all typed with descriptions and correct required arrays) that were previously being discarded. Backward-compatible: old-shape cache entries degrade gracefully to description-only rather than crashing, self-healing on next connect. 2. Global tool deny now enforced in preview. BuildPreviewPrompt previously hand-rolled a partial policy reimplementation (per-agent deny only, explicitly skipping the full PolicyEngine "because runtime state isn't available in preview") — but PolicyEngine.WouldAllow already handles this per-tool-name without needing channel context. A tool denied via the global config (not per-agent) would appear in preview despite being correctly denied in every real conversation. Preview now calls WouldAllow per candidate tool, with a graceful per-agent-deny-only fallback when no PolicyEngine is wired (e.g. in tests). 3. MCP tools now also subject to the same policy check — previously the MCP tool supplement (store-based, connection-free tool listing) added MCP tool names unconditionally, bypassing WouldAllow entirely, so a denied MCP tool could still appear in preview. Added tests for all three: real-schema presence, global-deny exclusion for both core and MCP tools, and backward-compat cache handling. * fix(prompt): remove redundant per-tool MCP enumeration from prompt text Now that MCP tool schemas in the tools: API parameter are real and complete (1290d4f1), the ## MCP Tools prompt-text section's per-tool "- mcp_x__y: description" enumeration is pure duplication with zero added value — the model already gets each tool's real schema (including description) via tools:. buildMCPToolsInlineSection now keeps only the behavioral instructions that aren't expressible via JSON schema and thus aren't duplicated: prefer-MCP-over-core-tools guidance, and the optional-parameter guidance (don't guess/fill optional fields). The per-tool name+ description enumeration loop is removed. Section still only appears when the agent has MCP tools (len(cfg.MCPToolDescs) > 0, unchanged gate). Updated tests to assert the enumeration is gone while the behavioral instructions remain; ToolDefs assertions are now the authoritative check for MCP tool allow/deny filtering behavior (prompt text no longer enumerates names at all). * test(agent): update TeamContextInjection test for removed Team Members section TestBuildSystemPrompt_TeamContextInjection asserted the presence of a 'Team Members' prompt-text section that was intentionally removed in 71d33180 (duplicate of the canonical TEAM.md-context-file Members section, which has better formatting). The test was never updated to match, causing it to fail on every run since. Moved the assertion from wantIn to wantNotIn for the 3 affected subtests -- BuildSystemPrompt correctly no longer renders team-member roster info directly; that info now comes exclusively from the TEAM.md context-file mechanism, outside this test's isolated scope. * fix(prompt): resolve real registry for WouldAllow calls in preview BuildPreviewPrompt's two WouldAllow calls hardcoded reg=nil, silently breaking group:* expansion (e.g. group:mcp) needed to resolve the AlsoAllow grant production actually uses to grant MCP tool access (resolver_helpers.go's agentToolPolicyWithMCP injects AlsoAllow: ["group:mcp"]). With reg=nil, WouldAllow could match literal tool names fine (the 18 core/static tools) but could never resolve group-based grants, so every MCP tool silently failed WouldAllow and was excluded from preview -- confirmed live via curl: 18 tools returned, zero mcp_* ones, for an agent with genuinely working MCP access in real conversations. Live conversations were never affected -- internal/mcp/bridge_server.go's WouldAllow call already correctly passes a real registry. Fix resolves a real *tools.Registry from deps.ToolLister via tools.ResolveConcreteRegistry (the same helper used at the live call site), passing it to both WouldAllow calls instead of nil. Falls back to nil gracefully for test mocks that don't implement the full ToolExecutor interface, preserving existing test behavior. Added a test proving an MCP tool granted via the exact production AlsoAllow: ["group:mcp"] pattern is now correctly included in preview ToolDefs, where the old reg=nil bug would have silently excluded it. * fix(prompt): use literal deny check for MCP tools in preview, not group expansion The MCP-tools policy gate added in 1290d4f1 called WouldAllow with a real registry (per 9b5fd2eb), which requires group:mcp expansion against that registry to grant access via the production AlsoAllow: ["group:mcp"] pattern. But MCP tools are only ever registered into ephemeral per-agent registry clones at live connection time (manager_connect.go) -- never into the shared/global registry preview uses. group:mcp always resolved empty in preview's connection-free context, so WouldAllow denied every MCP tool -- confirmed live via tool-name diff: live conversations correctly included all 6 MCP tools, preview included zero. MCP access-granting is already correctly handled by ListToolsForAgent's own per-server tool_allow/tool_deny grant logic (confirmed working correctly earlier this session). The preview gate only needs to catch the narrower case of a literally-denied tool name via global/per-agent policy config -- it never needed group expansion. Replaced WouldAllow with IsDenied(nil, name, agentPolicy), which forces a pure literal-name match with zero registry dependency, matching the existing usage pattern already established elsewhere in policy.go. This class of bug cannot recur: there's no registry-passing code path left in this check to silently reintroduce group-expansion dependence. Added a test proving MCP tool inclusion in preview is independent of group-expansion outcome (no AlsoAllow: group:mcp needed for a non-denied tool to appear). Also reverts the temporary loop.filtered_tool_names/ preview_prompt.filtered_tool_names diagnostic logging used to capture the live-vs-preview tool-name comparison that diagnosed this bug. * fix(http): preserve real MCP parameter schemas through HTTP preview adapter mcpPreviewAdapter.ListToolsForAgent (the HTTP-layer glue converting mcp.MCPToolPreviewInfo to agent.MCPToolPreviewInfo for BuildPreviewPrompt) only copied RegisteredName and Description, silently dropping Parameters -- a bug present since this adapter was introduced (2499d0be/7e250244), unrelated to today's other MCP preview fixes. This was masked until d5fc6344 fixed MCP tools being excluded from preview entirely (a separate bug) -- once MCP tools started appearing again, this pre-existing adapter gap became visible: tools showed up correctly, but always with the bare {"type":"object"} placeholder instead of their real cached schema (confirmed live: update_dns_record missing its 7 real properties). One-line fix: copy Parameters through in the adapter's struct literal. Added a regression test constructing a real *mcp.Manager with populated tool_cache, asserting the adapter's output preserves specific real schema properties (not just non-nil Parameters) -- verified this test fails without the fix and passes with it. --------- Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
747b58d24c |
fix(usage): repair cost analytics and display precision (#1330)
fix(usage): repair cost analytics and display precision - Automatic OpenRouter pricing sync with cost backfill for traces/snapshots/events - Live usage data merging for current-hour dashboard accuracy - 2-decimal API cost formatting across usage/overview pages - Comprehensive test coverage (PG, SQLite, HTTP, UI) Merged by github-maintain automation. |
||
|
|
2499d0be78 |
fix(mcp): show configured MCP tools in prompt preview (#1321)
* fix(mcp): show configured MCP tools in prompt preview MCP tools were only visible in prompts if currently loaded in registry (from active sessions). Prompt preview showed empty MCP tools section. Solution: Query MCP store for configured servers and tool lists. Show tools that would be available, not just currently loaded. - internal/mcp/manager.go: add ListToolsForAgent() for store-based tool discovery - internal/agent/preview_prompt.go: supplement live registry with store tools - internal/http/agents.go: add mcpPreviewMgr field and setter - internal/http/agents_prompt_preview.go: create adapter bridging mcp.Manager to preview - cmd/gateway.go: wire up MCPPreviewAdapter during setup * fix(mcp): load MCP servers from database at startup, not config file Single source of truth: MCP servers are now loaded from mcp_servers table at gateway startup, instead of from the config file. This ensures that MCP servers configured via web UI are automatically available without requiring config file updates. - cmd/gateway_setup.go: remove config fallback, add initMCPFromDB() helper - cmd/gateway.go: call initMCPFromDB() after store initialization - internal/mcp/manager.go: add SetConfigs() for runtime configuration Now mcpMgr will be non-nil when MCP servers are in the database, enabling MCP tools in prompt preview. --------- Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com> |
||
|
|
8963ec3177 |
fix(ollama): always send num_ctx to prevent context-window errors (#1307)
* feat(ollama): auto-detect context window, allow user override Query Ollama API (/api/show) to fetch model's actual context_length on provider startup/hot-registration. Use queried value by default, but allow user to override via num_ctx field in provider settings. - backend: query API, store in provider config, use in requests - frontend: add num_ctx input field to Ollama provider dialog (4 langs) - fallback: 131072 tokens (128K) if API unreachable Fixes context-window errors when using large-context models. * fix(ollama): route to /api/chat endpoint to honor num_ctx Ollama's /v1/chat/completions (OpenAI-compat shim) ignores options.num_ctx. Route to /api/chat (native endpoint) which respects num_ctx. Request body is identical, only endpoint URL changes: - Before: /v1/chat/completions (ignores options.num_ctx) - After: /api/chat (honors options.num_ctx) Fixes context-window errors when using models with large context windows. * refactor(providers): replace OpenAI-wrapped Ollama with native SDK Implement dedicated OllamaProvider using github.com/ollama/ollama/api official client library. Replaces the broken approach of wrapping Ollama behind OpenAI-compatible parsing. Benefits: - Proper response parsing (Ollama native format, not OpenAI) - Automatic /v1 suffix stripping - Built-in num_ctx support (injected into every request) - Tool call support via Ollama's native ToolCall schema - Token usage mapping from Ollama Metrics - Clean streaming implementation Fixes context-window errors by using Ollama native API directly. Removes dependency on OpenAI response format hacks. - new: internal/providers/ollama.go (OllamaProvider) - updated: cmd/gateway_providers.go, internal/http/providers.go - updated: go.mod (added github.com/ollama/ollama) * fix(ollama): fetch num_ctx for config-registered providers Config-file-registered Ollama providers (cfg.Providers.Ollama.Host) were hardcoding numCtx=nil, so num_ctx was never sent to Ollama. Add FetchOllamaModelContext call for config providers, same as DB-registered providers. Now config Ollama providers also auto-detect and send the model's context window to the Ollama API. Fixes context-window errors when using config-registered Ollama. * fix(ollama): always send num_ctx to prevent context window errors OllamaProvider.buildRequest only injected options.num_ctx when an explicit override was configured. When resolveOllamaNumCtx saw that /api/show returned the OllamaDefaultNumCtx value (131072), it returned nil, causing the provider to omit num_ctx entirely. Ollama then fell back to the model's baked-in default (4096 tokens), causing requests with long system prompts to fail with "request exceeds context size". Fix both layers: - resolveOllamaNumCtx now always returns &numCtx instead of nil when /api/show returns the default, so the explicit value is always passed. - OllamaProvider.buildRequest falls back to OllamaDefaultNumCtx (131072) when p.numCtx is nil, ensuring every request carries a large context window even if the startup query was skipped. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
969a5879ae |
fix: unify no reply detection on dev (#1303)
Co-authored-by: GoClaw Operator <operator@goclaw> |
||
|
|
1c53447b99 | feat(gateway): add provider-backed llm rpc | ||
|
|
4a79c8a208 |
feat(cron): deterministic command payloads (run a shell command, no LLM) (#1279)
* feat(cron): deterministic command payloads (run a shell command, no LLM) Cron jobs always run an agent turn today, so deterministic work (health probes, backups, syncs) pays model tokens on every fire. This adds a "command" payload kind that runs a shell command directly in the gateway process with zero model tokens, mirroring openclaw's command cron. - store: CronPayload.Command (*CronCommandSpec — argv/cwd/env/input/ timeouts/output cap). Persists in the existing payload JSON blob, so there is NO migration and no schema version bump. - internal/cronexec: in-process runner with wall-clock + no-output timeouts, per-stream output capping, and process-group termination so a timed-out command's forked children are also killed. - gateway_cron handler: command jobs run in-process and deliver stdout on success (honoring the NO_REPLY sentinel). A non-zero exit / timeout returns an error so the run is recorded as error and retried per cron.max_retries; failures are NOT delivered, mirroring the agent path (only successful output is announced — no channel spam). - surfaces: cron.create RPC, the agent `cron` tool, and a new `goclaw cron create` CLI all accept command payloads. - security: gated by cron.command_enabled (default false). Commands run with the gateway process's privileges, so the feature is opt-in per gateway; when disabled the RPC and tool reject command payloads and the handler refuses to run them. - i18n (en/vi/zh), docs (08-scheduling-cron.md), and tests for the runner and the handler command path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(cron): gate command payloads on the update surfaces too handleUpdate (RPC + agent tool) passed CronJobPatch.Command straight to UpdateJob, which switches the payload to command kind for any non-nil Command — without the command_enabled gate or ValidateCronCommandSpec that create enforces. A normal job could therefore be mutated into a command job (or persisted with an invalid spec, e.g. empty argv) on a gateway where command cron is disabled, breaking the disabled-gateway contract. Both update surfaces now require cron.command_enabled and validate the spec before UpdateJob, matching create. The agent tool parses the command via the same path as add and drops the raw keys so a shell-string command can't break the generic patch unmarshal. Regression tests added for RPC and tool update (command disabled + invalid argv), plus a positive enabled-valid case. Addresses review feedback from @mrgoonie on #1279. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ce0472b580 |
fix(gateway): re-apply allowed_paths to filesystem tools after system_configs overlay (#1274)
setupToolRegistry wires the filesystem tools' AllowPaths from the config/JSON5 default before ApplySystemConfigs overlays system_configs['allowed_paths'], so DB-configured allowed paths never reached read_file / list_files / write_file / edit / send_file. Only the rate limiter was re-applied after the overlay (#1111); the AllowPaths analogue was missing, so agents were denied access to configured shared directories outside their workspace even though the DB value was present (visible as a read_file "access denied" log whose allowedPrefixes omit the configured path). Extract the user-allowed-path application into applyUserAllowedPaths and re-run it from runGateway after the overlay, mirroring the rate-limiter re-apply. The helper is idempotent (AllowPaths is additive and the prefix check is membership-based), so the initial wiring call plus the re-apply is safe. Test: cmd/gateway_tools_wiring_test.go asserts a path outside the workspace is denied before the grant and readable after, that an unrelated path stays denied, that repeated application is safe, and that an empty list is a no-op. |
||
|
|
505b7a9d99 |
fix(agent): recover interrupted runs to failed on gateway startup (#1272)
A run records its terminal run.status (completed/failed/cancelled) only when the agent loop returns. If the gateway process is killed mid-run (restart, crash, or SIGKILL after a slow graceful shutdown), the terminal event is never emitted, so the run's last run.status item stays "started" forever — it shows as perpetually running in the timeline and is never counted as failed. Mirror the cron scheduler's startup reset (recomputeStaleJobs flips stale 'running' jobs to 'interrupted'): add RunTimelineStore.RecoverInterruptedRuns, invoked once on gateway startup. It finds runs that have a started run.status but no terminal sibling and appends a terminal failed run.status item, marked interrupted in metadata so it stays distinguishable from a genuine agent failure. Running only at startup means nothing from before is still executing, so there is no false-positive risk — unlike a periodic sweep (see the deliberately-disabled tracing stale-trace loop). - store: RecoverInterruptedRuns on the interface + PG and SQLite impls - cmd/gateway_managed: invoke once after the timeline recorder is built - tests: SQLite store test — orphan recovered, completed run untouched, idempotent on re-run |
||
|
|
e2ec3370dd |
feat(webhooks): stream provider responses for server-side runs to enable prompt caching (#1273)
Server-side webhook agent runs (sync, async, and admin test) used Stream:false, so OpenAI-compatible routers that only cache streaming requests never populated or served their prompt cache. Webhook runs paid full input-token price on every turn even with a stable session and an identical multi-turn prefix, while WS chat (Stream:true) got cache hits from the 3rd turn. Add gateway.webhook_stream (default true, env GOCLAW_WEBHOOK_STREAM) and apply it to all three run sites. ChatStream returns the fully assembled response, so the payload returned to the caller is unchanged. Set to false to restore non-streaming. |
||
|
|
a10464290d |
feat(webhooks): configurable agent run timeout (default 600s) (#1267)
Replace the hardcoded 30s webhook agent-run deadline with a configurable
timeout (default 600s, cap 3600s) for both the async worker and the
sync/test HTTP handler. Legitimate multi-step runs were dying at 30s
mid-tool-call; the same request over WebSocket completed fine.
- internal/webhooks/timeout.go: ResolveTimeoutSec helper (<=0 → 600s, cap 3600s)
- WorkerConfig.AsyncAgentTimeout (async worker) + WebhookLLMHandler.syncTimeout
(sync + admin test), wired from config
- config keys gateway.webhook_{async,sync}_timeout_sec, also settable via
GOCLAW_WEBHOOK_{ASYNC,SYNC}_TIMEOUT_SEC env (env overrides config)
- tests: timeout bounds + config file/env override
|
||
|
|
c0531e1e07 |
fix(cron): reset session for stateless cron jobs, not stateful ones (#1262)
The per-run session reset was gated on `!job.Stateless`, inverting the flag: stateless jobs (the token-saving default) skipped the reset and accumulated unbounded history, while stateful jobs were wiped every run. Reset for `job.Stateless` instead, and clear BOTH session layers — the goclaw session store AND the Claude CLI on-disk .jsonl. claude-cli resumes its own session by a deterministic per-key UUID, so without clearing the .jsonl a "stateless" run still replayed the entire accumulated history (and silently grew it run after run). |
||
|
|
bd5adc61c8 |
feat(bitrix24): imbot.v2 migration, 2-way media, openline sender-tag echo, and hardening (#1236)
* refactor(bitrix24): rename "Path B" framing to maintainer-specified naming [B24:2794] Per maintainer hard rule #10 (no generic "Path A/B" framing) from PR #1061 review. The Bitrix24 MCP auto-onboard flow is Bitrix-specific glue ("Bitrix24 OAuth -> existing mcp_user_credentials bridge"), NOT a generic MCP architecture pattern. Naming convention applied consistently: - First mention per file: full "Bitrix24 OAuth -> existing mcp_user_credentials bridge" (matches maintainer comment verbatim). - Subsequent mentions in same file: shortened "mcp_user_credentials bridge". - Test/log context referencing literal endpoint /api/auto-onboard: keep "auto-onboard" reference (it's the actual API endpoint name). Changes are documentation-only: - Rename in code comments + test descriptions + plan docs. - Clarify framing in mcp_client.go + provisioner.go doc comments to emphasize Bitrix-specific glue (not generic MCP infra). - Reuse existing mcp_user_credentials table + MCPServerStore methods (no schema / store / abstraction change). Files: - cmd/gateway.go (factory registration doc) - internal/channels/bitrix24/{channel,factory,mcp_client,provisioner}.go - internal/channels/bitrix24/{mcp_client,provisioner}_test.go - plan/goclaw-mcp-integration.md (21 occurrences) Verified: go build + MCP-related tests pass (TestProvision*, TestInitMCPProvisioner*, TestMCPClient*). Phase 1 of Path C execution per plans/reports/decision-log-260519-1555-bitrix24-pr-fork-decision.md. * fix: confine outbound media paths to agent workspace [B24:2794] Tool MEDIA:<path> output reached channel file-upload sinks (Bitrix imbot.v2.File.upload, Telegram sendDocument, etc.) verbatim via parseMediaResult, with no workspace-boundary check. A malicious or buggy tool emitting MEDIA:/etc/passwd could exfiltrate arbitrary files to chat. Extract the EvalSymlinks+Rel containment from extractMediaFromContent into a shared confineToWorkspace helper and apply it at the parseMediaResult sink in processToolResult. Fixing at the source/egress boundary protects every channel at once rather than per-channel. Paths that escape the workspace are dropped and logged (security.media_path_rejected). Add TestConfineToWorkspace (boundary unit) and TestParseMediaResultConfinedToWorkspace (sink regression for H2). * feat(bitrix24): support inbound + outbound media via imbot.v2 File API [B24:2794] Bitrix24 channel was text-only; attachments were parsed but dropped. - Inbound: download chat files via imbot.v2.File.download (one-time URL), forward to the agent with MIME preserved (internal/channels/bitrix24/download.go). - Outbound: upload agent media to the chat via imbot.v2.File.upload (internal/channels/bitrix24/send_media.go). - Add BaseChannel.HandleMessageMedia to preserve MIME/filename through the bus. - Per-channel media_max_mb cap (default 20) applies to both directions. Tests: 92 pass (internal/channels/bitrix24 + internal/channels), go vet clean (PG + sqliteonly). * refactor(bitrix24): migrate messaging/bot-list/unregister to imbot v2 API [B24:2794] Move outbound REST calls to the imbot v2 family (keeps register on v1): - imbot.message.add -> imbot.v2.Chat.Message.send (fields.message shape, live-verified) - imbot.bot.list (+ legacy imbot.list fallback) -> imbot.v2.Bot.list; add botListRows to normalize the v2 {bots:[...]} envelope, legacy array, and id-keyed map forms - imbot.unregister -> imbot.v2.Bot.unregister Bot registration stays on v1 imbot.register: v2 imbot.v2.Bot.register changes the event-delivery model (per-event handler URLs -> eventMode), which would require rewriting the inbound event parser. No user-facing behavior change. Tests: bitrix24 package green; go vet ./... clean. * feat(bitrix24): route whisper via v1 SKIP_CONNECTOR + add v2 replyId [B24:2794] Bot was leaking HiddenMessage (whisper) replies to the external Zalo connector because every outbound call went through imbot.v2.Chat.Message.send, which has no equivalent of the v1 SKIP_CONNECTOR flag. Branch the outbound path on inbound visibility: whisper → imbot.message.add + SKIP_CONNECTOR=Y (v1, send_v1.go) public → imbot.v2.Chat.Message.send + fields.replyId (v2, send_v2.go) Pipeline: events.go parse data[PARAMS][PARAMS][COMPONENT_ID]=HiddenMessage into EventParams.IsHiddenMessage (form + JSON variants) handle.go set bitrix_visibility on InboundMessage.Metadata consumer forward visibility + message_id into OutboundMessage send.go resolveSendOptions + sendChunk dispatcher + shared callWithRateLimitRetry helper metadata_keys.go single source of truth for the keys + values Defaults preserve pre-refactor behaviour: callers that don't populate bitrix_visibility still go through v2 public, and replyId is omitted unless a numeric bitrix_message_id arrives in metadata. Tests: TestParseEvent_FormURLEncoded_IsHiddenMessage (3 cases) TestParseEvent_JSON_IsHiddenMessage (3 cases) TestResolveSendOptions (8 cases) TestSend_BranchesOnVisibility (4 cases) * feat(bitrix24): openline sender-tag echo on replies [B24:2794] Openline sender-tag echo (this change): - Capture the connector sender tag ("[name #id]:" or "[name] #id:") from inbound openline group messages, strip it from the body the agent sees, and re-prepend the canonical "[name] #id:" form to the reply so the Open Channel connector routes the answer back to the right external user. - New sender_prefix.go helper (+ test) accepts both inbound layouts and emits one canonical form; scoped to messages carrying the tag, so plain chats are unaffected. - metadata_keys.go: MetaKeySenderPrefix; handle.go capture/strip/stash; gateway_consumer_normal.go forwards the key; send.go prepends it on the first chunk before chunking. Bundled bitrix24 channel-core work already on this branch: - handle.go: @mention is the sole trigger for both staff and connector customers; unmentioned traffic is dropped (was: drop all connector msgs). - isGroupMessageType: treat SONET_GROUP "B" as a group. - handle_test.go, mcp_client_test.go: cover the above. * feat(bitrix24): accept colon-less openline sender tag, echo [name] #id [B24:2794] The Open Channel connector dropped the trailing colon from its sender tag: inbound now arrives as "[Name] #id <msg>" (was "[Name] #id: <msg>"). The id-bearing patterns required the colon, so the tag fell through to the name-only branch and the reply echoed "[Name]" — dropping the #id the connector needs to route the answer back. - sender_prefix.go: make the trailing ":" optional on both id layouts ([name #id] / [name] #id, with or without colon) and echo the canonical "[name] #id" (no colon) to match the connector's current format. Bare "[name]" (no id) still echoes "[name]" for Open Channel only. - handle.go: gate the bare name-only layout to Open Channel (isOpenChannel) so ordinary group chats starting with "[x] ..." are left untouched. - sender_prefix_test.go: cover colon/no-colon x id-inside/id-outside, the name-only openline case, and the non-openline no-op. * fix: security and robustness fixes from the bitrix24 channel review [B24:2794] - download.go: block redirect-based SSRF on inbound media. CheckRedirect re-validates each hop (http(s) only, reject private/loopback/link-local hosts, cap hops); the initial portal-domain pin is no longer bypassable via a 3xx to an internal service. Public-host redirects still allowed. - handle.go: extract/echo the openline sender tag only for Open Channel sessions (was: any group chat), removing bogus prefixes in CRM group chats and narrowing the forged-tag misroute surface. - loop_tools.go + loop_media.go: confine result.Media to the agent / team / tenant-allowed roots (new confineToAnyRoot) before a channel uploads it, so a prompt-injected out-of-workspace path (e.g. /etc/passwd) cannot exfiltrate, while legitimate cross-workspace media (team files, delegatee output) still flows. - send_media.go: bounded outbound read via io.LimitReader replaces the os.Stat + os.ReadFile pair, closing the TOCTOU size-cap bypass; cap a single message's outbound attachments at 10 (mirrors inbound). - register.go: paginate imbot.v2.Bot.list (limit/offset + hasNextPage, capped at 40 pages) so verify/lookup see bots past the first 50. - mcp_client.go: redact access_token / refresh_token / client_secret from an echoed MCP error body before it is logged or returned (+ test). * fix(security): validate resolved dial IP on Bitrix media redirects [B24:2794] The inbound media download redirect guard only string-checked the redirect hostname (isPrivateOrLoopback on req.URL.Hostname()), so a redirect to a public hostname that resolves to 127.0.0.1 / 169.254.169.254 / an RFC1918 address — or a DNS-rebinding swap between check and dial — still passed the guard and the client would connect. Reported in PR review. Add security.NewRedirectFollowingSafeClient: it follows redirects but validates the RESOLVED destination IP of every hop at dial time via net.Dialer.Control, reusing the existing blocked-CIDR list. The IP it checks is the IP actually dialed, so both redirect-to-internal and DNS rebinding are refused, while legitimate public CDN redirects still succeed. download.go now uses it instead of the hostname-string guard. Tests: deterministic dial-control table (loopback / link-local / private / multicast / unspecified / public, v4 + v6), malformed/non-IP addr, test bypass, loopback-dial-blocked client wiring, and redirect cap + scheme checks. * feat(bitrix24): per-participant Zalo openline identity from 3-token sender tag [B24:2794] Parse the connector's "[Name] #uid #msgId" sender tag so each external customer in a shared Open Channel group gets its own contact + USER.md instead of collapsing onto the connector proxy id. Identity minting is gated on IS_CONNECTOR=Y to reject operator forged tags. Echo back the msgId only ("#msgId") on replies; keep the legacy single-number and name-only layouts unchanged. Zero DB migration. - sender_prefix.go: parseOpenlineSenderTag() classifies 3-token / legacy / name-only - handle.go: synthetic senderID "openlines:{instance}:{chat}:{uid}" + participant_user_id metadata, gated on FromIsConnector - gateway_consumer_normal.go: deriveGroupUserID() routes participant -> per-person scope, group fallback otherwise - send.go: buildAddressMention numeric-id guard so synthetic ids don't emit invalid [USER=...] BBCode - MetaKeyMessageID kept as Bitrix MESSAGE_ID (drives v2 fields.replyId); connector msgId surfaced only via echo prefix --------- Co-authored-by: DangTinh311 <dangtinh31193@gmail.com> Co-authored-by: Chinh Dang <chinhdang@192.168.68.104> |
||
|
|
c24677272c |
Fix/whatsapp connect (#1254)
* fix: recover whatsapp qr after deleted device * fix: whatsapp connect * fix whatsapp instance device scoping --------- Co-authored-by: Duy /zuey/ <duy@wearetopgroup.com> |
||
|
|
0ae55991bb |
feat(mcp): MCP OAuth 2.1 client for tool servers (#1196)
* feat(mcp): MCP OAuth 2.1 client — full implementation with tests
Implements a complete MCP OAuth 2.1 authorization flow for tool servers that
require user-delegated access, covering all layers from DB to UI.
- discovery.go: RFC 9728 protected-resource → RFC 8414 AS metadata → OIDC
fallback chain with 5-min in-memory cache and InvalidateCache()
- dcr.go: RFC 7591 Dynamic Client Registration with response size guard
- flow.go: PKCE (S256) authorization code flow — StartFlow(), ExchangeCode(),
ClientCredentials(), auto-cleanup of expired flows; carries AS issuer through
PendingFlow for status display
- refresher.go: OAuthTokenProvider with in-memory token cache, automatic refresh
on expiry, per-user vs global slot isolation, InvalidateCache/InvalidateServer
- migrations/000074 + SQLite schema: mcp_oauth_tokens with AES-256-GCM encrypted
access/refresh tokens, partial unique index for global vs per-user rows,
ON DELETE CASCADE from mcp_servers
- store.MCPOAuthTokenStore: Upsert, Get/GetUser, Delete/DeleteUser, and
DeleteServerOAuthTokens (purge all rows for a server)
- PostgreSQL + SQLite implementations
- POST /v1/mcp/oauth/start — discovery + optional DCR + PKCE redirect URL;
client_credentials completes server-side (no redirect) and returns completed=true
- GET /v1/mcp/oauth/callback — exchange code, persist token, publish WS event;
payload built via json.Marshal (no reflected XSS via error_description)
- GET /v1/mcp/oauth/status/{id}, DELETE /v1/mcp/oauth/token/{id} — admin-gated
- POST /v1/mcp/oauth/discover/{id} — on-demand discovery probe
- All outbound calls go through the SSRF-safe client with pinned IPs
- pkg/protocol/mcp_events.go: EventMCPOAuthComplete routed only to the initiating
user (admins in-tenant included); fail-closed across tenants
- getUserMCPTools() injects Authorization: Bearer from OAuthTokenProvider; on a
401 for OAuth servers it purges the cached token so the next turn re-resolves
- handleUpdateServer purges all OAuth tokens (global + per-user), drops the
refresher cache, and evicts the pool when a server's URL or OAuth config
(client_id / endpoints / grant_type / scope / auth_type) changes — so the
status UI and agent never use a token minted for the old resource/AS
- MCPOAuthDialog (WS-driven), unified user-credentials dialog, OAuth settings
fields; handles the no-redirect client_credentials completion
- internal/mcp/oauth/*_test.go: discovery cache, PKCE, DCR, refresher
- internal/http/mcp_oauth_test.go + mcp_update_oauth_purge_test.go: routes, auth
gating, WS event, purge-on-URL/OAuth-config-change
- tests/integration: store + encryption + tenant isolation, E2E start→callback,
DeleteServerOAuthTokens
- internal/gateway/event_filter_test.go, internal/agent/loop_mcp_user_test.go
* fix(mcp): return 400 on OAuth callback with code but missing state
The callback handler rendered a 200 HTML page whenever code or state was
absent. An auth code WITH a missing state is a malformed / CSRF-risk
callback (state is the CSRF token), so reject that case with HTTP 400.
A bare hit with neither code nor state (user opening the URL directly),
provider errors, and exchange failures keep their 200 HTML popup page.
Adds a status code parameter to writeCallbackHTML. Fixes the
TestOAuthCallbackMissingState integration regression while keeping
TestHandleCallbackMissingCodeAndState (no params -> 200) green.
* fix(mcp): scope-based OAuth auth + honor manual OAuth endpoints
Addresses the two MCP/OAuth security-review findings.
Finding 1 — authorization. mcp_oauth_tokens is tenant-scoped, but
start/status/revoke were gated only by requireAuth(RoleAdmin), an RBAC
role check, not tenant membership, so a RoleAdmin caller could act on a
tenant they don't administer. A blanket requireTenantAdmin would have
broken per-user self-service, which the UI exposes (the per-user
MCPUserCredentialsDialog shows an "Authorize" button to regular users for
their own credentials). Instead mirror the existing per-user MCP
credentials model (resolveTargetUserID in mcp_user_credentials.go):
- start/status/revoke accept any authenticated user; each handler calls
authorizeOAuthScope.
- a caller may manage their OWN per-user token (self-service); the
global/server token (user_id="") and other users' tokens require
tenant-admin (owner bypass), so a RoleAdmin that is not a tenant admin
is rejected.
- discover stays admin-only (it only previews AS metadata for a server).
Add a TenantStore dependency. Tests cover self-service, on-behalf-of-
another (403), and global-by-non-tenant-admin (403).
Finding 2 — honor manual OAuth config end-to-end. The UI sent use_dcr /
auth_endpoint / token_endpoint and the update path fingerprinted them for
purge, but handleStart always discovered + DCR'd and ignored them. Now:
- use_dcr=false (a *bool, so legacy/absent stays discover+DCR) skips
discovery/registration and uses the operator endpoints, SSRF-validated.
- token_endpoint is always required; auth_endpoint only for auth-code
grants — client_credentials needs no authorization URL, matching the UI
which hides that field for that grant.
- the refresher already refreshes against the stored token_endpoint and
the callback persists it, so manual-mode tokens refresh correctly.
- oauthFingerprint includes use_dcr (nil normalized to true) so toggling
DCR mode purges stale tokens.
- the web form only serializes manual endpoints when use_dcr is off.
Audited all MCP dialogs (form, global OAuth, per-user credentials, grants,
tools): OAuth dialogs handle completed/auth_url identically and read
config from stored server settings; runtime connect uses the stored token
via the refresher (no re-discovery).
Tests: manual auth-code + client_credentials endpoints, missing/SSRF
endpoints, and the full self/global/on-behalf authorization matrix.
|
||
|
|
21fb1f18ee |
feat(webhooks): add webhook management UI with delivery history & test (#1211)
* feat(webhooks): add webhook management UI with delivery history & test
Webhooks admin page on the web dashboard for managing inbound HTTP webhooks
(llm + message kinds), backed by new admin endpoints. No schema changes
(reuses migrations 000059-000061).
Backend (internal/http):
- GET /v1/webhooks/{id}/calls - paginated delivery history (trimmed
DTO; status/limit/offset filters; ownership + tenant scoped)
- GET /v1/webhooks/{id}/calls/{callId} - full single-call detail (request
payload, full response, callback URL, idempotency key, timestamps)
- POST /v1/webhooks/{id}/test - server-side test invocation using the
admin session (no secret); dispatched by kind via RunTest() on the llm/message
handlers; message tester nil-guarded (403 on Lite)
Injects WebhookCallStore + SetTesters() in cmd/gateway_http_wiring.go, adds i18n
key webhook.message_test_requires_standard (en/vi/zh), and unit tests for
list/detail (filter, tenant isolation) and test (success/error/edition gate).
Frontend (ui/web):
- /webhooks admin-only page: list with revoked filter + badge, create/edit form
(edition-gated message kind, Lite localhost_only lock), show-once secret dialog
(create + rotate), test dialog, and delivery-history dialog with server-side
pagination plus a click-through full call-detail dialog.
- use-webhooks hooks, query keys, types, routing, sidebar entry (Cable icon to
distinguish from the existing event "Hooks" page), webhooks i18n namespace
(en/vi/zh).
* fix(webhooks): validate admin message tests
---------
Co-authored-by: Goon <duy@wearetopgroup.com>
|
||
|
|
db98e477d4 |
fix(gateway): re-apply tool rate limiter after system_configs overlay (#1111)
setupToolRegistry creates the rate limiter from cfg.Tools.RateLimitPerHour during early bootstrap (gateway.go line 137). System_configs DB overlay runs ~50 lines later via cfg.ApplySystemConfigs (line 191), so any DB override of tools.rate_limit_per_hour was silently lost - the limiter object was already initialised from the JSON5 default. Symptom in production: editing tools.rate_limit_per_hour via system_configs table or the config HTTP API had no effect on running gateways. Operators had to inject a config.json file to change the value, defeating the DB-as-source-of-truth pattern that other tunables rely on. Re-apply the limiter after ApplySystemConfigs runs. Safe ordering: server has not started yet, no in-flight tool calls, and SetRateLimiter is a plain field assignment with no shutdown cost on the discarded limiter. nil case (rate_limit_per_hour <= 0) also handled so DB writes can disable the limiter without restart. Tests: new TestRegistry_SetRateLimiter_ReplacesPriorLimiter covers both the replace-with-higher-limit path and the nil-disables path. All existing rate limiter tests still pass. Note: the same ordering pattern likely affects other config consumed inside setupToolRegistry (tools.scrub_credentials, MCP server wiring). Out of scope for this hotfix - they need a deeper restructure. |
||
|
|
b5d5fce6e4 |
feat: update MiniMax and Z.AI provider defaults
Refresh MiniMax to MiniMax-M3, update Z.AI defaults to glm-5.2, and add focused provider catalog/runtime coverage. |
||
|
|
bd52dbaf6f |
fix(cli): use openai_compat instead of openai-compat in provider setup (#1062)
The CLI setup wizard and 'providers add' command were using 'openai-compat' (hyphen), but the API and database expect 'openai_compat' (underscore), causing provider creation to fail. Fixes #1046 |
||
|
|
2c41e0907e |
fix(announce): propagate sender + role through team-task announce re-ingress (#1128)
When a member calls team_tasks(action="complete"), goclaw fans the result
back to the Lead's session via the team-task announce queue. The Lead
resumes a turn whose initial user message is the synthesized
"[System Message] Team member ... completed task. Result: ..." string.
Inside that resumed turn, the Lead is a regular agent — it can call any
tool, including write_file and cron mutations. Those tools route through
CheckFileWriterPermission / CheckCronPermission, which in group-scope
sessions deny when the resumed RunRequest carries no SenderID:
permission denied: system context cannot write files in group chats.
If this is a legitimate user action, ensure the acting sender is
preserved through the tool chain.
team_tool_dispatch.go already stamps MetaOriginSenderID and MetaOriginRole
into the dispatch metadata. consumer_handlers.go already reads inMeta on
the completion-side teammate message. The gap was in between:
- announceRouting (cmd/gateway_announce_queue.go) had no field for
OriginSenderID / OriginRole
- The RunRequest it built had no SenderID / Role set
- loop_context.injectContext skips WithSenderID when req.SenderID is
empty — so the Lead's resumed ctx had no sender attribution, and
every group-scope permission check then tripped the deny path.
Subagent path (subagentAnnounceRouting in gateway_subagent_announce_queue.go)
already had these fields wired since #915. This brings the team-task
path to parity. No security relaxation: empty upstream → still empty
downstream (still denies, as intended for genuine system-initiated
turns); only legitimately-attributed dispatches now flow through.
- cmd/gateway_announce_queue.go: add OriginSenderID + OriginRole to
announceRouting; pass them into the RunRequest.
- cmd/gateway_consumer_handlers.go: read MetaOriginSenderID +
MetaOriginRole from inMeta when building the routing struct.
- cmd/gateway_announce_routing_test.go: 2 unit tests guarding both
"real human propagates through" and "empty stays empty" cases so
this regression can't re-land silently.
Tests: existing cmd/ + internal/tools/ suites pass; new tests pass.
|