30 Commits
Author SHA1 Message Date
Duc Nguyen e3dad3c60c fix(tts): unify voice resolution — dashboard settings affect tts tool (#1175)
The dashboard /tts page writes to system_configs[tts.<provider>.voice],
but the LLM-invoked tts tool was only checking args > agent OtherConfig
> builtin_tool_tenant_configs[tts].default_voice_id. Two different
storage locations → the user's chosen voice was ignored, Edge defaulted
to en-US-AriaNeural even when "HoaiMy" (vi-VN-HoaiMyNeural) was set
correctly in the dashboard.

Add system_configs as a 4th-level fallback so the dashboard becomes
the single source of truth.

- TtsTool gains SetSystemConfigStore(s) setter
- resolveVoiceAndModel takes providerName + looks up
  tts.<provider>.voice/model when no higher-precedence source set
- effectiveProvider resolved BEFORE voice/model so the right key is hit
- Wired at gateway boot in cmd/gateway.go (after pgStores ready)
- 4 unit tests covering: fallback, arg precedence, no-store, empty provider

Verified against production trace 019e6036-44bb-703b-85fa-dee34f7ab2c0
where the tts tool was called with provider=edge, no voice arg, and
defaulted to English instead of the configured Vietnamese voice.
2026-07-18 00:22:48 +07:00
Duc Nguyen 304ce72299 fix(tests): resolve integration test compile errors (#939)
* fix(tests): resolve integration test compile errors

- Remove duplicate allowLoopbackForTest in hooks_pipeline_test.go (canonical version lives in v3_test_helper.go)
- Remove unused fakeClient assignment in mcp_grant_revoke_test.go; fakeMCPClient type retained for future use

* ci: bump unit test timeout from 90s to 5m

The internal/hooks/handlers package binary under `-race -coverpkg=./...`
now runs up against the 90s cap because HTTPHandler retry uses a real
`time.After(1 * time.Second)` backoff across three HTTP test cases, and
goja-based memory-bomb sandbox tests have large allocations that run
inline before the sandbox deadline kicks in.

Recent main CI has been red with this timeout firing on different slow
tests each run (TestHTTP_5xxRetriesOnce, TestCorpus_MemoryBombString).
Bumping to 5m keeps the deadlock safety net (still half the 10-minute
Go default) while giving slow-but-non-deadlocked packages room.

Followup: make HTTPHandler backoff configurable so tests can override
with ms-scale delays and the 90s cap can come back.

Also update `make test` in Makefile to match.

* test(mcp): skip RevokeUserGrant test pending Phase 02 implementation

Commit 8b8da3a3 added this test alongside the grant-checker, but the
user-grant revocation semantics were never implemented: ListAccessible's
SQL treats an absent mcp_user_grants row as "allowed by default"
(WHERE mug.id IS NULL OR mug.enabled = true), so RevokeFromUser's DELETE
leaves the server accessible.

The test never actually ran in CI until the prior compile fix in this PR
unblocked the integration test binary. It's safe to skip — the agent-grant
counterpart test still exercises the grant-recheck code path. Re-enable
once user-grant-required semantics land.
2026-04-17 19:55:56 +07:00
Duc Nguyen 92dbb12c84 fix(tests): resolve integration test compile errors
- Remove duplicate allowLoopbackForTest in hooks_pipeline_test.go (canonical version lives in v3_test_helper.go)
- Remove unused fakeClient assignment in mcp_grant_revoke_test.go; fakeMCPClient type retained for future use
2026-04-17 19:02:53 +07:00
Duc Nguyen 52f48e5ea6 fix(backup): detect pg server major for version-aware pg_dump hints (#829) (#830)
* fix(backup): detect pg server major for version-aware pg_dump hints

pg_dump aborts when its major version is older than the server's, so
backup preflight now runs SHOW server_version_num against the live PG
server and uses the detected major to drive:

- the missing-pg_dump hint (names the exact postgresqlNN-client)
- a new compat check that flags an installed-but-too-old pg_dump as
  not-ready, instead of letting backup fail partway through the dump

Adds ParsePgDumpMajor helper for Debian/Homebrew/EDB output shapes,
covered by table-driven unit tests. Normalizes nil ctx once at
RunPreflight boundary so downstream exec.CommandContext calls are safe.
Stubs detectPGServerMajor + checkPgDumpServerCompat for the sqliteonly
build.

UI: removes the duplicate static hint from the preflight alert box so
the dynamic, actionable warning is the single source of truth. Drops
the obsolete pgDumpHint i18n key from en/vi/zh locale files.

Refs #829

* fix(docker): bump runtime base alpine 3.22 → 3.23

Alpine 3.23 main ships postgresql16-client, postgresql17-client, and
postgresql18-client simultaneously. Alpine 3.22 only shipped up to
postgresql17-client, so the backup preflight's dynamic hint to install
postgresql18-client previously failed with "no such package" on PG 18
deployments.

No client package is pre-bundled: the on-demand install via the
Packages page now resolves for any supported PG major.

Refs #829
2026-04-11 21:22:04 +07:00
Duc Nguyen 329b150902 fix(desktop): add defaultValues to form dialogs to prevent crash (#737)
useForm() without defaultValues causes watch() to return undefined on first render, crashing with undefined.trim(). Adds defaultValues to all 5 affected dialogs: Agent, Channel, MCP, Provider, TeamSettings.

Closes #732, closes #735
2026-04-08 15:26:33 +07:00
Duc Nguyen 0db1e93abf feat(whatsapp): add native WhatsApp channel with whatsmeow (#720)
Replace Node.js Baileys bridge with native go.mau.fi/whatsmeow — zero
external dependencies. QR auth, media support, markdown formatting,
typing indicators, dual JID/LID identity, group policies, pairing.

Resolves #703
2026-04-07 12:12:44 +07:00
Duc Nguyen 3039ce93e6 fix(scheduler): drain queued runs in DropOldPolicy test to prevent CI hang (#678)
The test closed blockCh then immediately exited, racing defer sched.Stop()
against scheduleNext(). Under -race on CI, Stop() could call wg.Wait()
while a goroutine was still being submitted, causing a 600s timeout hang.

Fix: wait for queued runs to complete before allowing Stop() to run.

Closes #677
2026-04-05 09:00:30 +07:00
Duc Nguyen e85545dc1b fix(gateway): add session ownership checks to chat.* WS methods (#676)
chat.history, chat.inject, chat.abort, and chat.session.status accepted
any sessionKey without verifying the caller owns the session. A non-admin
user could read, write, or disrupt another user's conversations by
supplying their sessionKey.

Apply the same requireSessionOwner() guard already used by sessions.*
methods: canSeeAll() bypass for admin/owner, sess.UserID match for
regular users. Extracted shared helper to access.go to reduce duplication.

Also fixes: handleSessionStatus i18n compliance (was hardcoded English),
and closes runId-only abort gap (non-admin must provide sessionKey).
2026-04-05 07:25:55 +07:00
Duc Nguyen 2b1180f59d fix(http): allow localhost URLs for local provider types (Ollama, Claude CLI) (#674)
SSRF validation in validateProviderURL() blocked all localhost/loopback
addresses, preventing local providers like Ollama from being configured
with http://localhost:11434/v1. Introduce localProviderTypes map to skip
SSRF checks for inherently local provider types (ollama, claude_cli, acp).

Closes #673
2026-04-04 07:06:34 +07:00
Duc Nguyen 532aab4387 fix(docker): pin Python/npm dependency versions (#663)
- Add docker/requirements-base.txt (edge-tts) and requirements-skills.txt (10 packages)
- Pin versions: pip ~= (patch only), npm ^, anthropic >=0.88<2.0
- Add .npmrc with supportedArchitectures for Alpine musl builds
- Restore --frozen-lockfile (lockfile regenerated with musl entries)
- Pin @anthropic-ai/claude-code@^1.0.47

Closes #662
2026-04-03 14:55:07 +07:00
Duc Nguyen 983f6184d9 fix(ui): dynamic searchable timezone picker with validation (#614)
Replace hardcoded 20-entry IANA_TIMEZONES with getAllIanaTimezones()
using Intl.supportedValuesOf (~400 zones). Switch Select dropdowns
to searchable Combobox in cron, heartbeat, and system config.

Add defense-in-depth timezone validation:
- Backend: validate in heartbeat.set handler and SetDefaultTimezone()
- Frontend: isValidIanaTimezone() guard before save in all 3 dialogs

Closes #614
2026-04-01 11:12:35 +07:00
Duc Nguyen 8b5a05a4b7 fix(web): sign media_refs paths in session history delivery (#519) (#520)
File downloads in web chat returned 401 when loading session history
because MediaRefs paths were not signed with ?ft= HMAC tokens.

- Add SignMediaPath() to clean legacy corrupted paths (stacked
  /v1/files/ prefixes, stale ?ft= tokens) and produce fresh signed URLs
- Sign MediaRefs in chat.history and sessions.get WS handlers
- Strip ?ft= from filename display in frontend
- Add path traversal defense-in-depth check
- Add unit tests for SignMediaPath legacy data healing

Closes #519
2026-03-28 13:33:25 +07:00
Duc Nguyen 19498bff79 fix(providers): register Ollama in-memory on HTTP create, Docker localhost rewrite (#483)
Closes #470. registerInMemory() skipped Ollama because APIKey was empty.
Add Ollama special case before the key guard (mirrors startup code).

When running in Docker, rewrite localhost → host.docker.internal so the
container can reach the host Ollama instance. Extract InDocker() and
DockerLocalhost() into config/runtime.go for reuse.

Add extra_hosts to docker-compose.yml for Linux compatibility.
2026-03-26 12:45:52 +07:00
Duc Nguyen 2445192819 fix(http): add missing fields to update allowlists (#479)
Channel display_name and MCP timeout_sec were silently dropped by
filterAllowedKeys, causing updates to return 200 OK without persisting.

Closes #463
2026-03-26 12:41:28 +07:00
Duc Nguyen 709d86c701 feat(ui): add stop button for running traces (#481)
* feat(ui): add stop button for running traces on traces page

Allows admins to abort running agent runs directly from the traces page,
useful for stopping channel-originated runs (Telegram, Discord, etc.)
without needing access to the chat page.

* style(ui): match stop button style to chat destructive button

* fix(ui): fix abort response handling and variable declaration order

- Check aborted field from chat.abort response for accurate feedback
- Move useTraces() before handleAbortRun to fix block-scoped variable error
- Add abortNotFound i18n key for when run already finished

* fix(ui): remove invalid pending status check, add cursor-pointer
2026-03-26 12:39:29 +07:00
Duc Nguyenandviettranx af3fe2fe08 fix: add panic recovery to tool, cron, and summarization goroutines (#420)
* fix: add panic recovery to tool, cron, and summarization goroutines

A panicking tool, cron job, or auto-summarization goroutine crashes
the entire server process — disconnecting all clients and aborting
all running agents. Add defer/recover at three levels:

- safeExecute() in tools/registry.go: catches panics from any tool's
  Execute method and returns an error result to the LLM
- Agent loop goroutine (loop.go): defense-in-depth for non-tool code
  in the parallel tool execution goroutine
- Cron execution goroutine (service_execution.go): prevents a single
  panicking cron job from taking down the server
- Summarization goroutine (loop_history.go): prevents background
  auto-summarization panics from crashing the process

All recovery points log the panic value and stack trace at ERROR level
for debugging. Tool execution panics return actionable error messages
to the LLM so the agent can recover gracefully.

* refactor: extract panic recovery into safego.Recover helper

DRY the duplicated recover+stack+log boilerplate across 4 goroutine
sites into a shared safego.Recover(onPanic, attrs...) function.
Stack buffer bumped from 4096 to 8192 bytes.

---------

Co-authored-by: viettranx <[email protected]>
2026-03-24 21:09:16 +07:00
Duc Nguyenandviettranx ebc82d3265 fix: clear cron session in DB when not cached (cold-cache after restart) (#424)
* fix: clear cron session in DB when not cached (cold-cache after restart)

Fixes #365

PGSessionStore.Reset() only cleared the in-memory cache. After a server
restart the cache is empty, so Reset was a no-op — the next GetOrCreate
loaded the full accumulated history from DB, causing LLM tool loops from
contradictory context.

Add a DB fallback: when the session isn't in cache, issue a direct
UPDATE to clear messages and summary in PostgreSQL. This ensures cron
sessions always start clean regardless of server restarts.

The #294 fix (Reset+Save before each cron run) already had the right
intent but only worked within the same server lifetime.

* fix: move DB call outside mutex and log errors in Reset cold-cache path

ExecContext was running under s.mu.Lock(), blocking all session cache
operations during DB round-trip. Release the lock before the DB call.
Also log ExecContext errors instead of silently discarding them —
silent failure defeats the purpose of the cold-cache fix.

---------

Co-authored-by: viettranx <[email protected]>
2026-03-24 20:47:44 +07:00
Duc Nguyen 39b7c689f2 fix(http): propagate cross-tenant context in HTTP auth middleware (#427)
The requireAuth and requireAuthBearer middlewares correctly detect
cross-tenant admin status (gateway token + owner user ID) but never
call store.WithCrossTenant(ctx). This causes all HTTP tenant management
endpoints (GET/POST /v1/tenants, etc.) to return 403 "insufficient role"
because handlers check store.IsCrossTenant(ctx) which was always false.

The WebSocket path works correctly because the gateway Client object
tracks cross-tenant status independently via client.IsCrossTenant().

Fix: Add store.WithCrossTenant(ctx) in both auth functions when
auth.CrossTenant is true, matching the design of the WS path.

Affected endpoints:
- GET/POST /v1/tenants (list, create)
- GET/PATCH /v1/tenants/{id} (get, update)
- GET/POST/DELETE /v1/tenants/{id}/users (membership)
- Any future endpoint checking store.IsCrossTenant(ctx)
2026-03-24 11:28:53 +07:00
Duc Nguyenandviettranx eab3766c5f fix: check errors in cron store Scan and Unmarshal to prevent data corruption (#422)
* fix: check errors in cron store Scan and Unmarshal to prevent data corruption

Three unchecked errors in the cron store could silently corrupt data:

1. cron_update.go:44 — QueryRowContext().Scan() error ignored when
   fetching current schedule during partial updates. If DB fails,
   all schedule fields default to zero values, corrupting schedule_kind
   to empty string and bypassing all validation.

2. cron_update.go:167 — json.Unmarshal error ignored when parsing
   existing payload for merge. Corrupted JSON leaves payload zero-valued,
   wiping all existing fields (Message, Deliver, Channel, To) when
   only one field was being updated.

3. cron_scan.go:63 — json.Unmarshal error ignored when scanning cron
   job rows. Malformed payload JSON silently returns zero-valued payload,
   masking data corruption.

All three now return descriptive errors instead of proceeding with
corrupted data. NULL payload (valid case) is handled separately.

* fix: add NULL payload guard and check Marshal error in cron update

- Add len(payloadJSON) > 0 guard before Unmarshal in UpdateJob payload
  merge path, consistent with the same guard in scanCronRow. Without
  this, NULL payload (valid for new jobs) would return a false error
  instead of proceeding with the merge.

- Check json.Marshal error instead of discarding it, consistent with
  the error-checking theme of the parent commit.

---------

Co-authored-by: viettranx <[email protected]>
2026-03-24 10:43:08 +07:00
Duc Nguyenandviettranx 23d0b5eb0b fix(providers): auto-clamp max_tokens on model rejection (#267)
* fix(providers): auto-clamp max_tokens on model rejection + fix verify for reasoning models

When OpenAI-compat models reject max_tokens as too large (e.g. gpt-3.5-turbo
supports 4096 but we send 8192), parse the model's stated limit from the 400
error, clamp the value, and retry once. This fixes agent creation for models
with lower output token limits without hardcoding model names.

Also increase the provider verify endpoint's max_tokens from 1 to 50 so
reasoning models (gpt-5, o-series) have enough headroom for internal
reasoning during the check call.

Closes #248, closes #245

* refactor(providers): extract chat retry closure + fix clamp log key

- Extract duplicate retry closure into chatRequestFn() to follow DRY
- Fix slog logging wrong key: body["max_tokens"] was nil for reasoning
  models that use max_completion_tokens — now uses clampedLimit() helper
- Remove unnecessary _ = resp in provider verify endpoint

---------

Co-authored-by: viettranx <[email protected]>
2026-03-19 08:41:20 +07:00
Duc Nguyenandviettranx 2cc9d68cdc fix(tts): config save, Edge provider, media dispatch + dark mode chat (#265)
* fix(tts): config save + Edge provider registration + dark mode chat bubbles

- Wrap TTS config payload in `raw` field for config.patch RPC (#229)
- Always register Edge TTS provider (free, no API key) instead of gating on `enabled` flag
- Fix low-contrast user message bubbles in dark mode chat

* fix(tts): skip duplicate media dispatch when temp file already delivered

When both the agent loop and the message tool dispatch the same TTS
temp file, the first dispatch succeeds and cleanup deletes it. Filter
out missing temp media files before sending to prevent "file not found"
errors and spurious error notifications on Telegram/Slack/Discord.

* feat(tts): include edge-tts in Docker image when Python enabled

Edge TTS is free (no API key) and serves as a universal TTS fallback.
Install it alongside Python in both ENABLE_PYTHON and ENABLE_FULL_SKILLS builds.

* chore(docker): expose build args from .env for compose builds

Pass ENABLE_OTEL, ENABLE_PYTHON, ENABLE_FULL_SKILLS as env-driven
build args so .env can control Docker build features without editing
docker-compose.yml directly.

* fix(tts): hot-reload TTS config on settings change via pub/sub

TTS providers were only registered at startup, so changing provider/API
key via the Web UI had no effect until container restart. Add a
tts-config-reload bus subscriber that rebuilds the TTS manager on
config changes, matching the pattern used by quota, cron, and web_fetch.
Always create a TtsTool at startup (even without providers) so the
reload subscriber can populate it when settings are first configured.

* fix(tts): protect TtsTool.UpdateManager with RWMutex to prevent data race

UpdateManager() can be called from the config reload goroutine while
Execute() reads t.manager concurrently from agent goroutines. Add
sync.RWMutex following the same pattern as WebFetchTool.UpdatePolicy().

Also update setupTTS doc comment which incorrectly stated it could
return nil — Edge TTS is now always registered.

---------

Co-authored-by: viettranx <[email protected]>
2026-03-19 08:21:06 +07:00
Duc Nguyenandviettranx dc51018563 fix: subagent provider routing + api_base fallback (#262)
* fix(subagent): inherit parent agent's provider instead of alphabetical fallback

Subagents previously used a fixed provider (alphabetically first from the
registry, often "anthropic") regardless of which provider the parent agent
used. This caused invalid combos like anthropic/glm-5 when a zai-coding
agent spawned subagents.

- Pass provider registry to SubagentManager for runtime resolution
- Inject parent provider name into context (WithParentProvider)
- Resolve activeProvider from parent context before LLM call
- Fix trace spans to show actual resolved provider, not default

* fix(providers): api_base fallback from config/env for DB providers

DB providers with empty api_base now inherit from config/env vars
(e.g., GOCLAW_ANTHROPIC_BASE_URL). Prevents proxy API keys from being
sent to the real provider API endpoint.

- Add APIBaseForType() method on ProvidersConfig
- registerProvidersFromDB falls back to config when api_base is empty
- ProvidersHandler uses resolveAPIBase() for model listing
- Add api_base, display_name, settings to provider validation whitelist

* fix(tracing): pass resolved provider name to subagent span emitters

- emitSubagentSpanStart now accepts providerName param instead of
  reading sm.provider.Name() — ensures root subagent span reflects
  the inherited parent provider, not the fallback default
- registerInMemory now uses resolveAPIBase() so DB providers with
  empty api_base inherit the config/env fallback (same as startup path)

---------

Co-authored-by: viettranx <[email protected]>
2026-03-18 22:40:49 +07:00
Duc Nguyen 7436d6b646 fix(i18n): add missing config.json locale files (#108)
* fix(i18n): add missing config.json locale files

The config namespace was registered in i18n/index.ts but the actual
JSON files were never created, breaking the UI Docker build (tsc error).

Also fix .gitignore: config.json → /config.json so only the root
gateway config is ignored, not the i18n locale files.

Closes #107

* feat(ui): replace language cycle button with dropdown select

Users can now pick a language directly from a dropdown instead of
cycling through all options on each click.

* fix(i18n): keep "Subagent" in Vietnamese translations

"Agent con" and "sinh sản" sound unnatural; keep the English term.

* chore: align .dockerignore config.json pattern with .gitignore
2026-03-10 09:50:52 +07:00
Duc Nguyen e63ff014cb feat: add Z.ai provider support (general API + coding plan) (#102)
* feat: add Z.ai provider support (general API + coding plan)

Add Z.ai (GLM) as a new LLM provider with two variants:
- `zai`: general API (api.z.ai/api/paas/v4)
- `zai_coding`: coding plan (api.z.ai/api/coding/paas/v4)

Reuses OpenAIProvider — Z.ai API is OpenAI-compatible with Bearer
token auth, SSE streaming, and reasoning_content support.

Includes: store constants, config struct fields, env var loading
(GOCLAW_ZAI_API_KEY, GOCLAW_ZAI_CODING_API_KEY), secret masking,
config + DB registration, onboard wizard, and UI provider types.

Default model: glm-5

Closes #100

* docs: add Z.ai provider entries to providers documentation
2026-03-09 22:28:18 +07:00
Duc Nguyen e05a4018c9 fix: use platform type instead of instance name in system prompt + Zalo group routing (#90)
* fix(agent): use ChannelType in system prompt for proper channel context

The system prompt was using the channel instance name (e.g. "zep-lao") instead
of the platform type (e.g. "zalo_personal"), causing the LLM to not understand
which messaging platform it's running on. This led to context confusion where
the bot would ask users which channel to send to instead of using the current one.

Changes:
- Add ChannelType field to RunRequest and SystemPromptConfig
- Thread channel type from consumer/cron → agent loop → system prompt
- Add WithToolChannelType/ToolChannelTypeFromCtx for tool context
- Register channel types for both config-based and DB-loaded instances
- Fix Zalo group thread type detection with approvedGroups cache
- Update cron handler to resolve channel type for cron-triggered runs

* refactor(channels): add Type() to Channel interface, remove channelTypes map

Move channel type from a separate map in Manager to the Channel interface
itself. BaseChannel.Type() falls back to Name() for config-based channels
where name == type. Extracts resolveChannelType helper to DRY up 6
repeated resolution blocks across consumer and cron handlers.

* feat(zalo): add pending group history for conversation context

Zalo personal groups now record non-@mentioned messages in a ring buffer
(default 50, configurable via history_limit). When the bot IS mentioned,
pending history is flushed as context — matching Telegram/Discord/Feishu.

Separated mention gating from policy gating in checkGroupPolicy for
cleaner control flow.
2026-03-09 08:30:45 +07:00
Duc Nguyen 137a986d4f feat(channels): add Slack channel (#83)
* feat(channels): add Slack channel via Socket Mode (#37)

Implement Slack integration using Socket Mode (xapp-/xoxb- tokens):
- Event-driven messaging via app_mention + message events
- Policy checks: open, pairing, allowlist, disabled (DM + group)
- Thread participation with configurable TTL
- Markdown-to-mrkdwn formatting pipeline
- Streaming support (edit-in-place + native ChatStreamer)
- SSRF-protected file downloads
- Debounce, dedup, reactions, group history context
- 170 unit tests (format, helpers, stream, SSRF)

Fix BaseChannel.HandleMessage allowlist to also check chatID,
enabling group allowlist with channel IDs across all channels.

Closes #37

* feat(slack): add file/media support and edit-to-mention handling

- Wire inbound file download into handleMessage (images, audio, documents)
- Add media.go with resolveMedia, classifyMime, buildMediaTags
- Extract shared ExtractDocumentContent to channels/media_utils.go (DRY with Telegram)
- Support file_share and message_changed subtypes
- Handle edit-to-mention: respond when user edits old message to add @bot
- Add MediaMaxBytes config field (default 20MB)
- Fix debounce media accumulation (was silently dropping files)
- Add 60s HTTP client timeout on file downloads
- Refactor downloadFile signature for slack.File compatibility
2026-03-09 07:02:37 +07:00
Duc Nguyen 84113cff2f fix(channels): fix Zalo personal group pairing bypass and reply thread type (#76)
Two bugs in Zalo personal channel policy:

1. Group pairing bypass: checkGroupPolicy() called IsAllowed(groupID)
   directly, which returns true when allowlist is empty — effectively
   skipping pairing for all groups. Fixed to match the DM pairing
   pattern: HasAllowList() && IsAllowed(groupID).

2. Wrong thread type for group pairing reply: sendPairingReply() always
   used ThreadTypeUser even when replying to a group. Now detects
   group sender IDs (prefixed "group:") and uses ThreadTypeGroup.
2026-03-07 07:12:47 +07:00
Duc Nguyen 0f5dd08f76 feat(channels): introduce Zalo Personal channel integration (#32)
* feat(channels): implement Zalo Personal Chat (ZCA) protocol layer

Implement complete Zalo Personal Chat integration including:
- Message protocol layer (request/response/event types)
- Connection management with auth flow
- Message sending/receiving with text and media support
- User/group management and sync
- Telegram-style contact and conversation handling
- Comprehensive unit tests with 85%+ coverage

Architecture follows existing channel patterns (Telegram, Feishu) with
raw API calls for session management and message delivery. Includes
error handling, rate limiting awareness, and logging.

* feat(channels): add Zalo Personal channel integration layer

Wire protocol package to GoClaw's channel system:
- channel.go: Channel struct, Start/Stop/Send, listenLoop, message handlers
- auth.go: credential resolution (preloaded > file > QR), persistence
- policy.go: DM/group policy, @mention gating, pairing with debounce
- factory.go: managed mode factory (requires credentials, no QR)
- cmd/gateway.go: register standalone + managed factory

* feat(ui): add Zalo Personal channel type to web dashboard

Add zalo_personal to channel type dropdown, credential fields
(IMEI, cookie, userAgent), and config schema (DM/group policy,
require_mention, allow_from).

* feat(channels): add WebSocket QR login for Zalo Personal channel

Add real-time QR code login flow for zalo_personal channel instances
in managed mode. Users create an instance without credentials, then
trigger QR login from the web dashboard.

Backend:
- New RPC method zalo.personal.qr.start with per-instance mutex
- QR PNG pushed via client-scoped WS events (not broadcast)
- Credentials encrypted and saved to DB on successful scan
- Cache invalidation triggers automatic channel reload/start
- Factory returns nil,nil for missing credentials (skip, not error)
- Instance loader handles nil-channel gracefully

Frontend:
- ZaloPersonalQRDialog with auto-start, retry, and auto-close
- QR button in channel instances table for zalo_personal type
- Credential fields no longer required (auto-populated via QR)

* fix(channels): skip redundant LoginWithCredentials after QR login

QR flow already validates session via qrCheckSession + qrGetUserInfo.
Calling LoginWithCredentials again conflicts with the active QR session
state, causing "empty response" errors. Credentials are validated when
the channel starts instead. Also rename log prefix from "zca" to
"Zalo Personal".

* fix(channels): fix Zalo Personal cookie domain for login API

BuildCookieJar only set cookies for chat.zalo.me but the login API
uses wpa.chat.zalo.me. Cookies weren't sent to the subdomain, causing
"empty response" on channel startup. Now sets cookies for both hosts.

* fix(channels): move UTF-8 check after gzip decompression in Zalo listener

The UTF-8 validity check in decryptAESGCMPayload ran on raw decrypted
bytes before gzip decompression, causing all encType=2 (AES-GCM+gzip)
messages to fail with "decrypted payload is not valid UTF-8".

Move the check to decryptEventData so it runs after all processing
(decryption + decompression) is complete.

* feat(channels): add QR-only onboarding and contacts picker for Zalo Personal

- Remove credential text fields for zalo_personal, show QR auth info banner
- Add has_credentials boolean to HTTP and WS mask functions
- Implement FetchFriends/FetchGroups protocol (encrypted Zalo API)
- Add zalo.personal.contacts WS RPC method with parallel fetch
- Create ZaloContactsPicker component with search, selection, manual entry
- Integrate picker in channel instance edit dialog for allow_from config

* refactor(channels): rename zca error prefix to zalo_personal across protocol package

* fix(channels): unwrap inner response envelope in Zalo contacts decryption

The Zalo API returns double-wrapped responses: outer envelope contains
encrypted base64 data, which when decrypted yields another Response
envelope with error_code and data fields. The decryptDataField helper
was returning the raw decrypted bytes without unwrapping the inner
envelope, causing json unmarshal failures when parsing friends/groups.

* fix(channels): pass version 0 for group details to get full data

The Zalo group info endpoint uses a version-based caching mechanism.
Passing the actual version from step 1 causes the server to return
the group in "unchangedsGroup" with empty "gridInfoMap". By passing
version 0 for all groups, we force the server to return full group
info including name, avatar, and member count.

* fix(ui): auto-load contacts on modal reopen to resolve display names

When the edit modal is reopened with already-selected contact IDs,
contacts are now auto-fetched so badges show display names instead
of raw numeric IDs.

* fix(channels): handle gzip-compressed response in Zalo SendMessage

SendMessage used io.ReadAll + json.Unmarshal directly but the response
is gzip-compressed (Accept-Encoding: gzip header). Use readJSON() which
handles gzip decompression, fixing "invalid character '\x1f'" errors.

* fix(channels): decrypt encrypted send response in Zalo SendMessage

The Zalo send message API response is encrypted like all other endpoints.
Parse outer envelope, decrypt the data field, then extract msgId from
the decrypted inner response.

* feat(channels): improve Zalo listener reliability and UI channel wizard

- Migrate WebSocket client from gorilla to coder/websocket, eliminating
  unsafe/reflect hacks for RSV1 decompression and buffer inspection
- Add channel-level restart with exponential backoff (2s→60s cap, max 10)
  so channels auto-recover instead of stopping permanently
- Reset listener retry counters after 60s stable connection to prevent
  long-lived connections from exhausting retry budget
- Add code 3000 (duplicate session) recovery with 60s initial delay
- Detect silent disconnects via read deadline (2.5x ping interval)
- Fix Stop() to always cancel context, preventing reconnect timer leaks
- Refactor UI channel form into wizard-based flow with registry pattern
- Auto-refresh channel status after create/update dialog closes

* refactor(channels): move Zalo RPC methods to zalomethods package

Move Zalo personal channel RPC handlers from internal/gateway/methods to
internal/channels/zalo/personal/zalomethods, improving code organization
and removing prefix redundancy. Rename types: ZaloPersonalQRMethods →
QRMethods, ZaloPersonalContactsMethods → ContactsMethods.

- Move zalo_personal_qr.go → zalomethods/qr.go
- Move zalo_personal_contacts.go → zalomethods/contacts.go
- Update imports in cmd/gateway.go (2 call sites)
- Update internal/channels/zalo/personal imports

* feat(channels): add typing indicator to Zalo Personal channel

Show "typing..." in Zalo while the LLM processes messages, matching
the Telegram/Discord pattern. Uses the shared typing.Controller with
4s keepalive (Zalo typing expires ~5s) and 60s TTL safety net.

* feat(channels): handle image attachments in Zalo Personal channel

- Add Raw field to Content struct to preserve non-string JSON payloads
- Add Attachment struct with IsImage() detection (ext + Zalo CDN paths)
- Add AttachmentText() for human-readable placeholders (image/file/other)
- Download image attachments to temp files for agent vision pipeline
- Non-image files get text placeholder only (no download)
- Fix URL query param stripping in file extension detection

* fix(channels): switch Zalo WS client to gorilla/websocket with cookie jar fix

coder/websocket did not propagate session cookies for wss:// URLs,
causing Zalo backend to reject connections with "zpw_sek not found".
Switch to gorilla/websocket which handles wss→https scheme conversion
natively. Add wsJar safety wrapper and fix Close() mutex consistency.

Also update Makefile `up` target to use --no-cache builds.

* fix(channels): inject cookies manually for Zalo WS connection

Replace wsJar wrapper with direct cookie injection from chat.zalo.me
base domain. Fixes host-only cookies (zpw_sek) not matching WS
subdomains (ws*-msg.chat.zalo.me) due to Go cookiejar limitations.

* fix(channels): harden Zalo Personal channel security and concurrency

- Add SSRF protection to downloadFile using CheckSSRF (URL validation,
  private IP blocking, DNS pinning) with context and 30s timeout
- Protect c.sess/c.listener with sync.RWMutex to eliminate data races
  during restart; add thread-safe session()/getListener() accessors
- Add stopped flag + reconnTimer to Listener to prevent zombie reconnects
  after Stop(); timer cancelled on Stop(), checked before Start()
- Fix QR flow using context.Background() detached from WS client; now
  derives from parent ctx so flow cancels on client disconnect
- Set initial 30s read deadline for cipher key handshake to prevent
  indefinite blocking before ping loop starts
- Use defer in WSClient.Close() to prevent connection leak on panic
- Document ReadMessage ctx limitation and two-layer reconnect design

* chore: remove unused gobwas/ws dependency from go.mod

gobwas/ws was a leftover from the previous coder/websocket usage,
no longer imported by any Go source files.

* fix(channels): align Zalo Personal policy defaults across UI and backend

Policy defaults were inconsistent across three layers causing group/DM
allowlist enforcement to silently fail. New() applied "allowlist" default
to local vars but never wrote back to config; checkGroupPolicy() then
read empty string and defaulted to "open", bypassing the allowlist.
UI Select components displayed schema defaults visually without
persisting them to configValues, so DB config never stored the policy.
2026-03-03 14:21:07 +07:00
Duc Nguyen f1397081d2 feat(skills): per-agent skill filtering with grant-based access control (#45)
* fix(store): expand tilde in skills storage directory path

The default skillsDir (~/.goclaw/skills-store) was not expanded,
causing os.MkdirAll to fail when creating skill upload directories.

* feat(skills): per-agent skill filtering with grant-based access control (#42)

Wire skill_agent_grants into the agent resolver so each agent only sees
skills explicitly granted to it. Add Skills tab to the web UI for
managing per-agent skill grants with toggle switches.

- Add SkillAccessStore interface to avoid import cycles
- Filter skills in resolver via ListAccessible + filesystem union
- Add GET /v1/agents/:id/skills endpoint with grant status
- Invoke onGrantChange callback to invalidate agent caches on grant/revoke
- Add agent-skills-tab React component with Switch toggles
- Allow read_file access to managed skills-store directory
- Fix rows.Err() propagation in ListAccessible/ListWithGrantStatus

Closes #42
2026-03-03 09:50:26 +07:00
Duc Nguyen 4c67dff24d feat(providers): support custom base URL for Anthropic provider (#16)
Allow overriding the Anthropic API base URL via GOCLAW_ANTHROPIC_BASE_URL
env var, config JSON, or DB provider record. Enables use of Anthropic-
compatible proxies and custom endpoints.

Also adds Makefile shortcuts for docker compose (up/down/logs).
2026-02-28 11:50:00 +07:00