setupToolRegistry creates the rate limiter from cfg.Tools.RateLimitPerHour
during early bootstrap (gateway.go line 137). System_configs DB overlay
runs ~50 lines later via cfg.ApplySystemConfigs (line 191), so any DB
override of tools.rate_limit_per_hour was silently lost - the limiter
object was already initialised from the JSON5 default.
Symptom in production: editing tools.rate_limit_per_hour via system_configs
table or the config HTTP API had no effect on running gateways. Operators
had to inject a config.json file to change the value, defeating the
DB-as-source-of-truth pattern that other tunables rely on.
Re-apply the limiter after ApplySystemConfigs runs. Safe ordering: server
has not started yet, no in-flight tool calls, and SetRateLimiter is a
plain field assignment with no shutdown cost on the discarded limiter.
nil case (rate_limit_per_hour <= 0) also handled so DB writes can disable
the limiter without restart.
Tests: new TestRegistry_SetRateLimiter_ReplacesPriorLimiter covers both
the replace-with-higher-limit path and the nil-disables path. All
existing rate limiter tests still pass.
Note: the same ordering pattern likely affects other config consumed
inside setupToolRegistry (tools.scrub_credentials, MCP server wiring).
Out of scope for this hotfix - they need a deeper restructure.
* feat(providers): explicit prompt cache for DashScope/Qwen
Extend OpenAI-compat path with Anthropic-style cache_control:ephemeral
inline blocks for Alibaba DashScope endpoints (provider_type=bailian or
URL contains "dashscope"). Verified live: qwen3.6-plus / qwen3.5-plus /
qwen3-coder-plus all return 99.7% cache hit rate on 6K-token prefix,
yielding ~89% token cost reduction on cached prefix (90% Alibaba discount).
- isDashScope(): 3-source detection (URL + providerType + name) handles
reverse-proxied endpoints; covers both "dashscope" and "bailian"
- buildRequestBody wraps system content with SplitSystemPromptForCache
(exported from anthropic_request.go) using <!-- GOCLAW_CACHE_BOUNDARY -->
- Tool prefix cache: cache_control on last tool definition, with 4-marker
budget guard
- Parse cache_creation_input_tokens from prompt_tokens_details into
Usage.CacheCreationTokens; propagates via existing span metadata writers
- Runtime escape hatch: GOCLAW_DISABLE_DASHSCOPE_CACHE=true
- Live integration smoke test (build tag integration, env-gated)
* chore(providers): polish per PR #1127 review
- Use strings.Repeat instead of custom repeat() helper in smoke test
- Clarify BuildRequestBodyForTest is test-only, not public API
* feat(providers): enable thinking for Qwen 3.7/3.6, observe cache in smoke test
Qwen3.7-plus and qwen3.6-plus support deep thinking but were missing from
dashscopeThinkingModels, so enable_thinking/thinking_budget was silently
skipped. Add both to the whitelist and test.
Add qwen3.7-plus to the cache smoke test as cache-optional: Alibaba's
context-cache doc does not yet list 3.7/3.6 for explicit cache, and the
cache wrap is a safe no-op when unsupported, so a no-cache result is logged
rather than failed to avoid a flaky live assertion.
Claude-Session: https://claude.ai/code/session_01X3jkrc7N3ar8ZGzUJWyNS5
* fix(providers): preserve DashScope detection for proxy routes
Add TikTok account type dropdown (Livestream AIO, Business Messaging,
TikTok Shop) to Pancake channel config with prefix-aware webhook routing.
- Add tt, ttm, tts prefixes to platformPrefixes map for convID parsing
- Add TikTokType field to pancakeInstanceConfig (optional, backward compat)
- Add tiktok_type field to UI with showWhen platform=tiktok
- Add i18n labels for en/vi/zh locales
Co-authored-by: viettranx <viettranx@gmail.com>
* feat(tools): add vault tool group to policy
Add "vault" group containing vault_search and vault_read tools.
Include group:vault in messaging and coding profiles so agents
using these profiles can search and read vault documents.
Previously, vault tools were only in the goclaw composite group
but not exposed to profile-restricted agents like Telegram bots.
* fix(tests): use strings.Contains to avoid redeclaration
---------
Co-authored-by: viettranx <edu@200lab.io>
* feat(pancake): add Shopee platform support
- Fix Accept header for JSON response negotiation (Shopee GETs)
- Platform-prefix-aware pageID parsing (spo_ convID format)
- 500-char message limit for Shopee DMs
- Rune-safe truncation with slog.Warn on truncation
Phase 3 (integration smoke) pending: requires real credentials.
* fix(pancake): address PR review feedback
- Add RegisterPlatformPrefix() for extensibility (issue #1)
- Rename setCommonHeaders → setAcceptJSONHeader (issue #2)
- Add testdata fixture with page_id for priority testing (issue #3)
* chore: retrigger CI
* fix(permissions): classify sessions.compact in RBAC policy
Add MethodSessionsCompact to isWriteMethod to fix drift coverage test.
Method was added in #958 but missing from permissions classification.
* fix(tests): use strings.Contains to avoid redeclaration
Remove duplicate contains helper in tts_gemini_live_test.go that
conflicts with mcp_grant_revoke_test.go in same package.
* refactor(pancake): guard platformPrefixes with RWMutex for concurrent safety
RegisterPlatformPrefix is exported for future use (Lazada, Tokopedia) but
webhook handler reads the map concurrently. Protect with sync.RWMutex so
callers can register prefixes at any time without data races.
- Add platformPrefixesMu sync.RWMutex
- Wrap RegisterPlatformPrefix with write lock
- Extract isKnownPlatformPrefix helper with read lock
- Document that RegisterPlatformPrefix is currently unused (extension point)
---------
Co-authored-by: viettranx <viettranx@gmail.com>
* feat(pancake): private-reply funnel — 2 modes, scope filter, DB dedup, metrics
Refactor the pancake channel's one-time DM feature from the 5-day-old minimal
`first_inbox` toggle into a complete comment → DM funnel. Breaking rename
(first_inbox → private_reply) matches Pancake + Facebook Graph terminology.
Channel logic
- Two modes: `after_reply` (default — public reply then DM) and `standalone`
(DM-only; bypasses keyword filter, publishes synthetic outbound via
`ch.Bus().PublishOutbound` so the LLM pipeline is skipped for template DMs).
- Post-level scope filter (`private_reply_options.allow_post_ids` /
`deny_post_ids`, deny beats allow).
- Template vars `{{commenter_name}}` / `{{post_title}}` with pre-sanitized
literal-replace (strips `{{`/`}}` from values — no nested substitution).
- Locale-aware default text via `i18n.T(locale, MsgPancakePrivateReplyDefault)`
with en/vi/zh catalogs and English fallback.
DB dedup (replaces in-memory sync.Map)
- New `store.PancakePrivateReplyStore` (PG + SQLite impls), tenant-scoped via
`store.TenantIDFromContext`, fail-closed on missing tenant.
- Atomic `TryClaim` / `Unclaim` (INSERT ... ON CONFLICT DO UPDATE WHERE stale)
eliminates the concurrent-comment TOCTOU that would have fired duplicate
DMs — claim first, release on API failure so the next comment can retry.
- Configurable TTL (`private_reply_ttl_days`, default 7).
- Dual-DB: migration 000056 on PG; schema v25 on SQLite.
Bumps RequiredSchemaVersion 55 → 56.
Wiring
- `pancake.FactoryWithStores(ppReplyStore)` follows the existing
`FactoryWithStores*` pattern (telegram, discord, whatsapp) and replaces
`pancake.Factory` in `cmd/gateway.go`.
- `Channel.Send` injects `store.WithTenantID(ctx, ch.TenantID())` before
dispatch so the outbound path sees a non-nil tenant (regression guard:
without this the new store fails-closed and drops DMs in production — unit
fakes masked this by bypassing the tenant check).
- Routing metadata whitelist extended: `private_reply_mode`,
`private_reply_only`, `post_id`, `display_name` survive inbound → outbound.
Observability
- New `internal/metrics` package with `pancake_private_reply_total{
page_id, result, reason}` counter and `/metrics` endpoint mounted on the
gateway.
- Endpoint gated behind Bearer auth when `gateway.token` is configured; open
(with warning) when unset for local dev scraping. Operators should set a
token or bind `GOCLAW_HOST=127.0.0.1` to avoid leaking page identifiers.
UI
- 6 additive Pancake fields in `channel-schemas.ts` with `showWhen` gating on
`features.private_reply` (itself gated on platform = fb/ig).
- 10 schema tests; full web build passes.
- i18n entries in en/vi/zh `channels.json` + 2 new Go keys
(`MsgPancakePrivateReplyWindowExpired`, `MsgPancakePrivateReplyDefault`).
Tests
- 18 unit tests in pancake package (render/filter/mode switch/scope/dedup/
metrics/tenant-ctx regression); 6 store-level tests; 5 E2E tests against
live PG covering happy path, dedup-across-restart, TTL expiry via SQL,
scope filter, FB 7-day policy error → no claim retained.
- All green on both PG and SQLite build tags with race detector.
* refactor(pancake): simplify private-reply to stateless DM
Roll back the dedup funnel from PR 951 — no table, no in-memory state,
no metrics, no modes, no scope filter. Keep only feature flag + message
template.
Rationale: each platform owning its own dedup table is an anti-pattern
that inflates multi-tenant data footprint. Use the platform itself
(Facebook Graph /comment/private_replies is per-comment idempotent)
combined with the existing webhook comment_id dedup in
comment_handler.go. Zero GoClaw state required.
Removed:
- DB table pancake_private_reply_sent (PG migration 000056, SQLite v25)
- Store interface + PG + SQLite implementations + tests
- internal/metrics package + /metrics endpoint + prometheus dep
- Two modes (after_reply/standalone) + standalone fast-path
- Scope filter (allow/deny post IDs) + PrivateReplyOptions struct
- TTL config (private_reply_ttl_days)
- Locale-aware i18n default (MsgPancakePrivateReplyDefault,
MsgPancakePrivateReplyWindowExpired)
- FactoryWithStores wiring + tenant ctx injection in Send()
- SetHTTPClientForTest dead code
Kept:
- features.private_reply flag
- private_reply_message template with {{commenter_name}} and
{{post_title}} vars (literal-replace, injection-safe)
- After-reply flow: comment -> public reply -> DM
- Hard-coded English fallback when message is empty
Schema: RequiredSchemaVersion 56 -> 55; SQLite SchemaVersion 25 -> 24;
schema.sql DDL block removed.
Net delta: -2661 lines across 44 files. Build (PG + sqliteonly) +
vet + race tests clean.
* test(agent): bump none-mode prompt size budget to 3100
vault_read wiring (#948) added ~95 chars to read_file tool summary,
pushing none-mode prompt from <3000 to 3075 chars. Bump budget to
3100 (~775 tokens) to match the intentional addition.
* test(integration): ensure data_migrations table exists in reset helper
The reset helper runs before RunPendingHooks, but RunPendingHooks is
what normally creates data_migrations. On a fresh CI database the
DELETE fails with 'relation does not exist'. Create the table
defensively so reset works regardless of execution order.
* chore(sqlite): remove accidental trailing blank line in schema.sql
---------
Co-authored-by: viettranx <viettranx@gmail.com>
* fix(vault): prevent vault_read id-namespace collision
vault_search was leaking KG/episodic entity ids into result sets even when
narrow `types` were requested, and callers then passed those ids to
vault_read which returned a generic "document not found". The cause was
threefold:
1. `types` filter was only applied to the vault fan-out; KG and episodic
ran unconditionally. Now gated by shouldFanout(types, key).
2. vault_search output lacked a per-source tool hint. Each result now ends
with " → use <tool>" naming the correct follow-up (vault_read,
knowledge_graph_search, or memory_search).
3. vault_read miss returned "document not found" without checking whether
the id belonged to a foreign namespace. It now probes KG then episodic
and returns a namespace-specific redirect error. Stores are injected
via SetKGStore/SetEpisodicStore, nil-safe, tenant-scoped.
Adds red→green characterization tests plus an end-to-end integration
scenario seeding a vault doc + KG entity with identical basenames.
* test(agent): bump none-mode prompt size budget to 3100
vault_read wiring (#948) added ~95 chars to read_file tool summary,
pushing none-mode prompt from <3000 to 3075 chars. Bump budget to
3100 (~775 tokens) to match the intentional addition.
* test(integration): ensure data_migrations table exists in reset helper
The reset helper runs before RunPendingHooks, but RunPendingHooks is
what normally creates data_migrations. On a fresh CI database the
DELETE fails with 'relation does not exist'. Create the table
defensively so reset works regardless of execution order.
* refactor(vault): per-source id fields + wire episodic into search
Align vault_search output fields with downstream tool input params:
doc_id (vault_read), entity_id (knowledge_graph_search), episodic_id
(memory_expand). Prevents LLMs from pattern-matching a generic `id:`
and misrouting a foreign-namespace uuid into vault_read. Fallback
redirect in vault_read now quotes id + names the correct param so
the LLM can self-correct in one turn.
Also wire stores.Episodic into VaultSearchService (stale comment
claimed pending-impl; PGEpisodicStore has existed and been in use
since v3). Unifies search fan-out with vault_read namespace probe.
---------
Co-authored-by: viettranx <viettranx@gmail.com>
Opt-in feature (features.auto_react: true) that automatically likes
a user's Facebook comment via Pancake /likes endpoint when the comment
webhook is received, as an engagement signal.
- Add AutoReact bool to Features config struct
- Add ReactComment() to APIClient (multipart POST to /likes, auth via api_key)
- Add reactCommentAsync() to comment handler; fires before keyword filter,
independent of comment_reply feature
- Bound with 10-slot semaphore + 5s stopCtx timeout per goroutine
- Guard: only fires on platform=facebook with non-empty convID + msgID
- Startup warning when auto_react=true without webhook_secret (F9)
- Update feature gate: exit only when BOTH auto_react AND comment_reply disabled
- Add 8 tests (ReactComment contract, error body, invalid IDs, handler variants)
* feat(channels): make Pancake platform a required select with 11 options
- Add mandatory platform select field (11 options) to Pancake channel schema
- Add config required validation on channel instance form (create-only)
- Add i18n fieldOptions/fieldConfig for platform in en/vi/zh locales
- Add TDD tests for Pancake config schema (channel-schemas.test.ts)
- Add backend test TestFactoryExplicitPlatformPreserved
- Add slog.Debug for auto-detect path in pancake.go
- Update Platform field comment in types.go for clarity
- Update plan status to completed; add changelog entry
* feat(channels): hide comment_reply for non-social Pancake platforms
Extend showWhen to accept string | string[] values. Apply it to
features.comment_reply so the field only appears for platforms that
support public posts/comments (facebook, instagram, threads, tiktok,
youtube). E-commerce platforms (shopee, lazada, tokopedia) and
messaging-only platforms (line, google, chat_plugin) no longer show
the irrelevant Comment Reply toggle.
Also guard depValue before String() coercion to avoid the "undefined"
literal matching hazard.
* feat(channels): Pancake comment reply fix + platform select + UI improvements
- Fix comment reply: pass message_id (reply_to_comment_id) to Pancake API
ReplyComment now requires messageID param; guard added for empty ID
- Add SendMessageRequest.MessageID field (omitempty) for reply_comment action
- Platform select: required field with 11 options, showWhen gate for comment_reply
- Webhook Page ID moved to Advanced collapsible section in channel form
(auto-expands when existing value is configured)
- routing_metadata.go: centralize routing metadata key constants
- config-flatten.ts: flatten/unflatten nested config for form state
Closes: Pancake comment reply returns 'Missing required field: message_id'
* fix(security): harden exec path exemption matching (#721)
- Add absolute path exemption for dataDir/skills-store/ (fixes skill
scripts using absolute paths like /app/data/skills-store/ being denied)
- Strip surrounding quotes before prefix matching (LLMs often quote paths)
- Reject path traversal ("..") in exempt fields to prevent escape
- Switch from "any field exempt → skip" to per-field matching: only exempt
if ALL fields that match the deny pattern are individually exempt
- Closes pipe/comment bypass vectors where an exempt path in one argument
would exempt the entire command including non-exempt paths
Includes 27 test cases covering: legitimate access, quoted paths,
path traversal, unicode bypass, pipe/comment bypass, mixed args.
* fix(permissions): use cron-specific permission check for cron tool
Cron tool was hardcoded to check `file_writer` configType via
CheckFileWriterPermission(), ignoring the `cron` configType that
the UI actually saves when granting cron permissions. This caused
agents in group chats to be denied cron access even with correct
permission configured.
Add ConfigTypeCron constant and CheckCronPermission() that checks
`cron` configType first, falling back to `file_writer`.
---------
Co-authored-by: Viet Tran <viettranx@gmail.com>
Added ui/web/.npmrc with supportedArchitectures for musl+glibc/arm64+x64.
Updated Dockerfile to use --no-frozen-lockfile so pnpm fetches native rollup
binding compatible with Alpine's musl libc. Lockfile still pinned by copy order.
* fix(discord): preserve attachment source URL in media tags (#602)
- Add SourceURL field to MediaInfo struct
- Populate SourceURL from Discord attachment URL in resolveMedia()
- Emit <media:image url="..."> in BuildMediaTags() when SourceURL is set
- Refactor enrichImageIDs/enrichImagePaths with helper functions
- Add comprehensive unit tests for media URL handling
* fix: update system prompt for url attribute + add enrichment regression tests
- System prompt now documents the url attribute in <media:image> tags
- Add TestEnrichImageIDs_BareTag for non-Discord channel enrichment
- Add TestEnrichImageIDs_SkipsAlreadyEnriched for double-enrichment safety
- Add TestEnrichImagePaths_NoDoubleEnrich for historical message safety
- Add TestEnrichImagePaths_AttributeOrderIndependence for url-before-id tags
---------
Co-authored-by: viettranx <viettranx@gmail.com>
* fix(cron): Run Now executes correctly and reset stale running status (#568)
- PG store RunJob now sets next_run_at = NULL before executeOneJob so
loadClaimedJob (which requires next_run_at IS NULL) succeeds instead
of silently skipping the forced run
- recomputeStaleJobs on startup resets last_status 'running' → 'interrupted'
for jobs that were mid-execution when the server crashed
- UI: add typed CronJobPatch interface matching backend store.CronJobPatch
(deliverChannel/deliverTo/stateless) — replaces Record<string, unknown>
on updateJob and onUpdate props across cron detail components
- UI: fix enabled/disabled label in overview tab to reflect local state
instead of stale job prop
* fix(cron): add error handling for ExecContext calls per code review
- RunJob: return error if claiming job (next_run_at = NULL update) fails
instead of silently proceeding to executeOneJob which would then skip
- recomputeStaleJobs: log warning on failure, log count of reset jobs
- Add log/slog import to cron_exec.go
* fix(providers): claude-cli ChatSync image format mismatch (#577)
When images are present, Chat() now uses stream-json output format to
match the stream-json input format required by Claude CLI >= v2.1.87.
Also adds concurrent execution guard to RunJob() preventing
double-execution when Run Now is clicked rapidly.
---------
Co-authored-by: viettranx <viettranx@gmail.com>