6 Commits
Author SHA1 Message Date
thotam 0c1ededc92 feat(webhooks): return per-call usage breakdown with provider/model/cost (#1421)
feat(webhooks): per-call usage breakdown with provider/model/cost (#1421)
2026-07-10 22:56:05 +07:00
thotam f87cdaf616 feat(webhooks): split webhook usage tokens by cache vs non-cache (#1345)
Surface cache-read and cache-creation input tokens (plus the
prompt_tokens_include_cached_segments flag) in webhook usage responses,
so callers can distinguish cached from non-cached input tokens.

The data already flowed end-to-end via providers.Usage and
agent.RunResult.Usage; the two webhook envelope structs (webhookLLMUsage
for the sync LLM webhook, callbackUsage for the async callback payload)
copied only 3 of the fields, dropping the cache data. Add the 3 cache
fields (mirroring providers.Usage JSON tags exactly, all omitempty) and
populate them. The async callbackUsage mapping is extracted into a pure
newCallbackUsage helper for unit testing.

Additive and non-breaking: omitempty keeps responses without caching
byte-identical. No schema change.
2026-07-04 23:26:49 +07:00
thotam e2ec3370dd feat(webhooks): stream provider responses for server-side runs to enable prompt caching (#1273)
Server-side webhook agent runs (sync, async, and admin test) used Stream:false,
so OpenAI-compatible routers that only cache streaming requests never populated
or served their prompt cache. Webhook runs paid full input-token price on every
turn even with a stable session and an identical multi-turn prefix, while WS chat
(Stream:true) got cache hits from the 3rd turn.

Add gateway.webhook_stream (default true, env GOCLAW_WEBHOOK_STREAM) and apply it
to all three run sites. ChatStream returns the fully assembled response, so the
payload returned to the caller is unchanged. Set to false to restore non-streaming.
2026-06-24 13:42:20 +07:00
thotam 06ef06ca37 feat(webhooks): paginate list + call-history endpoints (server + web UI) (#1268)
Add total-count pagination to the webhook admin endpoints and the web UI.

Store:
- WebhookStore/WebhookCallStore gain Count; WebhookListFilter gains
  IncludeRevoked + Query (PG + SQLite, parameterized, tenant-scoped)

API:
- GET /v1/webhooks and GET /v1/webhooks/{id}/calls return
  {items, total, limit, offset} with server-side search + revoked filtering

Web UI:
- server-driven list pager + search/revoked filter; call-history dialog uses
  the real total (fixes the full-page "has more" boundary bug)
- i18n pager labels (en/vi/zh)

Tests: store pagination integration test; mock stores implement Count.
No schema migration (read-only COUNT).
2026-06-24 10:39:24 +07:00
thotam a10464290d feat(webhooks): configurable agent run timeout (default 600s) (#1267)
Replace the hardcoded 30s webhook agent-run deadline with a configurable
timeout (default 600s, cap 3600s) for both the async worker and the
sync/test HTTP handler. Legitimate multi-step runs were dying at 30s
mid-tool-call; the same request over WebSocket completed fine.

- internal/webhooks/timeout.go: ResolveTimeoutSec helper (<=0 → 600s, cap 3600s)
- WorkerConfig.AsyncAgentTimeout (async worker) + WebhookLLMHandler.syncTimeout
  (sync + admin test), wired from config
- config keys gateway.webhook_{async,sync}_timeout_sec, also settable via
  GOCLAW_WEBHOOK_{ASYNC,SYNC}_TIMEOUT_SEC env (env overrides config)
- tests: timeout bounds + config file/env override
2026-06-24 09:22:06 +07:00
Duy /zuey/ ddf8e1099f feat(webhooks): HTTP webhooks to trigger agents with HMAC auth + durable callbacks (#2)
* feat(webhooks): HTTP webhooks to trigger agents with HMAC auth and durable callbacks

Add multi-tenant HTTP webhook endpoints for agent triggering:
- /v1/webhooks/message: send messages to channels
- /v1/webhooks/llm: sync/async LLM prompts with HMAC-signed callbacks
- HMAC-256 + bearer token authentication
- Rate limiting and tenant isolation
- Durable callback worker with exponential backoff
- PG 000056 + SQLite schema v25 migrations
- Unit + integration tests, P0 tenant isolation invariants
- Channel media capability helpers for attachment routing
- Comprehensive webhook documentation and i18n strings

* fix(webhooks): address post-review findings (K1-K10)

Comprehensive post-merge fixes addressing 10 blocking code review issues
and 2 adversarial re-audit findings in webhook-agent-triggering feature:

K1: Fix auth middleware tenant context lookup sequencing — move
    tenant context injection before authenticate() call to prevent
    unscoped secret lookups.

K2: Canonicalize JSON payload format for jsonb compatibility across
    PostgreSQL and SQLite — ensure consistent serialization without
    whitespace variance to prevent hash mismatches.

K3: Add fail-closed JSON parsing in body hash extraction with explicit
    error handling for malformed payloads before HMAC verification.

K4: Fix worker queue wedge by properly draining slot reservations
    when delivery succeeds, preventing permanent slot occupancy.

K5: Implement lease-token optimistic concurrency control to prevent
    duplicate webhook delivery under high concurrency or retry storms.

K6: Add AES-256-GCM encrypted secret storage at rest with fail-fast
    skip-mount when GOCLAW_ENCRYPTION_KEY environment variable unset.

K7: Implement IP allowlist enforcement supporting both CIDR ranges
    and exact IP matching with proper X-Forwarded-For parsing.

K8: Add HMAC replay nonce cache (5min expiry, non-blocking async flush)
    to prevent request replay attacks on webhook handler.

K9: Fix invariant test schema selection — replace hardcoded assumption
    with explicit schema name from config to support multi-schema testing.

K10: Consolidate rate limiters into single shared instance to prevent
     per-endpoint limiter starvation and ensure fair rate limiting.

New database migrations:
- 000057: webhook_calls.lease_token for optimistic concurrency
- 000058: webhooks.encrypted_secret_key for AES-256-GCM encryption

New i18n keys: MsgWebhookIPDenied, MsgWebhookEncryptionUnavailable
(with English, Vietnamese, Chinese translations).

New modules:
- internal/http/webhooks_payload.go: JSON canonicalization + body hash
- internal/http/webhooks_nonce.go: Replay nonce cache implementation
- internal/http/webhooks_idempotency_test.go: Integration tests

Documentation updates:
- docs/webhooks.md: §13-14 security sections, encryption flow
- docs/00-architecture-overview.md: webhook subsystem security overview
- docs/codebase-summary.md: webhook security patterns
- docs/project-changelog.md: webhook fixes changelog

Test coverage: 53 webhook tests + 4 P0 invariant tests all passing.
No tenant isolation violations. All security gates enforced.

* docs(journals): webhook feature ship + fix cycle entries

* fix(webhooks): address Claude review findings

- webhooks_llm.go: remove misleading ptr() helper; use &completedAt
  pattern for error-path audit rows (matches success path)
- webhooks_auth.go: wrap TouchLastUsed context in WithoutCancel so
  background DB update isn't cancelled when HTTP response completes
- store GetByIDUnscoped (PG+SQLite): add NOT revoked / revoked = 0
  filter for defense-in-depth parity with GetByHashUnscoped
- webhooks/sign.go: fix package doc — HMAC key is raw plaintext
  secret bytes, not hex-decoded SHA-256
- webhooks_admin.go: check auth before encKey guard to avoid leaking
  config state to unauthenticated callers
- webhooks_ratelimit.go: two-phase Load→LoadOrStore to avoid per-call
  entry allocation on the hot path

* docs(webhooks): fix Sign() function doc to match actual key input

Function-level comment still referenced hex-decoded SecretHash after
the package-level doc was corrected. Align with actual caller usage
([]byte(rawSecret)).

* fix(webhooks): use WithoutCancel for worker execute DB updates

Terminal status writes in execute() ran through the worker main-loop
ctx, which is cancelled on graceful shutdown. If the outbound send
completed but the status update raced with shutdown, the row stayed
in 'running' and got re-delivered via reclaimStale. WithoutCancel
lets the DB write survive worker cancellation while preserving
propagated values (tenant ID, etc.).

* fix(webhooks): move tctx init before panic defer in worker execute

Panic recovery called updateRetry with raw ctx (no tenant ID), making
requireTenantID fail and the reset-to-retry DB write silently drop.
Row stayed 'running' until reclaimStale (~90s delay). Init tctx first
so defer closure captures tenant-scoped non-cancellable context.

* fix(webhooks): pass tenant-scoped tctx to invokeAgent in worker

execute() was passing the raw worker-loop ctx (no tenant ID) to
invokeAgent → router.Get → PGAgentStore.GetByID. GetByID reads
TenantIDFromContext which returned uuid.Nil, making every lookup
return 'agent not found'. Async LLM webhook calls silently failed
all retries. Pass tctx (already tenant-scoped + WithoutCancel) so
the router resolves the agent correctly.

* fix(tests): resolve integration test compile errors

- Remove duplicate contains() in mcp_grant_revoke_test.go (already
  defined in tts_gemini_live_test.go)
- Update webhooks_admin_test.go RotateSecret call to match current
  5-arg signature (newSecretHash, newPrefix, newEncryptedSecret)

* fix(webhooks): default nil scopes/ip_allowlist to empty slice in Create

PG columns are NOT NULL DEFAULT '{}'. Explicit NULL from pqStringArray(nil)
violated the constraint, breaking TestWebhookAdminCRUD/TenantIsolation.
Coerce nil slices to empty []string{} so the default applies at the DB layer.

* chore: trigger CI on digitopvn/goclaw fork

* ci: retrigger workflows

* fix(webhooks): renumber migrations to 000059-000061 for merge train
2026-05-11 13:29:24 +07:00