19 Commits
Author SHA1 Message Date
thotam 3f90057c24 fix(pipeline): stop aborting runs on heuristic context budget estimates (#1587)
Runs on models without a registered tokenizer (e.g. 9router brand models)
ended with the generic "Agent couldn't generate a response" fallback even
though the real request used about 55% of the context window.

PruneStage counted history with TokenCounter, which falls back to a
chars/2 heuristic for unregistered models and overcounted about 1.8x.
Once over budget it ran memory flush (~35s, invisible in traces), then
mid-loop compaction, which cannot summarize a history made only of tool
call/result pairs. The callback reported the untouched history as
compacted, PruneStage still saw it over budget and returned AbortRun
before any LLM call, and FinalizeStage replaced the empty reply with the
fallback.

- PruneStage and ContextStage overhead count with the request guard's
  BudgetCounter. PruneStage no longer controls loop flow; the final
  request guard in ThinkStage decides.
- CompactMessages returns ErrNotCompacted when history is unchanged.
  Callers stop counting it as a compaction and do not retry it in the
  same run, while post-run summarization still sees the pressure.
- When the guard exhausts every reduction step, ThinkStage stops the run
  with a localized chat.context_budget_exceeded notice instead of an
  error, so the run's tool results are still persisted. The stop reason
  marks the trace and agent span as error; team tasks, cron and
  heartbeat treat it as a failure via RunOutcome.Failure().
- Memory flush and mid-loop compaction emit event spans.
- Web and desktop UIs treat an unset context_pruning as enabled (the
  backend default since 7639a8c0), keep it unset when untouched, and can
  re-enable pruning after it was turned off.
2026-09-29 18:04:08 +07:00
thotam ce9e7af3fb fix(tools): price out-of-band media by measured duration, not by its bytes (#1564)
read_video called with a url parameter always failed under an agent budget:

    tool:read_video: cannot verify streamed native media against the agent
    context budget (no in-memory payload to count); refusing to send

read_audio and read_document appeared to work but only for very small files.
Both symptoms share one root cause.

5ca433b8 introduced the complete-input invariant and, to keep out-of-band
media honest, appended the standard-base64 encoding of the payload as a
synthetic guard-only message. That message is counted as text. Measured against
the bundled BudgetCounter, base64 costs 0.956 tokens per raw byte, so an 11.6 MB
video counted 11,161,058 tokens and exceeded a 200k, a 1M and a 2M window alike.
The practical ceiling was roughly 200 KB at a 200k window while videoMaxBytes is
100 MB. The read_video URL transport streams to the File API without buffering,
so it had no bytes to count at all and was routed to a helper that refused to
send whenever an agent budget was present.

Media is now priced by what a provider actually bills for it. Every number below
is either a published per-unit rate or a published provider limit; none is
derived from a byte size, because a static-image video compresses arbitrarily
small and no bitrate floor exists. An earlier revision of this branch tried one
and a 120-second, 20,627-byte clip priced at 526 tokens against a true 31,560.

  internal/mediabudget prices one payload:
  - Video and audio: ffprobe measures duration. 263 tokens/second for video,
    32 for audio, both published for static processing at 1 FPS.
  - PDF: pdfinfo counts pages at 258 tokens/page. When the page count cannot be
    read, the charge is the proven ceiling of 1000 pages, which is the most a
    provider will accept and therefore the most it can bill.
  - Video and audio that cannot be measured are refused. Upstream already
    refused unverifiable native media for every URL on the Gemini streamed path;
    this narrows that refusal from every URL to only what genuinely cannot be
    measured, rather than removing it. PDF differs because its page count is
    cheaply measurable and its ceiling is small enough to stay usable.

A remote video is measured without downloading it: two ranged GETs, 512 KB from
the head and 512 KB from the tail, written into a sparse temp file sized to the
declared total and handed to ffprobe. The tail matters because every container
that puts its index at the end keeps it there: head-only probing under-reports
mpeg by 98% and ogg by 79%, and a plain prefix makes ffprobe under-report a
30-second WAV as 0.74 seconds because it clamps to the bytes it can see. Sizing
the temp file to the real total fixes that. Verified end to end on an 11,673,105
byte MP4 served by nginx: 1 MB of ranged reads yielded duration 40.000000,
identical to ffprobe reading the whole URL, for a charge of 10,520 tokens. Those
requests reuse the existing SSRF-safe path, security.WithPinnedIP plus
security.NewSafeClient(0), and no URL is ever handed to an external binary.

Beyond the reported bug, two pre-existing gaps let large media reach a provider
almost unpriced. ExecuteWithChain treated every callProvider error as a provider
failure and advanced to the next entry, and the non-Gemini branches of read_video
and read_document reserved without pricing their payload at all. Measured under a
20,000-token window before this change, a 40 MB video and a 1000-page PDF each
reached a provider charged about 1,600 tokens. Budget refusals are now terminal
in the chain and every media branch prices its payload, so both reach no provider
at all. Genuine provider failures still fail over.

Known limits, stated rather than discovered:
- The /Type /Page scan that guards against a forged /Count is a floor, not a
  bound. Pages inside a compressed object stream are invisible to it, and a
  9,484-byte PDF built that way is charged 258 tokens for 1000 pages. A real
  pdfinfo reads such files correctly; the scan only ever raises a probed count.
- A video URL whose origin does not serve byte ranges is now refused on the
  non-Gemini path too, and the refusal is terminal. Upstream forwarded such URLs
  unpriced. A HEAD giving only a size is not enough to price one.
- read_video and read_audio require ffprobe. Docker images install it by default
  except the base variant; bare binaries and the desktop build do not ship it.
- A hostile origin can craft a container ffprobe reads as about one second.
  Reservation.Reconcile overwrites the estimate with the provider's reported
  usage, so this weakens the gate rather than defeating it.

Byte ceilings videoMaxBytes, audioMaxBytes and documentMaxBytes are unchanged.
No new module dependency; ffprobe and pdfinfo are optional runtime probes.
2026-09-12 10:31:12 +07:00
thotamandClaude Opus 5 c2c5124a0c feat(image): add GPT Image 2.5 and default to flare (#1562)
OpenAI shipped GPT Image 2.5 on 2026-09-08 as two model IDs rather than one:
flare is tuned for speed, sunburst for editing precision. Both accept text and
image input and run on /images/generations, /images/edits and the Responses
API, which is the path the native image_generation tool already uses.

Flare replaces gpt-image-2 as the default because OpenAI reports higher quality
at roughly half the latency, which suits create_image's one-shot call. Both
gpt-image-2 and gpt-image-1.5 stay on the whitelist, so a saved image_model
keeps working untouched.

Also in this change:

- The rejection message is now built from the whitelist itself instead of a
  hand-written string that drifts whenever a model is added.
- codex_build.go uses the DefaultImageModel constant instead of hardcoding
  "gpt-image-2", so the inline chat path cannot drift from the shared default.
- Both new models join the isEditModel branch so reference images route to
  /images/edits.
- ParamField gains an optional labelKey, letting the dropdown read its label
  from i18n instead of showing English inside a translated UI. It passes a
  defaultValue, so a missing key degrades to the literal rather than breaking.
- Locale keys added to all five catalogs; ko was missing the whole
  mediaChain.imageModel* group.

Surface parity: the desktop UI is unchanged because sortable-provider-card
renders no param form and passes params through as Record<string, unknown>.
The CLI is unchanged because cmd/ exposes no image-model flag. The API contract
is unchanged because image_model lives in the free-form params map; only the
backend whitelist widened.

Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 10:31:48 +07:00
thotam 0c1ededc92 feat(webhooks): return per-call usage breakdown with provider/model/cost (#1421)
feat(webhooks): per-call usage breakdown with provider/model/cost (#1421)
2026-07-10 22:56:05 +07:00
thotam e95851e690 fix(pipeline): aggregate cache tokens across agent turns (#1419)
ThinkStage folded each turn's usage into state.Think.TotalUsage but summed
only prompt/completion/total/thinking tokens, dropping CacheReadTokens,
CacheCreationTokens and the PromptTokensIncludeCachedSegments flag. The
aggregated RunResult.Usage is the source for webhook usage responses and
usage_events analytics, so async multi-turn agent runs reported zero cache
tokens even when traces showed heavy prompt-cache use (cache is read
per-response for spans, but never summed into the aggregate).

Sum cache read/creation across turns and OR the include-cached flag (a
provider-level property, consistent across a run). Restores correct cache
accounting for webhook usage and usage_events without any envelope change.
2026-07-10 20:25:06 +07:00
thotam 84f06c0443 feat(web): restore agent Shares tab for per-user access sharing (#1360)
Approved by github-maintain bot. Clean UI restoration of agent Shares tab.
2026-07-05 13:56:24 +07:00
thotam dcf000ab6a fix(http): allow Operator to read providers and OAuth status (#1075) (#1359)
Browser pairing authenticates users as RoleOperator, but every provider
and OAuth route was wrapped in requireAuth(RoleAdmin). The /setup flow reads
GET /v1/providers, so paired browsers got 403 and hung on the setup page.

Split provider/OAuth route auth by operation:
- providers.go: GET list/get/models/status/embedding/codex-activity now use a
  readAuth wrapper (method-derived min → Viewer); POST/PUT/DELETE and
  reconnect/verify stay Admin. Responses already mask API keys and queries
  stay tenant-scoped, so no secrets are exposed.
- oauth.go: GET status/quota use readAuth (they return only connection state
  and quota, never tokens); start/callback/logout stay Admin (preserving the
  #450 admin-only-management decision for mutations).

Updates the two auth tests to assert the new contract: an Operator can read
providers/OAuth status (200) but still cannot mutate them (403).
2026-07-05 12:56:13 +07:00
thotam cab639d15e fix(migrate): use iofs source so migrations load on Windows (#1358)
`goclaw migrate up` (and every other migrate subcommand) failed on Windows
with `create migrator: failed to open source, "file:///D:/.../migrations":
open .: The filename, directory name, or volume label syntax is incorrect.`

golang-migrate's file source driver mis-parses absolute drive-letter file://
URLs; the drive-letter formatting in absoluteToFileURI produced a URL the
driver could not open. Replace the file:// URL with an iofs source over
os.DirFS, which uses native OS path handling and works identically on every
platform. Drop the now-unused absoluteToFileURI/migrationsSourceURL helpers
and their URL-shape tests, and add a DB-free regression test that opens the
real migrations directory via newMigrationSource().

Verified end-to-end on Windows: `migrate version` and `migrate up` now run.
2026-07-05 12:26:42 +07:00
thotam 25eaa0166f fix(backup): repair tenant backup/restore (config_secrets order, hooks registry, restore conn check) (#1357)
Closes #1076, #1338.

Fix A — tenant backup aborted with SQLSTATE 42703 "column id does not
exist" because exportQuery() hardcoded ORDER BY id. Several tenant-scoped
tables have composite PKs and no id column. Add a TableDef.OrderBy field,
honor it in exportQuery(), and set it for every id-less registry table:
config_secrets, agent_team_members, tenant_hook_budget, system_configs,
builtin_tool_tenant_configs, skill_tenant_configs, user_agent_profiles.

Fix B — hooks, tenant_hook_budget and webhook config were missing from
the backup registry, silently dropping their data on backup/restore. Add
hooks, tenant_hook_budget, webhooks (preserve) plus the hook_agents junction
(via ParentJoin through hooks, composite PK), and mark hook_executions and
webhook_calls as ephemeral in the skipped list.

Fix C — restoring on a fresh server failed with "N active DB connection(s)
detected" because the gateway's own pool connections were counted as active
clients. Tag pool connections with application_name='goclaw' (pg.OpenDB) and
exclude them in CheckActiveConnections; genuine external clients still block.

Adds unit + integration regression tests, including an export-over-every-
registered-table test that surfaced the additional id-less tables.
2026-07-05 11:56:36 +07:00
thotam 7475ff3d6c fix(vault): preserve team-scoped docs on team delete (#1356)
Deleting a team failed with SQLSTATE 23514 whenever it owned a
team-scoped vault document. The vault_documents.team_id FK is
ON DELETE SET NULL; when team_id became NULL the
vault_docs_team_null_scope_fix() trigger (migration 000043)
unconditionally set scope='personal'. Team docs have agent_id IS NULL,
so the resulting (personal, agent_id NULL) row violates the
vault_documents_scope_consistency CHECK and aborts the whole delete.

PostgreSQL: migration 000089 replaces the trigger function to pick a
valid target scope by ownership, and only rewrite genuinely team-scoped
rows so scope='custom' docs that merely carry a team_id are preserved:
  agent_id IS NOT NULL -> 'personal' (defensive: legacy dirty rows)
  agent_id IS NULL     -> 'shared'   (normal team docs)

SQLite: the same class of bug exists (FK SET NULL leaves scope='team'
with a NULL team_id and aborts on the CHECK), but SQLite fires no
trigger during the FK action, so a DB-level fix is impossible.
SQLiteTeamStore.DeleteTeam now converts team-scoped docs in the same
transaction before removing the team, scoped by tenant_id in the
tenant path to keep tenant isolation.

Adds regression tests on both engines and bumps RequiredSchemaVersion
to 89. SQLite needs no schema migration since the fix is in Go code.
2026-07-05 11:26:21 +07:00
thotam f87cdaf616 feat(webhooks): split webhook usage tokens by cache vs non-cache (#1345)
Surface cache-read and cache-creation input tokens (plus the
prompt_tokens_include_cached_segments flag) in webhook usage responses,
so callers can distinguish cached from non-cached input tokens.

The data already flowed end-to-end via providers.Usage and
agent.RunResult.Usage; the two webhook envelope structs (webhookLLMUsage
for the sync LLM webhook, callbackUsage for the async callback payload)
copied only 3 of the fields, dropping the cache data. Add the 3 cache
fields (mirroring providers.Usage JSON tags exactly, all omitempty) and
populate them. The async callbackUsage mapping is extracted into a pure
newCallbackUsage helper for unit testing.

Additive and non-breaking: omitempty keeps responses without caching
byte-identical. No schema change.
2026-07-04 23:26:49 +07:00
thotam cd5ad845f5 feat(webhooks): lease heartbeat to prevent duplicate async processing (#1275)
The async webhook worker double-processed long-running agent calls: the
90s stale-running sweep reclaimed rows whose agent ran longer than 90s
(no heartbeat), re-running the agent and duplicating MCP side-effects.

Add a last_heartbeat_at column + a Heartbeat lease-renewal method (CAS on
lease_token). While an agent runs, the worker renews the lease every 30s;
ReclaimStale now keys off last_heartbeat_at instead of started_at, so a
live run is never reclaimed while a dead worker is still recovered within
the stale window. If a run loses its lease (reclaimed), the heartbeat
cancels the run context so the agent stops immediately and writes no
further side-effects.

- migrations/000085 (PG) + SQLite schema v52 + RequiredSchemaVersion 85
- WebhookCallStore.Heartbeat (PG + SQLite impls); ClaimNext/ReclaimStale
  switched to last_heartbeat_at
- worker heartbeat goroutine + cancel-on-lease-loss; invokeAgent honors ctx
- store + worker tests
2026-06-24 16:10:38 +07:00
thotam e2ec3370dd feat(webhooks): stream provider responses for server-side runs to enable prompt caching (#1273)
Server-side webhook agent runs (sync, async, and admin test) used Stream:false,
so OpenAI-compatible routers that only cache streaming requests never populated
or served their prompt cache. Webhook runs paid full input-token price on every
turn even with a stable session and an identical multi-turn prefix, while WS chat
(Stream:true) got cache hits from the 3rd turn.

Add gateway.webhook_stream (default true, env GOCLAW_WEBHOOK_STREAM) and apply it
to all three run sites. ChatStream returns the fully assembled response, so the
payload returned to the caller is unchanged. Set to false to restore non-streaming.
2026-06-24 13:42:20 +07:00
thotam 06ef06ca37 feat(webhooks): paginate list + call-history endpoints (server + web UI) (#1268)
Add total-count pagination to the webhook admin endpoints and the web UI.

Store:
- WebhookStore/WebhookCallStore gain Count; WebhookListFilter gains
  IncludeRevoked + Query (PG + SQLite, parameterized, tenant-scoped)

API:
- GET /v1/webhooks and GET /v1/webhooks/{id}/calls return
  {items, total, limit, offset} with server-side search + revoked filtering

Web UI:
- server-driven list pager + search/revoked filter; call-history dialog uses
  the real total (fixes the full-page "has more" boundary bug)
- i18n pager labels (en/vi/zh)

Tests: store pagination integration test; mock stores implement Count.
No schema migration (read-only COUNT).
2026-06-24 10:39:24 +07:00
thotam a10464290d feat(webhooks): configurable agent run timeout (default 600s) (#1267)
Replace the hardcoded 30s webhook agent-run deadline with a configurable
timeout (default 600s, cap 3600s) for both the async worker and the
sync/test HTTP handler. Legitimate multi-step runs were dying at 30s
mid-tool-call; the same request over WebSocket completed fine.

- internal/webhooks/timeout.go: ResolveTimeoutSec helper (<=0 → 600s, cap 3600s)
- WorkerConfig.AsyncAgentTimeout (async worker) + WebhookLLMHandler.syncTimeout
  (sync + admin test), wired from config
- config keys gateway.webhook_{async,sync}_timeout_sec, also settable via
  GOCLAW_WEBHOOK_{ASYNC,SYNC}_TIMEOUT_SEC env (env overrides config)
- tests: timeout bounds + config file/env override
2026-06-24 09:22:06 +07:00
thotam 0ae55991bb feat(mcp): MCP OAuth 2.1 client for tool servers (#1196)
* feat(mcp): MCP OAuth 2.1 client — full implementation with tests

Implements a complete MCP OAuth 2.1 authorization flow for tool servers that
require user-delegated access, covering all layers from DB to UI.

- discovery.go: RFC 9728 protected-resource → RFC 8414 AS metadata → OIDC
  fallback chain with 5-min in-memory cache and InvalidateCache()
- dcr.go: RFC 7591 Dynamic Client Registration with response size guard
- flow.go: PKCE (S256) authorization code flow — StartFlow(), ExchangeCode(),
  ClientCredentials(), auto-cleanup of expired flows; carries AS issuer through
  PendingFlow for status display
- refresher.go: OAuthTokenProvider with in-memory token cache, automatic refresh
  on expiry, per-user vs global slot isolation, InvalidateCache/InvalidateServer

- migrations/000074 + SQLite schema: mcp_oauth_tokens with AES-256-GCM encrypted
  access/refresh tokens, partial unique index for global vs per-user rows,
  ON DELETE CASCADE from mcp_servers

- store.MCPOAuthTokenStore: Upsert, Get/GetUser, Delete/DeleteUser, and
  DeleteServerOAuthTokens (purge all rows for a server)
- PostgreSQL + SQLite implementations

- POST   /v1/mcp/oauth/start      — discovery + optional DCR + PKCE redirect URL;
  client_credentials completes server-side (no redirect) and returns completed=true
- GET    /v1/mcp/oauth/callback   — exchange code, persist token, publish WS event;
  payload built via json.Marshal (no reflected XSS via error_description)
- GET    /v1/mcp/oauth/status/{id}, DELETE /v1/mcp/oauth/token/{id} — admin-gated
- POST   /v1/mcp/oauth/discover/{id} — on-demand discovery probe
- All outbound calls go through the SSRF-safe client with pinned IPs

- pkg/protocol/mcp_events.go: EventMCPOAuthComplete routed only to the initiating
  user (admins in-tenant included); fail-closed across tenants

- getUserMCPTools() injects Authorization: Bearer from OAuthTokenProvider; on a
  401 for OAuth servers it purges the cached token so the next turn re-resolves

- handleUpdateServer purges all OAuth tokens (global + per-user), drops the
  refresher cache, and evicts the pool when a server's URL or OAuth config
  (client_id / endpoints / grant_type / scope / auth_type) changes — so the
  status UI and agent never use a token minted for the old resource/AS

- MCPOAuthDialog (WS-driven), unified user-credentials dialog, OAuth settings
  fields; handles the no-redirect client_credentials completion

- internal/mcp/oauth/*_test.go: discovery cache, PKCE, DCR, refresher
- internal/http/mcp_oauth_test.go + mcp_update_oauth_purge_test.go: routes, auth
  gating, WS event, purge-on-URL/OAuth-config-change
- tests/integration: store + encryption + tenant isolation, E2E start→callback,
  DeleteServerOAuthTokens
- internal/gateway/event_filter_test.go, internal/agent/loop_mcp_user_test.go

* fix(mcp): return 400 on OAuth callback with code but missing state

The callback handler rendered a 200 HTML page whenever code or state was
absent. An auth code WITH a missing state is a malformed / CSRF-risk
callback (state is the CSRF token), so reject that case with HTTP 400.
A bare hit with neither code nor state (user opening the URL directly),
provider errors, and exchange failures keep their 200 HTML popup page.

Adds a status code parameter to writeCallbackHTML. Fixes the
TestOAuthCallbackMissingState integration regression while keeping
TestHandleCallbackMissingCodeAndState (no params -> 200) green.

* fix(mcp): scope-based OAuth auth + honor manual OAuth endpoints

Addresses the two MCP/OAuth security-review findings.

Finding 1 — authorization. mcp_oauth_tokens is tenant-scoped, but
start/status/revoke were gated only by requireAuth(RoleAdmin), an RBAC
role check, not tenant membership, so a RoleAdmin caller could act on a
tenant they don't administer. A blanket requireTenantAdmin would have
broken per-user self-service, which the UI exposes (the per-user
MCPUserCredentialsDialog shows an "Authorize" button to regular users for
their own credentials). Instead mirror the existing per-user MCP
credentials model (resolveTargetUserID in mcp_user_credentials.go):
- start/status/revoke accept any authenticated user; each handler calls
  authorizeOAuthScope.
- a caller may manage their OWN per-user token (self-service); the
  global/server token (user_id="") and other users' tokens require
  tenant-admin (owner bypass), so a RoleAdmin that is not a tenant admin
  is rejected.
- discover stays admin-only (it only previews AS metadata for a server).
Add a TenantStore dependency. Tests cover self-service, on-behalf-of-
another (403), and global-by-non-tenant-admin (403).

Finding 2 — honor manual OAuth config end-to-end. The UI sent use_dcr /
auth_endpoint / token_endpoint and the update path fingerprinted them for
purge, but handleStart always discovered + DCR'd and ignored them. Now:
- use_dcr=false (a *bool, so legacy/absent stays discover+DCR) skips
  discovery/registration and uses the operator endpoints, SSRF-validated.
- token_endpoint is always required; auth_endpoint only for auth-code
  grants — client_credentials needs no authorization URL, matching the UI
  which hides that field for that grant.
- the refresher already refreshes against the stored token_endpoint and
  the callback persists it, so manual-mode tokens refresh correctly.
- oauthFingerprint includes use_dcr (nil normalized to true) so toggling
  DCR mode purges stale tokens.
- the web form only serializes manual endpoints when use_dcr is off.

Audited all MCP dialogs (form, global OAuth, per-user credentials, grants,
tools): OAuth dialogs handle completed/auth_url identically and read
config from stored server settings; runtime connect uses the stored token
via the refresher (no re-discovery).

Tests: manual auth-code + client_credentials endpoints, missing/SSRF
endpoints, and the full self/global/on-behalf authorization matrix.
2026-06-21 22:27:44 +07:00
thotamandGoon 21fb1f18ee feat(webhooks): add webhook management UI with delivery history & test (#1211)
* feat(webhooks): add webhook management UI with delivery history & test

Webhooks admin page on the web dashboard for managing inbound HTTP webhooks
(llm + message kinds), backed by new admin endpoints. No schema changes
(reuses migrations 000059-000061).

Backend (internal/http):
- GET  /v1/webhooks/{id}/calls           - paginated delivery history (trimmed
  DTO; status/limit/offset filters; ownership + tenant scoped)
- GET  /v1/webhooks/{id}/calls/{callId}  - full single-call detail (request
  payload, full response, callback URL, idempotency key, timestamps)
- POST /v1/webhooks/{id}/test            - server-side test invocation using the
  admin session (no secret); dispatched by kind via RunTest() on the llm/message
  handlers; message tester nil-guarded (403 on Lite)
Injects WebhookCallStore + SetTesters() in cmd/gateway_http_wiring.go, adds i18n
key webhook.message_test_requires_standard (en/vi/zh), and unit tests for
list/detail (filter, tenant isolation) and test (success/error/edition gate).

Frontend (ui/web):
- /webhooks admin-only page: list with revoked filter + badge, create/edit form
  (edition-gated message kind, Lite localhost_only lock), show-once secret dialog
  (create + rotate), test dialog, and delivery-history dialog with server-side
  pagination plus a click-through full call-detail dialog.
- use-webhooks hooks, query keys, types, routing, sidebar entry (Cable icon to
  distinguish from the existing event "Hooks" page), webhooks i18n namespace
  (en/vi/zh).

* fix(webhooks): validate admin message tests

---------

Co-authored-by: Goon <[email protected]>
2026-06-21 18:46:37 +07:00
thotam a5a853f461 feat(tools): image reference processing and native provider support (#1251)
* feat(tools): implement image reference processing and native provider support

- Support OpenAI image edits via both Multipart form-data and JSON payloads
- Automatically append reference image descriptions to prompt under [Reference Image Roles]
- Support downloading image URLs for Gemini native image generation
- Deduplicate reference images to optimize API request size
- Add unit tests for Codex, DashScope, MiniMax, BytePlus and local/remote path resolution

* fix(tools): SSRF-guard reference-image URL downloads in create_image

downloadImageBytes fetched caller-supplied ref_images[].url with a plain
http.Client and unbounded io.ReadAll — no SSRF validation, redirect
policy, or size cap, letting the gateway dial loopback/private/metadata
hosts or read arbitrarily large responses.

- Validate the URL via security.Validate and pin the resolved IP, then
  download through security.NewSafeClient (pinned dial, no redirects).
- Cap the response with a bounded read (refImageMaxBytes, 20 MB).
- Reject non-HTTP(S) reference URLs up front (file://, data:, gopher://)
  so provider-forwarded URLs stay HTTP(S)-only; document the trust
  boundary between gateway-side fetch and provider-forwarded URLs.
- Add regression tests: blocked loopback/private/metadata, unfollowed
  redirect, oversized response, and non-http(s) scheme rejection.
2026-06-21 18:27:51 +07:00
thotam 389640ae51 fix(agent): check nil td.Function to prevent panic on native tools (#1192)
* fix(agent): check nil td.Function to prevent panic on native tools

Fix nil pointer dereference panic when agent processes native tools (such as image_generation) which do not contain function schema wrappers.

Key changes:
- loop_pipeline_callbacks.go & think_stage.go: add checks for td.Function != nil before accessing td.Function.Name.
- loop_tool_filter.go & loop_history_toolnames.go: add safety checks when filtering and extracting tool names.
- anthropic_request.go & codex_build.go: verify t.Function != nil when formatting tools for provider API requests.

* test: add regression tests for native tools nil function panic

Add targeted regression tests to verify that:
- ThinkStage builds AllowedTools without panicking when mixed with native tools
- Loop's buildFilteredTools filters native tools correctly without panicking across all filters
- Anthropic request builder skips native tools safely
- Codex request builder handles native tools with nil Function fields gracefully
2026-06-21 18:04:07 +07:00