Cron job execution set the tenant ID on its context but not the tenant
slug. Tenant-scoped filesystem paths (skills-store, workspace, media via
config.TenantScopedDir) key off the slug and fall back to an id-based
path when it is absent — a different directory than where HTTP/WS skill
upload materialized the files (which sets the slug). As a result a cron
agent turn in a non-master tenant saw NONE of its tenant's managed
skills: skill_search returned 0 results and the agent, unable to run the
skill, produced an ungrounded answer.
Add cronTenantContext() which resolves the tenant slug via TenantStore
and sets both WithTenantID and WithTenantSlug. Master tenant and
nil-store/lookup-failure paths fall back to id-only (prior behavior).
Thread TenantStore into makeCronJobHandler and runCommandCronJob.
Tested: added unit tests for cronTenantContext (slug injected for
non-master; master skips lookup; nil store and lookup error fall back to
id-only). Verified end-to-end on a live tenant: before, a daily-agenda
cron guessed an empty day; after, it read the real event from the DB.
Note: other background executors that build a context from a tenant ID
(e.g. heartbeat) likely share this gap and are worth an audit.
* feat(cron): deterministic command payloads (run a shell command, no LLM)
Cron jobs always run an agent turn today, so deterministic work (health
probes, backups, syncs) pays model tokens on every fire. This adds a
"command" payload kind that runs a shell command directly in the gateway
process with zero model tokens, mirroring openclaw's command cron.
- store: CronPayload.Command (*CronCommandSpec — argv/cwd/env/input/
timeouts/output cap). Persists in the existing payload JSON blob, so
there is NO migration and no schema version bump.
- internal/cronexec: in-process runner with wall-clock + no-output
timeouts, per-stream output capping, and process-group termination so a
timed-out command's forked children are also killed.
- gateway_cron handler: command jobs run in-process and deliver stdout on
success (honoring the NO_REPLY sentinel). A non-zero exit / timeout
returns an error so the run is recorded as error and retried per
cron.max_retries; failures are NOT delivered, mirroring the agent path
(only successful output is announced — no channel spam).
- surfaces: cron.create RPC, the agent `cron` tool, and a new
`goclaw cron create` CLI all accept command payloads.
- security: gated by cron.command_enabled (default false). Commands run
with the gateway process's privileges, so the feature is opt-in per
gateway; when disabled the RPC and tool reject command payloads and the
handler refuses to run them.
- i18n (en/vi/zh), docs (08-scheduling-cron.md), and tests for the runner
and the handler command path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cron): gate command payloads on the update surfaces too
handleUpdate (RPC + agent tool) passed CronJobPatch.Command straight to
UpdateJob, which switches the payload to command kind for any non-nil
Command — without the command_enabled gate or ValidateCronCommandSpec that
create enforces. A normal job could therefore be mutated into a command job
(or persisted with an invalid spec, e.g. empty argv) on a gateway where
command cron is disabled, breaking the disabled-gateway contract.
Both update surfaces now require cron.command_enabled and validate the spec
before UpdateJob, matching create. The agent tool parses the command via the
same path as add and drops the raw keys so a shell-string command can't break
the generic patch unmarshal. Regression tests added for RPC and tool update
(command disabled + invalid argv), plus a positive enabled-valid case.
Addresses review feedback from @mrgoonie on #1279.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>