Claude Code sends the system prompt as the top-level `system` field and,
separately, sends skill/plugin listings as `role: "system"` entries inside
`messages` (#1459 made the transformer accept those). `transform()`
unconditionally prepends the top-level field, so once both are present the
OpenAI-compat payload ends up with two `system` messages that are not
adjacent. `coalesceMessages` only merges consecutive same-role messages and
explicitly skips `system`, so it cannot fix this.
Strict OpenAI-compatible backends (LiteLLM among them) reject that shape with:
400 A 'system' message can only appear at index 0 of the messages array.
Add `hoistSystemMessages`, run before `coalesceMessages`, which extracts every
`system` message in encounter order and reinserts a single merged one at
index 0. Content-preserving, no behavior change when at most one system
message is present.
Claude Code's status line and the session log showed 0 tokens for every
streaming response routed through the OpenAI-compatible proxy. The root
cause was that the upstream request did not include
`stream_options: { include_usage: true }`, so the provider (LiteLLM and
other OpenAI-compatible endpoints) never emitted the final SSE chunk
that carries `usage`. The proxy's stream parser was already wired to
forward that chunk into an Anthropic-shaped `message_delta` event but
had nothing to forward.
Set `stream_options.include_usage` whenever the incoming Anthropic
request is streamed. Non-streamed requests are unaffected.
Ref: https://platform.openai.com/docs/api-reference/chat-streaming
(search 'include_usage')
Sets undici headersTimeout/bodyTimeout to request_timeout+30s so the AbortController is the single authority on upstream request lifetime, preventing premature socket closes on slow self-hosted upstreams. Verified undici v5 ProxyAgent object-signature compat.
Some providers (e.g. Kimi, Anthropic-API mirrors) reject OpenAI-format
chat-completions requests and/or only accept requests from a recognized
coding-agent User-Agent (e.g. Claude Code, Roo Code, Kilo Code). The
OpenAI-compat proxy previously translated every profile's request to
OpenAI format and overwrote the User-Agent with a fixed sentinel,
which made these providers unreachable.
This change adds an opt-in Anthropic passthrough mode:
- New CCS_OPENAI_PROXY_PASSTHROUGH=1 env var on a profile opts it in.
- The base URL is auto-detected as Anthropic-style for known hosts
(api.kimi.com, api.minimax.com, api.anthropic.com) or any base URL
ending in /v1.
- In passthrough mode the proxy forwards the incoming Anthropic body
verbatim to the upstream /v1/messages endpoint, preserving the
original User-Agent (or x-stainless-user-agent) so coding-agent
provider checks pass.
- The Anthropic-format response is streamed back unchanged.
Adds:
- isAnthropicPassthroughProfile() + passthrough option on
resolveOpenAIChatCompletionsUrl/resolveOpenAIModelsUrl
- CCS_OPENAI_PROXY_PASSTHROUGH env var on OpenAICompatProfileConfig
- readRawBody() helper for the passthrough path
- Preserved User-Agent (or x-stainless-user-agent) on the upstream
request, falling back to CCS-OpenAI-Compat-Proxy/1.0
- Skip SSE response transformation in passthrough mode (upstream
already returns Anthropic-format bytes)
Tests:
- 9 new tests in upstream-url.test.ts covering auto-detection and the
passthrough URL contract
- 2 new tests in profile-router.test.ts covering the env var
Verified end-to-end against api.kimi.com: a request through the
modified proxy returned a real Kimi response (model kimi-k2p7-coding)
with the original claude-cli/2.1.170 User-Agent preserved.
Issue #1161. Sweeps 127 files to import from
src/config/config-loader-facade.ts instead of unified-config-loader or
utils/config-manager directly.
WRITE callers (32 files): replaced raw saveUnifiedConfig /
mutateUnifiedConfig / updateUnifiedConfig calls with the facade's
cache-coherent wrappers saveConfig / mutateConfig / updateConfig. This
fixes a latent stale-cache window where direct writes through the
underlying loader bypassed the facade's memoization.
READ callers (95 files): mechanical import-path migration only —
function names unchanged because the facade re-exports them. No
behavior change.
Also updated:
- tests/unit/utils/browser/browser-setup.test.ts (DI interface rename)
- src/management/checks/image-analysis-check.ts (dynamic import rename)
- src/web-server/health-service.ts (dynamic require rename)
- src/ccs.ts (path prefix fix from sweep script)
After sweep: zero raw write callers remain outside src/config/. Direct
imports of config-manager remain only for symbols not in the facade
(getConfigPath, getCcsDirSource, etc). Behavior unchanged; full suite
passes 1824/1824.
Out of scope: switching loadOrCreateUnifiedConfig() callers to
getCachedConfig() — needs per-callsite cache-safety analysis. Tracked
as follow-up.
Refs #1161
Wrap proxy server entry edge in withRequestContext so every inbound request
gets a requestId (reused from x-ccs-request-id header when valid UUID-ish,
freshly minted otherwise). messages-route emits 7 stages: intake / auth /
transform / route / dispatch / upstream / respond, each with latencyMs and
structured error metadata on failure.
CCS CLI entry (ccs.ts) wraps main() in runWithRequestId and emits
cli.command.start / complete / failed stages so command lifecycle is
correlatable end-to-end.
Refs #1141, #1138
- reject pending tool_result layouts that cannot be translated without reordering user content
- keep interleaved GLMT tool_use blocks open until finalization instead of stopping early
- cover leading/interleaved tool_result regressions and interleaved streaming tool fragments
- enforce strict tool_result ordering and pairing against assistant tool_use ids
- reject tool_result image payloads that cannot map to OpenAI tool messages
- preserve raw tool schemas on the /v1/messages proxy path instead of silently tightening them
- forward Anthropic tool_choice semantics and cover adaptive routing plus upstream payload checks