fix(proxy): forward stream_options.include_usage to OpenAI-compatible upstreams

Claude Code's status line and the session log showed 0 tokens for every
streaming response routed through the OpenAI-compatible proxy. The root
cause was that the upstream request did not include
`stream_options: { include_usage: true }`, so the provider (LiteLLM and
other OpenAI-compatible endpoints) never emitted the final SSE chunk
that carries `usage`. The proxy's stream parser was already wired to
forward that chunk into an Anthropic-shaped `message_delta` event but
had nothing to forward.

Set `stream_options.include_usage` whenever the incoming Anthropic
request is streamed. Non-streamed requests are unaffected.

Ref: https://platform.openai.com/docs/api-reference/chat-streaming
     (search 'include_usage')
This commit is contained in:
ReWiG committed 2026-06-19 21:45:47 +03:00
1 parent 0d856fc0ab
commit 2aaad43fd7
2 files changed
+45

No files matched your search

@@ -101,6 +101,9 @@ interface OpenAIMessage {
export interface ProxyOpenAIRequest {
model?: string;
stream: boolean;
stream_options?: {
include_usage: boolean;
};
reasoning_effort?: string;
reasoning?: {
enabled: boolean;
@@ -728,6 +731,16 @@ export class ProxyRequestTransformer {
? source.model.trim()
: undefined,
stream: source.stream === true,
// Ask the upstream provider to emit a final SSE chunk with token usage.
// OpenAI-compatible endpoints (and most proxies that follow the spec,
// including LiteLLM) only include `usage` in the streamed response when
// `stream_options.include_usage` is set on the request. Without this
// flag the proxy's stream parser still writes Anthropic-shaped events
// with `usage: { input_tokens: 0, output_tokens: 0 }`, so Claude Code's
// status line and the session log show 0 tokens even though the call
// was billed. See:
// https://platform.openai.com/docs/api-reference/chat-streaming
...(source.stream === true ? { stream_options: { include_usage: true } } : {}),
messages: coalesceMessages(allMessages),
max_tokens: asNumber(source.max_tokens),
temperature: asNumber(source.temperature),
@@ -115,4 +115,36 @@ describe('ProxyRequestTransformer', () => {
expect(result.stop).toEqual(['A']);
expect(result.metadata).toBeUndefined();
});
it('sets stream_options.include_usage for streaming requests so upstreams return token usage', () => {
const transformer = new ProxyRequestTransformer();
const result = transformer.transform({
stream: true,
messages: [{ role: 'user', content: 'ping' }],
});
expect(result.stream).toBe(true);
expect(result.stream_options).toEqual({ include_usage: true });
});
it('omits stream_options for non-streaming requests', () => {
const transformer = new ProxyRequestTransformer();
const result = transformer.transform({
stream: false,
messages: [{ role: 'user', content: 'ping' }],
});
expect(result.stream).toBe(false);
expect(result.stream_options).toBeUndefined();
});
it('omits stream_options when the upstream flag is absent', () => {
const transformer = new ProxyRequestTransformer();
const result = transformer.transform({
messages: [{ role: 'user', content: 'ping' }],
});
expect(result.stream).toBe(false);
expect(result.stream_options).toBeUndefined();
});
});