mirror of
https://github.com/tiennm99/goclaw.git
synced 2026-10-11 03:13:24 +00:00
The v3 pipeline compacts session history mid-loop (prune_stage + final-request guard) but only mutates the run's message buffer, never the session store. Each turn reloads full history and re-compacts from scratch: message_tokens climb 129k->156k across turns while every turn compacts back down to ~60k. The lossy compaction differs per run, degrading the agent. The same missing persistence stalls episodic memory: the cumulative compaction count never advances, so the episodic worker's idempotency key (sessionKey:count) is pinned and every cycle after the first is skipped. Observed on live traffic: 8 run.completed since deploy, 0 new episodic. Fixes, all reusing existing machinery (no new store methods, no migrations): - Bug A: emitSessionCompleted reads cumulative GetCompactionCount (matching the legacy v2 path) instead of the per-run counter that resets to 0. - Bug B/anti-loop: finalize passes state.Prune.MidLoopCompacted into maybeSummarize; under pressure it lowers the trigger to a unit-aligned threshold (compactionInputCap - overhead, same MaxRequestShare the guard uses) so the compaction is PERSISTED via the existing TruncateHistory + IncrementCompaction path. Defensive floor prevents over-compaction on pathological config; tool-result-only bloat still skips (history-only). - Bug C: SourceID embeds the count (sessionKey:count) so the eventbus dedup key advances per compaction cycle instead of swallowing rapid same-session turns within the 5m TTL. Tests: episodic compaction, maybe_summarize pressure, request budget. go build (PG + sqliteonly), go vet, go test -race all green.
32 lines
1.2 KiB
Go
32 lines
1.2 KiB
Go
package pipeline
|
|
|
|
import "github.com/nextlevelbuilder/goclaw/internal/providers"
|
|
|
|
// InputContextTokens returns the provider-reported input occupancy for one
|
|
// request. Anthropic-style usage reports cached segments separately, while
|
|
// OpenAI-style prompt tokens already include them.
|
|
func InputContextTokens(usage providers.Usage) int {
|
|
if usage.PromptTokensIncludeCachedSegments {
|
|
return usage.PromptTokens
|
|
}
|
|
return usage.PromptTokens + usage.CacheReadTokens + usage.CacheCreationTokens
|
|
}
|
|
|
|
// AddUsage accumulates provider usage while preserving telemetry fields that
|
|
// are easy to drop when callers hand-roll partial sums.
|
|
func AddUsage(dst *providers.Usage, src providers.Usage) {
|
|
if dst == nil {
|
|
return
|
|
}
|
|
dst.PromptTokens += src.PromptTokens
|
|
dst.CompletionTokens += src.CompletionTokens
|
|
dst.TotalTokens += src.TotalTokens
|
|
dst.CacheCreationTokens += src.CacheCreationTokens
|
|
dst.CacheReadTokens += src.CacheReadTokens
|
|
dst.PromptTokensIncludeCachedSegments = dst.PromptTokensIncludeCachedSegments || src.PromptTokensIncludeCachedSegments
|
|
dst.ThinkingTokens += src.ThinkingTokens
|
|
dst.RequestCount += src.RequestCount
|
|
dst.ImageCount += src.ImageCount
|
|
dst.WebSearchCount += src.WebSearchCount
|
|
}
|