mirror of
https://github.com/tiennm99/goclaw.git
synced 2026-10-11 12:18:59 +00:00
The v3 pipeline compacts session history mid-loop (prune_stage + final-request guard) but only mutates the run's message buffer, never the session store. Each turn reloads full history and re-compacts from scratch: message_tokens climb 129k->156k across turns while every turn compacts back down to ~60k. The lossy compaction differs per run, degrading the agent. The same missing persistence stalls episodic memory: the cumulative compaction count never advances, so the episodic worker's idempotency key (sessionKey:count) is pinned and every cycle after the first is skipped. Observed on live traffic: 8 run.completed since deploy, 0 new episodic. Fixes, all reusing existing machinery (no new store methods, no migrations): - Bug A: emitSessionCompleted reads cumulative GetCompactionCount (matching the legacy v2 path) instead of the per-run counter that resets to 0. - Bug B/anti-loop: finalize passes state.Prune.MidLoopCompacted into maybeSummarize; under pressure it lowers the trigger to a unit-aligned threshold (compactionInputCap - overhead, same MaxRequestShare the guard uses) so the compaction is PERSISTED via the existing TruncateHistory + IncrementCompaction path. Defensive floor prevents over-compaction on pathological config; tool-result-only bloat still skips (history-only). - Bug C: SourceID embeds the count (sessionKey:count) so the eventbus dedup key advances per compaction cycle instead of swallowing rapid same-session turns within the 5m TTL. Tests: episodic compaction, maybe_summarize pressure, request budget. go build (PG + sqliteonly), go vet, go test -race all green.
51 lines
1.9 KiB
Go
51 lines
1.9 KiB
Go
package consolidation
|
|
|
|
import (
|
|
"context"
|
|
"log/slog"
|
|
|
|
"github.com/google/uuid"
|
|
"github.com/nextlevelbuilder/goclaw/internal/config"
|
|
"github.com/nextlevelbuilder/goclaw/internal/store"
|
|
)
|
|
|
|
// withAgentRequestBudget wires the calling agent's configured request
|
|
// budget (context_window, max_tokens) into ctx so background LLM calls
|
|
// (episodic summarization, dreaming synthesis) pass the agent-only preflight
|
|
// guard instead of failing closed with an AgentBudgetWiringError.
|
|
//
|
|
// Background workers run off a session.completed / episodic.created event and
|
|
// do not carry a live agent Loop, so the budget cannot come from
|
|
// agent.WithAgentBudget. We load it straight from the agent row. If the store
|
|
// is unavailable or the row cannot be read, we fall back to the operator
|
|
// defaults rather than let a transient lookup miss silently kill memory
|
|
// consolidation — the guard's job is to bound requests to the agent window,
|
|
// and the defaults are the same window the agent would use unconfigured.
|
|
func withAgentRequestBudget(ctx context.Context, agents store.AgentCRUDStore, agentID uuid.UUID, purpose string) context.Context {
|
|
contextWindow := config.DefaultContextWindow
|
|
maxTokens := config.DefaultMaxTokens
|
|
|
|
if agents != nil && agentID != uuid.Nil {
|
|
if ag, err := agents.GetByIDUnscoped(ctx, agentID); err != nil {
|
|
slog.Warn("consolidation: agent budget lookup failed, using defaults",
|
|
"purpose", purpose, "agent", agentID, "err", err,
|
|
"context_window", contextWindow, "max_tokens", maxTokens)
|
|
} else {
|
|
if ag.ContextWindow > 0 {
|
|
contextWindow = ag.ContextWindow
|
|
}
|
|
if ag.MaxTokens > 0 {
|
|
maxTokens = ag.MaxTokens
|
|
}
|
|
}
|
|
} else {
|
|
slog.Warn("consolidation: no agent store or agent id, using default budget",
|
|
"purpose", purpose, "agent", agentID,
|
|
"context_window", contextWindow, "max_tokens", maxTokens)
|
|
}
|
|
|
|
ctx = store.WithAgentContextWindow(ctx, contextWindow)
|
|
ctx = store.WithAgentMaxTokens(ctx, maxTokens)
|
|
return ctx
|
|
}
|