Files
Clark Cant 5ca433b8ac fix: persist mid-loop compaction to stop re-compaction loop and revive episodic
The v3 pipeline compacts session history mid-loop (prune_stage +
final-request guard) but only mutates the run's message buffer, never the
session store. Each turn reloads full history and re-compacts from scratch:
message_tokens climb 129k->156k across turns while every turn compacts back
down to ~60k. The lossy compaction differs per run, degrading the agent.

The same missing persistence stalls episodic memory: the cumulative
compaction count never advances, so the episodic worker's idempotency key
(sessionKey:count) is pinned and every cycle after the first is skipped.
Observed on live traffic: 8 run.completed since deploy, 0 new episodic.

Fixes, all reusing existing machinery (no new store methods, no migrations):

- Bug A: emitSessionCompleted reads cumulative GetCompactionCount (matching
  the legacy v2 path) instead of the per-run counter that resets to 0.
- Bug B/anti-loop: finalize passes state.Prune.MidLoopCompacted into
  maybeSummarize; under pressure it lowers the trigger to a unit-aligned
  threshold (compactionInputCap - overhead, same MaxRequestShare the guard
  uses) so the compaction is PERSISTED via the existing TruncateHistory +
  IncrementCompaction path. Defensive floor prevents over-compaction on
  pathological config; tool-result-only bloat still skips (history-only).
- Bug C: SourceID embeds the count (sessionKey:count) so the eventbus dedup
  key advances per compaction cycle instead of swallowing rapid same-session
  turns within the 5m TTL.

Tests: episodic compaction, maybe_summarize pressure, request budget.
go build (PG + sqliteonly), go vet, go test -race all green.
2026-07-25 00:39:02 +07:00

51 lines
1.9 KiB
Go

package consolidation
import (
"context"
"log/slog"
"github.com/google/uuid"
"github.com/nextlevelbuilder/goclaw/internal/config"
"github.com/nextlevelbuilder/goclaw/internal/store"
)
// withAgentRequestBudget wires the calling agent's configured request
// budget (context_window, max_tokens) into ctx so background LLM calls
// (episodic summarization, dreaming synthesis) pass the agent-only preflight
// guard instead of failing closed with an AgentBudgetWiringError.
//
// Background workers run off a session.completed / episodic.created event and
// do not carry a live agent Loop, so the budget cannot come from
// agent.WithAgentBudget. We load it straight from the agent row. If the store
// is unavailable or the row cannot be read, we fall back to the operator
// defaults rather than let a transient lookup miss silently kill memory
// consolidation — the guard's job is to bound requests to the agent window,
// and the defaults are the same window the agent would use unconfigured.
func withAgentRequestBudget(ctx context.Context, agents store.AgentCRUDStore, agentID uuid.UUID, purpose string) context.Context {
contextWindow := config.DefaultContextWindow
maxTokens := config.DefaultMaxTokens
if agents != nil && agentID != uuid.Nil {
if ag, err := agents.GetByIDUnscoped(ctx, agentID); err != nil {
slog.Warn("consolidation: agent budget lookup failed, using defaults",
"purpose", purpose, "agent", agentID, "err", err,
"context_window", contextWindow, "max_tokens", maxTokens)
} else {
if ag.ContextWindow > 0 {
contextWindow = ag.ContextWindow
}
if ag.MaxTokens > 0 {
maxTokens = ag.MaxTokens
}
}
} else {
slog.Warn("consolidation: no agent store or agent id, using default budget",
"purpose", purpose, "agent", agentID,
"context_window", contextWindow, "max_tokens", maxTokens)
}
ctx = store.WithAgentContextWindow(ctx, contextWindow)
ctx = store.WithAgentMaxTokens(ctx, maxTokens)
return ctx
}