The Ollama integration never applied a correct context window, so agents
with real prompts (20k-100k tokens) were rejected with HTTP 400
exceed_context_size_error against a 4096-token default.
Three root causes fixed:
1. FetchOllamaModelContext issued a GET to /api/show, which Ollama answers
with 405 (the endpoint is POST-only). Now POSTs {"model": "<name>"}.
2. The response parser expected a flat model_info.context_length, but a real
Ollama server namespaces the key by architecture (gemma4.context_length,
qwen3.context_length, ...). extractContextLength now matches "context_length"
or any "*.context_length" key.
3. num_ctx was resolved once at startup for a hardcoded "llama3.3" model and
never for the model an agent actually uses. Resolution now happens per
request for the real model inside OllamaProvider.resolveNumCtx, cached under
an RWMutex, with an explicit settings override winning and the fetched value
bounded by OllamaDefaultNumCtx so an enormous advertised window (Qwen3.5
reports 262144) cannot balloon the KV cache beyond VRAM.
Also classify Ollama's "exceed_context_size" 400 as a context-overflow error so
the pipeline's emergency-compaction+retry path (Issue 958) engages gracefully
instead of surfacing a raw error.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>