diff --git a/docs/research/opencode-auto-compact.md b/docs/research/opencode-auto-compact.md new file mode 100644 index 0000000..d1646d2 --- /dev/null +++ b/docs/research/opencode-auto-compact.md @@ -0,0 +1,358 @@ +# OpenCode CLI: how auto-compact actually decides to trigger + +**Date:** 2026-09-03 +**Scope:** Resolves haylan/LLM-Server issue #27 (part of #26) — the config key(s), trigger +threshold/formula, reasoning-token accounting, and per-model tunability of OpenCode's +context-compaction behavior, and whether any of it is specific to hosted providers vs. the +`@ai-sdk/openai-compatible` path this repo's `docs/research/opencode-cli-setup.md` documents for +llama.cpp/litellm-style backends. + +**Freshness note:** `opencode-cli-setup.md` (dated 2026-08-24) does not cover compaction at all +beyond one sentence ("§5: OpenCode's `limit.context`/`limit.output` fields ... don't change what the +server actually accepts ... may miscalculate when to compact/summarize"). This document supersedes +that gap. It is based on: +1. Live docs fetches from https://opencode.ai/docs/ (2026-09-03), and +2. A fresh `git clone` of https://github.com/anomalyco/opencode at commit + `b578b7261fc9ec4917fe272df5cc4bd8a056cd5d` (2026-09-03T09:47:21+08:00 — same day as this + research), `package.json` version `1.18.27`. Source-code claims below are cited by file path in + that clone and are the **highest-confidence source in this doc** — they're what actually ships, + not a doc description or a third-party claim. + +## Confidence scheme + +- **High** — read directly from the current source code in the repo, or a docs page quoted + verbatim. +- **Medium** — inferred from source code behavior that isn't spelled out in a single line/comment + (i.e., I traced call sites to confirm it, rather than reading one authoritative line). +- **Low** — plausible but not directly confirmed in the sources checked; flagged as an open + question. + +--- + +## Verdict summary (answers to the four questions asked in #27) + +| Question | Answer | Confidence | +|---|---|---| +| Config key(s) controlling threshold/behavior | Single global top-level `compaction` object in `opencode.json`: `auto`, `prune`, `reserved`, `tail_turns`, `preserve_recent_tokens`. **No `threshold` percentage key exists.** | High | +| Is it a hardcoded percentage of `limit.context`? | **No.** It's `usedTokens >= (contextLimit − reservedBuffer)`, where `reservedBuffer` defaults to `min(20_000, min(model.limit.output, 32_768) or default)`, i.e. compaction reserves room for one more max-size reply, not a flat 75%/95% cutoff. Some closed GitHub feature requests describe it as "hardcoded 75%" — that claim is **not what the current source does** (see §2). | High (source), contradicts a stale community claim (see §2) | +| Does compaction count reasoning/`reasoning_content` tokens? | **Yes, in practice**, via the provider's `usage.total_tokens` (which for llama.cpp/litellm includes every generated token, reasoning or not) — but the token bookkeeping OpenCode itself derives (`tokens.output`, `tokens.reasoning`) explicitly **splits reasoning out of `output`**, and the compaction trigger's own fallback arithmetic (used only if the provider omits `total_tokens`) **omits `tokens.reasoning` entirely**. See §3 for the exact mechanism and the one edge case where reasoning tokens could be undercounted. | High (source), Medium (edge-case behavior when `total_tokens` is absent) | +| Per-model or single global behavior? | **Global only.** The `compaction` block is a top-level config key, not nested under `provider..models.` or any per-model schema. It cannot be disabled or tuned for one model while enabled for another. The *indirect* lever is each model's own `limit.context`/`limit.output` (which you already set per-model for custom providers), since those numbers feed the same global formula per-model. Multiple GitHub feature requests (#11314, #11930, #8140, #16375) ask for per-model/per-agent configurability; all are open or closed-not-planned as of this check. | High | + +--- + +## 1. The config surface (verbatim from source + docs) + +Global (or project) `opencode.json`: + +```json +{ + "compaction": { + "auto": true, + "prune": false, + "reserved": 20000, + "tail_turns": null, + "preserve_recent_tokens": null + } +} +``` + +Field descriptions, quoted verbatim from the config schema +(`packages/core/src/v1/config/config.ts`, lines ~149–166 in the cloned repo): + +- `auto` — *"Enable automatic compaction when context is full (default: true)"* +- `prune` — *"Enable pruning of old tool outputs (default: false)"* +- `tail_turns` — *"Maximum number of recent user turns, including their following + assistant/tool responses, to keep verbatim during compaction. By default retention is limited + only by the preserved token budget."* +- `preserve_recent_tokens` — *"Maximum number of tokens from recent turns to preserve verbatim + after compaction"* +- `reserved` — *"Token buffer for compaction. Leaves enough window to avoid overflow during + compaction."* + +Sources: +- https://opencode.ai/docs/config/ (fetched 2026-09-03) confirms `auto`/`prune`/`reserved` with the + same defaults and descriptions; the docs page does **not** mention `tail_turns` or + `preserve_recent_tokens` — those two are documented only in the source schema, not (yet) on the + public docs page. **Confidence: High** on all five keys existing and their defaults; the + docs-vs-source gap on the last two is itself notable (docs page is behind the schema). +- `packages/core/src/v1/config/config.ts` (cloned repo, ~line 149) — schema + descriptions. + +There is no `threshold`, `percent`, or similarly-named key anywhere in the config schema. I grepped +`packages/core/src/v1/config/` and `packages/core/src/config.ts` for `compaction` and found only the +struct above — no percentage field exists in the current schema (High confidence; direct grep of +current source). + +**Not nested under provider/model.** The `compaction` key sits at the top level of `opencode.json`, +a sibling of `provider`, `model`, `agent`, etc. — not inside +`provider..models.` (the block this repo's `opencode-cli-setup.md` §3 documents for +declaring the llama.cpp provider). Confirmed by reading the full top-level config struct in +`packages/core/src/v1/config/config.ts` (~lines 95–170): `compaction` is a direct sibling of +`provider`, not a child of it. + +There *is* a separate, easily-confused concept: `agent.compaction` (also in that same file, ~line +105) — this only lets you assign a **different agent/model to perform the summarization step +itself** (e.g. run the compaction LLM call on a cheaper model), not a per-model *threshold* +override. It does not change when compaction triggers. + +## 2. The actual trigger formula (not a flat percentage) + +Live code, `packages/opencode/src/session/overflow.ts` (the shipped/default compaction-trigger +path — see "which code path ships" note at the end of this section): + +```ts +const COMPACTION_BUFFER = 20_000 + +export function usable(input: { cfg: ConfigV1.Info; model: Provider.Model; outputTokenMax?: number }) { + const context = input.model.limit.context + if (context === 0) return 0 + + const reserved = + input.cfg.compaction?.reserved ?? + Math.min(COMPACTION_BUFFER, ProviderTransform.maxOutputTokens(input.model, input.outputTokenMax)) + return input.model.limit.input + ? Math.max(0, input.model.limit.input - reserved) + : Math.max(0, context - ProviderTransform.maxOutputTokens(input.model, input.outputTokenMax)) +} + +export function isOverflow(input: { + cfg: ConfigV1.Info + tokens: SessionV1.Assistant["tokens"] + model: Provider.Model + outputTokenMax?: number +}) { + if (input.cfg.compaction?.auto === false) return false + if (input.model.limit.context === 0) return false + + const count = + input.tokens.total || input.tokens.input + input.tokens.output + input.tokens.cache.read + input.tokens.cache.write + return count >= usable(input) +} +``` + +In plain terms: +- Compaction triggers when the running token count (`count`) reaches or exceeds + `usable = contextLimit − reservedBuffer` (or, if the model declares a separate + `limit.input`, `usable = limit.input − reservedBuffer` instead). +- `reservedBuffer` defaults to `min(20_000, maxOutputTokens)`, where `maxOutputTokens = min(model.limit.output, 32_768) || 32_768` + (`OUTPUT_TOKEN_MAX` constant, `packages/opencode/src/provider/transform.ts` — `maxOutputTokens()` + at ~line 1468). It can be overridden with `compaction.reserved`. +- If `model.limit.context` is `0` or unset, compaction is **silently disabled entirely** — this + matters for custom `@ai-sdk/openai-compatible` providers where `limit.context` is a value the + user must supply by hand (per `opencode-cli-setup.md` §3); omit it, and auto-compact never fires + for that model. +- Compaction is disabled outright if `compaction.auto === false`. + +This is **not** a fixed 75%/90%/95%-of-context cutoff. It's `context − reservedOutputBuffer`, which +in practice usually lands somewhere in the 80–95%+ range depending on the model's own +`limit.output` relative to `limit.context` — but it's a token-count subtraction, not a percentage +multiplication, and there is no config key to set a percentage. + +**On the "hardcoded 75%" claim**: a closed GitHub feature request +(https://github.com/anomalyco/opencode/issues/11314, "Feature Request: Configurable Context +Compaction Threshold") asserts *"Currently, OpenCode triggers context compaction at a hardcoded 75% +threshold of a model's context window."* That is a **user's claim in a feature-request issue, not a +maintainer statement or a source citation**, and it does not match the formula actually in the +current source (which is a token-count subtraction tied to output-buffer size, not a flat +percentage). Treat that issue as evidence that *users perceive/want configurability*, not as an +accurate description of the mechanism. A second, similarly-shaped request +(https://github.com/anomalyco/opencode/issues/11930) instead describes the current default as "100% +threshold" — the two community reports disagree with each other, which is itself a signal neither +is a reliable description of internals; the source code is what should be trusted here. +**Confidence: High** on the source-code formula; **High** that the 75%/100% claims in those two +issues are user speculation, not verified facts — both are closed as "not planned" with no +maintainer confirmation of the underlying mechanism in the content fetched. + +**Which code path ships**: the repo contains two parallel implementations — the one quoted above +(`packages/opencode/src/session/overflow.ts` + `packages/opencode/src/session/compaction.ts`, +using `ConfigV1`/`SessionV1` types) and a newer, structurally different one +(`packages/core/src/session/compaction.ts` + `packages/core/src/session/runner/llm.ts`, using +`Config.Entry`/`SessionMessage` types, with its own `compactIfNeeded`/`compactAfterOverflow` +functions and a `buffer` config field instead of `reserved`). Tracing the call path: the newer +runner is only reached when `flags.experimentalNativeLlm` is true, gated by the +`OPENCODE_EXPERIMENTAL_NATIVE_LLM` env var (`packages/opencode/src/effect/runtime-flags.ts`, +`packages/opencode/src/session/llm.ts` line ~226) — off by default. **The `overflow.ts`/V1 path +quoted above is what ships by default in v1.18.27.** If this repo's OpenCode setup ever sets +`OPENCODE_EXPERIMENTAL_NATIVE_LLM=1` (or a future release flips the default), the newer engine's +`compactIfNeeded()` uses a structurally similar but not identical check: +`estimatedPromptTokens <= context − max(output, config.buffer)` — same shape (context minus a +reserved buffer, no percentage), different field name (`buffer` vs `reserved`). **Confidence: +High** on which path ships by default; **Medium** on exact behavior differences if the experimental +flag is ever turned on, since I did not exhaustively diff every line of the newer engine. + +## 3. Reasoning/thinking-token accounting — the Qwen3.8 angle + +This repo's `litellm-config.yaml` flags that Qwen3.8-27B (`qwen3.8-27b-local`) spends generation +tokens on `reasoning_content` before `content`: + +```yaml +# Qwen3 is a reasoning model — it spends output tokens on +# reasoning_content before ever writing content. ... +max_tokens: 16384 +``` +(`/home/haylan/Projects/LLM-Server/litellm-config.yaml`, `qwen3.8-27b-local` model block.) + +OpenCode's own token bookkeeping (`packages/opencode/src/session/session.ts`, ~lines 340–375, the +function that turns a provider's raw `usage` object into the `tokens` struct stored on each +assistant message) does this: + +```ts +const inputTokens = safe(input.usage.inputTokens ?? 0) +const outputTokens = safe(input.usage.outputTokens ?? 0) +const reasoningTokens = safe(input.usage.reasoningTokens ?? 0) +... +const total = input.usage.totalTokens + +const tokens = { + total, + input: adjustedInputTokens, + output: safe(outputTokens - reasoningTokens), // reasoning is subtracted OUT of "output" + reasoning: reasoningTokens, // tracked as its own field + cache: { write: cacheWriteInputTokens, read: cacheReadInputTokens }, +} +``` + +And the compaction trigger (`overflow.ts`, quoted in §2) computes: + +```ts +const count = input.tokens.total || input.tokens.input + input.tokens.output + input.tokens.cache.read + input.tokens.cache.write +``` + +Two things follow: + +1. **`tokens.reasoning` is never added back in** by the fallback sum + (`input + output + cache.read + cache.write`) — that expression has no `+ reasoning` term. If + the compaction trigger ever fell back to this sum (i.e., the provider didn't return + `totalTokens`), reasoning tokens spent on `reasoning_content` would be **excluded** from the + overflow calculation, undercounting real context usage. +2. **In the normal case, `total` is used instead of the fallback sum**, and `total = usage.totalTokens` + straight from the provider's raw response — computed by the provider/AI-SDK *before* OpenCode + splits `outputTokens` into `output`/`reasoning`. Since the AI SDK's OpenAI-compatible adapter (and + litellm/llama.cpp underneath it) counts every generated token — reasoning and content alike — as + part of `completion_tokens`/`total_tokens`, `total` **does include reasoning tokens** in this + normal path. So in practice, for a llama.cpp/litellm backend that reports `usage.total_tokens` on + every response (the OpenAI chat-completions spec requires this field), **reasoning tokens are + accounted for in the compaction trigger via `total`, not via the explicit `reasoning` field.** + +The edge case that would matter for this repo: if litellm or llama.cpp's OpenAI-compatible endpoint +ever omitted `usage.total_tokens` from a response (malformed/incomplete usage block — this has +happened with some llama.cpp server versions/flags), OpenCode's fallback sum would silently +undercount by the full `reasoning` amount, delaying compaction past the point it should have +triggered and increasing risk of a hard `context_length_exceeded` — exactly the failure mode +reported by an unrelated user in +https://github.com/anomalyco/opencode/issues/8089 ("Auto-compaction enabled by default, but +context_length_exceeded errors still occur in agent workflows"), though that issue's cause was not +confirmed to be this specific gap (it involved OpenAI's GPT-5.2 and multi-agent/subagent workflows, +not a local llama.cpp backend, and the issue thread contains no maintainer diagnosis of root cause +in the content fetched). + +**Confidence: High** on the source-code mechanics described (the `output = outputTokens − +reasoningTokens` split, the `total || sum` fallback, and the missing `+ reasoning` term in the +fallback). **Medium** on whether this repo's specific llama.cpp/litellm stack reliably returns +`usage.total_tokens` on every response for the `qwen3.8-27b-local` model — this was not verified +against a live request/response in this research pass (would need an empirical check: hit +`http://localhost:8080/v1/chat/completions` directly or via the litellm proxy and inspect the +`usage` block of an actual reasoning response). Recommend that empirical check as a fast follow if +this matters operationally. + +## 4. Per-model tunability + +Confirmed absent, both from the schema (§1) and from community feature requests asking for exactly +this and not getting it: + +- https://github.com/anomalyco/opencode/issues/11314 — "Feature Request: Configurable Context + Compaction Threshold" — requests a `compaction.threshold` with "optional per-model overrides." + Closed as not planned (per WebFetch of the issue). +- https://github.com/anomalyco/opencode/issues/11930 — "Feature: Configurable compaction threshold + and model (global + per-model)" — explicitly requests global **and** per-model threshold config. + Closed as not planned, no maintainer reply visible in the content fetched. +- https://github.com/anomalyco/opencode/issues/8140 — "Feature Request: Configurable context limit + and auto-compaction threshold" — same theme (title only confirmed via search; not individually + fetched in this pass). +- https://github.com/anomalyco/opencode/issues/16375 — "[FEATURE]: Per-agent compaction config + (disable compaction for specific agents)" — same theme, per-agent instead of per-model (title only + confirmed via search; not individually fetched in this pass). + +All four are open/closed-not-planned as of 2026-09-03 — i.e., **as of this check, none of this has +shipped**: compaction remains a single global on/off + buffer-size knob, with no per-model or +per-agent threshold override. **Confidence: High** that the feature doesn't exist in the schema +(direct source read); **Medium** on the exact current status of #8140/#16375 specifically since only +their titles were confirmed via search results, not their full issue bodies. + +The one *indirect* per-model lever that does exist: since `usable()` (§2) reads +`input.model.limit.context` / `input.model.limit.input` / `input.model.limit.output` — all +per-model fields already documented in `opencode-cli-setup.md` §3/§5 for custom providers — setting +those numbers differently per model in the `provider..models..limit` block changes +where that model's compaction fires, without needing a dedicated per-model compaction key. This +is bookkeeping-hint tuning, not a first-class "compaction threshold" feature. + +## 5. Is any of this specific to hosted/built-in providers vs. `@ai-sdk/openai-compatible`? + +**No.** The entire trigger path (`overflow.ts`, `compaction.ts`) operates only on `Provider.Model` +(a normalized model descriptor with `limit.context`/`limit.input`/`limit.output`) and the message +`tokens` struct built from the SDK's generic `usage` object (`session.ts`, §3) — nothing in the +compaction code branches on `model.api.npm` or provider identity. The mechanism is provider-agnostic +by construction: any provider adapter that populates `usage` (inputTokens/outputTokens/totalTokens) +and any model entry that has a nonzero `limit.context` gets the same compaction behavior, including +a hand-declared `@ai-sdk/openai-compatible` provider block like this repo's `llamacpp` provider in +`opencode-cli-setup.md` §3. **Confidence: High** — read directly from the trigger/accounting source, +which takes no provider-specific branch. + +The one place this repo needs to be careful about, restated from §2: for a custom +`@ai-sdk/openai-compatible` provider, `limit.context` (and ideally `limit.output`) must be set by +hand in `opencode.json` to match the real `--ctx-size` the llama.cpp container is launched with — if +left unset (`limit.context` defaults to `0` for an unrecognized custom model), compaction is +silently disabled for that model rather than silently misfiring. + +## Sources consulted (primary) + +- https://opencode.ai/docs/ — nav/sitemap fetch (2026-09-03); confirms no dedicated "context + management"/"compaction" page exists in the current docs nav — it lives only in the Config + reference. +- https://opencode.ai/docs/config/ — fetched 2026-09-03; source of the `compaction.auto/prune/reserved` + descriptions and defaults quoted in §1. +- https://opencode.ai/docs/models/ — fetched 2026-09-03; confirms no compaction/context-limit + content on that page (only `reasoningEffort`/`thinking` keys, unrelated to compaction). +- `github.com/anomalyco/opencode` @ `b578b7261fc9ec4917fe272df5cc4bd8a056cd5d` (cloned 2026-09-03, + `package.json` version `1.18.27`) — primary source for all source-code claims: + - `packages/opencode/src/session/overflow.ts` — trigger formula (`usable`/`isOverflow`), §2 + - `packages/opencode/src/session/compaction.ts` — shipped compaction service, tail-turn selection, + pruning, §2/§4 + - `packages/opencode/src/session/session.ts` (~lines 340–375) — `usage` → `tokens` mapping, + reasoning-token split, §3 + - `packages/core/src/v1/config/config.ts` (~lines 95–170) — config schema, field descriptions, §1 + - `packages/core/src/config/compaction.ts`, `packages/core/src/session/compaction.ts`, + `packages/core/src/session/runner/llm.ts` — the experimental/newer compaction engine, gated + behind `OPENCODE_EXPERIMENTAL_NATIVE_LLM`, §2 + - `packages/opencode/src/effect/runtime-flags.ts`, `packages/opencode/src/session/llm.ts` (~line + 226) — confirms which engine ships by default, §2 + - `packages/opencode/src/provider/transform.ts` (~line 1468, `maxOutputTokens`) — output-buffer + sizing used in the reserved-token default, §2 +- https://github.com/anomalyco/opencode/issues/11314 — "Configurable Context Compaction Threshold" + (closed, not planned) — source of the "hardcoded 75%" community claim, §2/§4 +- https://github.com/anomalyco/opencode/issues/11930 — "Configurable compaction threshold and model + (global + per-model)" (closed, not planned) — source of the conflicting "100% threshold" claim, + §2/§4 +- https://github.com/anomalyco/opencode/issues/8089 — "Auto-compaction enabled by default, but + context_length_exceeded errors still occur in agent workflows" (closed, not planned) — cited in §3 + as a related-but-unconfirmed failure report +- https://github.com/anomalyco/opencode/issues/8140, #16375 — titles only, confirmed via + `WebSearch`, not individually fetched; §4 +- This repo: `docs/research/opencode-cli-setup.md` — structural template, and source of the + `limit.context`/`limit.output` per-model config shape referenced throughout +- This repo: `litellm-config.yaml` — `qwen3.8-27b-local` model block, `max_tokens` comment on + `reasoning_content`, §3 + +## Confidence summary + +| Claim | Confidence | +|---|---| +| `compaction` is a single global top-level config key (`auto`/`prune`/`reserved`/`tail_turns`/`preserve_recent_tokens`) | High | +| No percentage-threshold config key exists | High | +| Trigger formula is `usedTokens >= context − reservedBuffer`, not a flat percentage | High | +| "Hardcoded 75%" (issue #11314) and "100% threshold" (issue #11930) are unverified community claims, not confirmed mechanism | High (that they're unverified/conflicting); the actual mechanism per source is definitive | +| `overflow.ts`/V1 path ships by default; newer `core` engine is gated behind `OPENCODE_EXPERIMENTAL_NATIVE_LLM` | High (default path); Medium (exact newer-engine behavior if enabled) | +| Reasoning tokens counted via `usage.total_tokens` in the normal (non-fallback) path | High | +| Reasoning tokens excluded from the fallback sum if `total_tokens` is ever absent | High (source); Medium (whether this repo's llama.cpp/litellm stack ever hits that fallback in practice) | +| No per-model/per-agent compaction threshold override exists; confirmed by rejected feature requests | High | +| Compaction mechanism is provider-agnostic — applies identically to `@ai-sdk/openai-compatible` custom providers | High | +| `limit.context` unset/0 on a custom model silently disables compaction for it | High |