Files
LLM-Server/docs/research/opencode-auto-compact.md
haylanandClaude-Bot 1eaa2a0d2e docs(research): document opencode auto-compact trigger and config surface
Resolves #27 (part of #26). Findings from opencode.ai/docs (fetched
2026-09-03) plus a fresh clone of anomalyco/opencode @ b578b72
(v1.18.27):

- No percentage-threshold config key exists; compaction.{auto,prune,
  reserved,tail_turns,preserve_recent_tokens} is the full global
  config surface (opencode.json top-level, not per-provider/model).
- Trigger is usedTokens >= context - reservedBuffer, not a hardcoded
  75%/95% cutoff — contradicts an unverified claim in a closed
  GitHub feature request (#11314).
- Reasoning tokens (Qwen3's reasoning_content) are counted via the
  provider's usage.total_tokens in the normal path, but excluded
  from the fallback sum if a provider ever omits total_tokens.
- No per-model/per-agent threshold override exists (confirmed by
  several closed-not-planned feature requests); the only per-model
  lever is each model's own limit.context/limit.output.
- Mechanism is provider-agnostic: applies identically to a hand-
  declared @ai-sdk/openai-compatible provider (this repo's llamacpp
  setup) as to hosted providers, provided limit.context is set.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FCAUsjGNSoJTtK8hyLKg5m
2026-09-03 06:25:57 +02:00

359 lines
23 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# OpenCode CLI: how auto-compact actually decides to trigger
**Date:** 2026-09-03
**Scope:** Resolves haylan/LLM-Server issue #27 (part of #26) — the config key(s), trigger
threshold/formula, reasoning-token accounting, and per-model tunability of OpenCode's
context-compaction behavior, and whether any of it is specific to hosted providers vs. the
`@ai-sdk/openai-compatible` path this repo's `docs/research/opencode-cli-setup.md` documents for
llama.cpp/litellm-style backends.
**Freshness note:** `opencode-cli-setup.md` (dated 2026-08-24) does not cover compaction at all
beyond one sentence ("§5: OpenCode's `limit.context`/`limit.output` fields ... don't change what the
server actually accepts ... may miscalculate when to compact/summarize"). This document supersedes
that gap. It is based on:
1. Live docs fetches from https://opencode.ai/docs/ (2026-09-03), and
2. A fresh `git clone` of https://github.com/anomalyco/opencode at commit
`b578b7261fc9ec4917fe272df5cc4bd8a056cd5d` (2026-09-03T09:47:21+08:00 — same day as this
research), `package.json` version `1.18.27`. Source-code claims below are cited by file path in
that clone and are the **highest-confidence source in this doc** — they're what actually ships,
not a doc description or a third-party claim.
## Confidence scheme
- **High** — read directly from the current source code in the repo, or a docs page quoted
verbatim.
- **Medium** — inferred from source code behavior that isn't spelled out in a single line/comment
(i.e., I traced call sites to confirm it, rather than reading one authoritative line).
- **Low** — plausible but not directly confirmed in the sources checked; flagged as an open
question.
---
## Verdict summary (answers to the four questions asked in #27)
| Question | Answer | Confidence |
|---|---|---|
| Config key(s) controlling threshold/behavior | Single global top-level `compaction` object in `opencode.json`: `auto`, `prune`, `reserved`, `tail_turns`, `preserve_recent_tokens`. **No `threshold` percentage key exists.** | High |
| Is it a hardcoded percentage of `limit.context`? | **No.** It's `usedTokens >= (contextLimit reservedBuffer)`, where `reservedBuffer` defaults to `min(20_000, min(model.limit.output, 32_768) or default)`, i.e. compaction reserves room for one more max-size reply, not a flat 75%/95% cutoff. Some closed GitHub feature requests describe it as "hardcoded 75%" — that claim is **not what the current source does** (see §2). | High (source), contradicts a stale community claim (see §2) |
| Does compaction count reasoning/`reasoning_content` tokens? | **Yes, in practice**, via the provider's `usage.total_tokens` (which for llama.cpp/litellm includes every generated token, reasoning or not) — but the token bookkeeping OpenCode itself derives (`tokens.output`, `tokens.reasoning`) explicitly **splits reasoning out of `output`**, and the compaction trigger's own fallback arithmetic (used only if the provider omits `total_tokens`) **omits `tokens.reasoning` entirely**. See §3 for the exact mechanism and the one edge case where reasoning tokens could be undercounted. | High (source), Medium (edge-case behavior when `total_tokens` is absent) |
| Per-model or single global behavior? | **Global only.** The `compaction` block is a top-level config key, not nested under `provider.<id>.models.<id>` or any per-model schema. It cannot be disabled or tuned for one model while enabled for another. The *indirect* lever is each model's own `limit.context`/`limit.output` (which you already set per-model for custom providers), since those numbers feed the same global formula per-model. Multiple GitHub feature requests (#11314, #11930, #8140, #16375) ask for per-model/per-agent configurability; all are open or closed-not-planned as of this check. | High |
---
## 1. The config surface (verbatim from source + docs)
Global (or project) `opencode.json`:
```json
{
"compaction": {
"auto": true,
"prune": false,
"reserved": 20000,
"tail_turns": null,
"preserve_recent_tokens": null
}
}
```
Field descriptions, quoted verbatim from the config schema
(`packages/core/src/v1/config/config.ts`, lines ~149166 in the cloned repo):
- `auto`*"Enable automatic compaction when context is full (default: true)"*
- `prune`*"Enable pruning of old tool outputs (default: false)"*
- `tail_turns` — *"Maximum number of recent user turns, including their following
assistant/tool responses, to keep verbatim during compaction. By default retention is limited
only by the preserved token budget."*
- `preserve_recent_tokens` — *"Maximum number of tokens from recent turns to preserve verbatim
after compaction"*
- `reserved` — *"Token buffer for compaction. Leaves enough window to avoid overflow during
compaction."*
Sources:
- https://opencode.ai/docs/config/ (fetched 2026-09-03) confirms `auto`/`prune`/`reserved` with the
same defaults and descriptions; the docs page does **not** mention `tail_turns` or
`preserve_recent_tokens` — those two are documented only in the source schema, not (yet) on the
public docs page. **Confidence: High** on all five keys existing and their defaults; the
docs-vs-source gap on the last two is itself notable (docs page is behind the schema).
- `packages/core/src/v1/config/config.ts` (cloned repo, ~line 149) — schema + descriptions.
There is no `threshold`, `percent`, or similarly-named key anywhere in the config schema. I grepped
`packages/core/src/v1/config/` and `packages/core/src/config.ts` for `compaction` and found only the
struct above — no percentage field exists in the current schema (High confidence; direct grep of
current source).
**Not nested under provider/model.** The `compaction` key sits at the top level of `opencode.json`,
a sibling of `provider`, `model`, `agent`, etc. — not inside
`provider.<id>.models.<model-id>` (the block this repo's `opencode-cli-setup.md` §3 documents for
declaring the llama.cpp provider). Confirmed by reading the full top-level config struct in
`packages/core/src/v1/config/config.ts` (~lines 95170): `compaction` is a direct sibling of
`provider`, not a child of it.
There *is* a separate, easily-confused concept: `agent.compaction` (also in that same file, ~line
105) — this only lets you assign a **different agent/model to perform the summarization step
itself** (e.g. run the compaction LLM call on a cheaper model), not a per-model *threshold*
override. It does not change when compaction triggers.
## 2. The actual trigger formula (not a flat percentage)
Live code, `packages/opencode/src/session/overflow.ts` (the shipped/default compaction-trigger
path — see "which code path ships" note at the end of this section):
```ts
const COMPACTION_BUFFER = 20_000
export function usable(input: { cfg: ConfigV1.Info; model: Provider.Model; outputTokenMax?: number }) {
const context = input.model.limit.context
if (context === 0) return 0
const reserved =
input.cfg.compaction?.reserved ??
Math.min(COMPACTION_BUFFER, ProviderTransform.maxOutputTokens(input.model, input.outputTokenMax))
return input.model.limit.input
? Math.max(0, input.model.limit.input - reserved)
: Math.max(0, context - ProviderTransform.maxOutputTokens(input.model, input.outputTokenMax))
}
export function isOverflow(input: {
cfg: ConfigV1.Info
tokens: SessionV1.Assistant["tokens"]
model: Provider.Model
outputTokenMax?: number
}) {
if (input.cfg.compaction?.auto === false) return false
if (input.model.limit.context === 0) return false
const count =
input.tokens.total || input.tokens.input + input.tokens.output + input.tokens.cache.read + input.tokens.cache.write
return count >= usable(input)
}
```
In plain terms:
- Compaction triggers when the running token count (`count`) reaches or exceeds
`usable = contextLimit reservedBuffer` (or, if the model declares a separate
`limit.input`, `usable = limit.input reservedBuffer` instead).
- `reservedBuffer` defaults to `min(20_000, maxOutputTokens)`, where `maxOutputTokens = min(model.limit.output, 32_768) || 32_768`
(`OUTPUT_TOKEN_MAX` constant, `packages/opencode/src/provider/transform.ts``maxOutputTokens()`
at ~line 1468). It can be overridden with `compaction.reserved`.
- If `model.limit.context` is `0` or unset, compaction is **silently disabled entirely** — this
matters for custom `@ai-sdk/openai-compatible` providers where `limit.context` is a value the
user must supply by hand (per `opencode-cli-setup.md` §3); omit it, and auto-compact never fires
for that model.
- Compaction is disabled outright if `compaction.auto === false`.
This is **not** a fixed 75%/90%/95%-of-context cutoff. It's `context reservedOutputBuffer`, which
in practice usually lands somewhere in the 8095%+ range depending on the model's own
`limit.output` relative to `limit.context` — but it's a token-count subtraction, not a percentage
multiplication, and there is no config key to set a percentage.
**On the "hardcoded 75%" claim**: a closed GitHub feature request
(https://github.com/anomalyco/opencode/issues/11314, "Feature Request: Configurable Context
Compaction Threshold") asserts *"Currently, OpenCode triggers context compaction at a hardcoded 75%
threshold of a model's context window."* That is a **user's claim in a feature-request issue, not a
maintainer statement or a source citation**, and it does not match the formula actually in the
current source (which is a token-count subtraction tied to output-buffer size, not a flat
percentage). Treat that issue as evidence that *users perceive/want configurability*, not as an
accurate description of the mechanism. A second, similarly-shaped request
(https://github.com/anomalyco/opencode/issues/11930) instead describes the current default as "100%
threshold" — the two community reports disagree with each other, which is itself a signal neither
is a reliable description of internals; the source code is what should be trusted here.
**Confidence: High** on the source-code formula; **High** that the 75%/100% claims in those two
issues are user speculation, not verified facts — both are closed as "not planned" with no
maintainer confirmation of the underlying mechanism in the content fetched.
**Which code path ships**: the repo contains two parallel implementations — the one quoted above
(`packages/opencode/src/session/overflow.ts` + `packages/opencode/src/session/compaction.ts`,
using `ConfigV1`/`SessionV1` types) and a newer, structurally different one
(`packages/core/src/session/compaction.ts` + `packages/core/src/session/runner/llm.ts`, using
`Config.Entry`/`SessionMessage` types, with its own `compactIfNeeded`/`compactAfterOverflow`
functions and a `buffer` config field instead of `reserved`). Tracing the call path: the newer
runner is only reached when `flags.experimentalNativeLlm` is true, gated by the
`OPENCODE_EXPERIMENTAL_NATIVE_LLM` env var (`packages/opencode/src/effect/runtime-flags.ts`,
`packages/opencode/src/session/llm.ts` line ~226) — off by default. **The `overflow.ts`/V1 path
quoted above is what ships by default in v1.18.27.** If this repo's OpenCode setup ever sets
`OPENCODE_EXPERIMENTAL_NATIVE_LLM=1` (or a future release flips the default), the newer engine's
`compactIfNeeded()` uses a structurally similar but not identical check:
`estimatedPromptTokens <= context max(output, config.buffer)` — same shape (context minus a
reserved buffer, no percentage), different field name (`buffer` vs `reserved`). **Confidence:
High** on which path ships by default; **Medium** on exact behavior differences if the experimental
flag is ever turned on, since I did not exhaustively diff every line of the newer engine.
## 3. Reasoning/thinking-token accounting — the Qwen3.8 angle
This repo's `litellm-config.yaml` flags that Qwen3.8-27B (`qwen3.8-27b-local`) spends generation
tokens on `reasoning_content` before `content`:
```yaml
# Qwen3 is a reasoning model — it spends output tokens on
# reasoning_content before ever writing content. ...
max_tokens: 16384
```
(`/home/haylan/Projects/LLM-Server/litellm-config.yaml`, `qwen3.8-27b-local` model block.)
OpenCode's own token bookkeeping (`packages/opencode/src/session/session.ts`, ~lines 340375, the
function that turns a provider's raw `usage` object into the `tokens` struct stored on each
assistant message) does this:
```ts
const inputTokens = safe(input.usage.inputTokens ?? 0)
const outputTokens = safe(input.usage.outputTokens ?? 0)
const reasoningTokens = safe(input.usage.reasoningTokens ?? 0)
...
const total = input.usage.totalTokens
const tokens = {
total,
input: adjustedInputTokens,
output: safe(outputTokens - reasoningTokens), // reasoning is subtracted OUT of "output"
reasoning: reasoningTokens, // tracked as its own field
cache: { write: cacheWriteInputTokens, read: cacheReadInputTokens },
}
```
And the compaction trigger (`overflow.ts`, quoted in §2) computes:
```ts
const count = input.tokens.total || input.tokens.input + input.tokens.output + input.tokens.cache.read + input.tokens.cache.write
```
Two things follow:
1. **`tokens.reasoning` is never added back in** by the fallback sum
(`input + output + cache.read + cache.write`) — that expression has no `+ reasoning` term. If
the compaction trigger ever fell back to this sum (i.e., the provider didn't return
`totalTokens`), reasoning tokens spent on `reasoning_content` would be **excluded** from the
overflow calculation, undercounting real context usage.
2. **In the normal case, `total` is used instead of the fallback sum**, and `total = usage.totalTokens`
straight from the provider's raw response — computed by the provider/AI-SDK *before* OpenCode
splits `outputTokens` into `output`/`reasoning`. Since the AI SDK's OpenAI-compatible adapter (and
litellm/llama.cpp underneath it) counts every generated token — reasoning and content alike — as
part of `completion_tokens`/`total_tokens`, `total` **does include reasoning tokens** in this
normal path. So in practice, for a llama.cpp/litellm backend that reports `usage.total_tokens` on
every response (the OpenAI chat-completions spec requires this field), **reasoning tokens are
accounted for in the compaction trigger via `total`, not via the explicit `reasoning` field.**
The edge case that would matter for this repo: if litellm or llama.cpp's OpenAI-compatible endpoint
ever omitted `usage.total_tokens` from a response (malformed/incomplete usage block — this has
happened with some llama.cpp server versions/flags), OpenCode's fallback sum would silently
undercount by the full `reasoning` amount, delaying compaction past the point it should have
triggered and increasing risk of a hard `context_length_exceeded` — exactly the failure mode
reported by an unrelated user in
https://github.com/anomalyco/opencode/issues/8089 ("Auto-compaction enabled by default, but
context_length_exceeded errors still occur in agent workflows"), though that issue's cause was not
confirmed to be this specific gap (it involved OpenAI's GPT-5.2 and multi-agent/subagent workflows,
not a local llama.cpp backend, and the issue thread contains no maintainer diagnosis of root cause
in the content fetched).
**Confidence: High** on the source-code mechanics described (the `output = outputTokens
reasoningTokens` split, the `total || sum` fallback, and the missing `+ reasoning` term in the
fallback). **Medium** on whether this repo's specific llama.cpp/litellm stack reliably returns
`usage.total_tokens` on every response for the `qwen3.8-27b-local` model — this was not verified
against a live request/response in this research pass (would need an empirical check: hit
`http://localhost:8080/v1/chat/completions` directly or via the litellm proxy and inspect the
`usage` block of an actual reasoning response). Recommend that empirical check as a fast follow if
this matters operationally.
## 4. Per-model tunability
Confirmed absent, both from the schema (§1) and from community feature requests asking for exactly
this and not getting it:
- https://github.com/anomalyco/opencode/issues/11314 — "Feature Request: Configurable Context
Compaction Threshold" — requests a `compaction.threshold` with "optional per-model overrides."
Closed as not planned (per WebFetch of the issue).
- https://github.com/anomalyco/opencode/issues/11930 — "Feature: Configurable compaction threshold
and model (global + per-model)" — explicitly requests global **and** per-model threshold config.
Closed as not planned, no maintainer reply visible in the content fetched.
- https://github.com/anomalyco/opencode/issues/8140 — "Feature Request: Configurable context limit
and auto-compaction threshold" — same theme (title only confirmed via search; not individually
fetched in this pass).
- https://github.com/anomalyco/opencode/issues/16375 — "[FEATURE]: Per-agent compaction config
(disable compaction for specific agents)" — same theme, per-agent instead of per-model (title only
confirmed via search; not individually fetched in this pass).
All four are open/closed-not-planned as of 2026-09-03 — i.e., **as of this check, none of this has
shipped**: compaction remains a single global on/off + buffer-size knob, with no per-model or
per-agent threshold override. **Confidence: High** that the feature doesn't exist in the schema
(direct source read); **Medium** on the exact current status of #8140/#16375 specifically since only
their titles were confirmed via search results, not their full issue bodies.
The one *indirect* per-model lever that does exist: since `usable()` (§2) reads
`input.model.limit.context` / `input.model.limit.input` / `input.model.limit.output` — all
per-model fields already documented in `opencode-cli-setup.md` §3/§5 for custom providers — setting
those numbers differently per model in the `provider.<id>.models.<model-id>.limit` block changes
where that model's compaction fires, without needing a dedicated per-model compaction key. This
is bookkeeping-hint tuning, not a first-class "compaction threshold" feature.
## 5. Is any of this specific to hosted/built-in providers vs. `@ai-sdk/openai-compatible`?
**No.** The entire trigger path (`overflow.ts`, `compaction.ts`) operates only on `Provider.Model`
(a normalized model descriptor with `limit.context`/`limit.input`/`limit.output`) and the message
`tokens` struct built from the SDK's generic `usage` object (`session.ts`, §3) — nothing in the
compaction code branches on `model.api.npm` or provider identity. The mechanism is provider-agnostic
by construction: any provider adapter that populates `usage` (inputTokens/outputTokens/totalTokens)
and any model entry that has a nonzero `limit.context` gets the same compaction behavior, including
a hand-declared `@ai-sdk/openai-compatible` provider block like this repo's `llamacpp` provider in
`opencode-cli-setup.md` §3. **Confidence: High** — read directly from the trigger/accounting source,
which takes no provider-specific branch.
The one place this repo needs to be careful about, restated from §2: for a custom
`@ai-sdk/openai-compatible` provider, `limit.context` (and ideally `limit.output`) must be set by
hand in `opencode.json` to match the real `--ctx-size` the llama.cpp container is launched with — if
left unset (`limit.context` defaults to `0` for an unrecognized custom model), compaction is
silently disabled for that model rather than silently misfiring.
## Sources consulted (primary)
- https://opencode.ai/docs/ — nav/sitemap fetch (2026-09-03); confirms no dedicated "context
management"/"compaction" page exists in the current docs nav — it lives only in the Config
reference.
- https://opencode.ai/docs/config/ — fetched 2026-09-03; source of the `compaction.auto/prune/reserved`
descriptions and defaults quoted in §1.
- https://opencode.ai/docs/models/ — fetched 2026-09-03; confirms no compaction/context-limit
content on that page (only `reasoningEffort`/`thinking` keys, unrelated to compaction).
- `github.com/anomalyco/opencode` @ `b578b7261fc9ec4917fe272df5cc4bd8a056cd5d` (cloned 2026-09-03,
`package.json` version `1.18.27`) — primary source for all source-code claims:
- `packages/opencode/src/session/overflow.ts` — trigger formula (`usable`/`isOverflow`), §2
- `packages/opencode/src/session/compaction.ts` — shipped compaction service, tail-turn selection,
pruning, §2/§4
- `packages/opencode/src/session/session.ts` (~lines 340375) — `usage``tokens` mapping,
reasoning-token split, §3
- `packages/core/src/v1/config/config.ts` (~lines 95170) — config schema, field descriptions, §1
- `packages/core/src/config/compaction.ts`, `packages/core/src/session/compaction.ts`,
`packages/core/src/session/runner/llm.ts` — the experimental/newer compaction engine, gated
behind `OPENCODE_EXPERIMENTAL_NATIVE_LLM`, §2
- `packages/opencode/src/effect/runtime-flags.ts`, `packages/opencode/src/session/llm.ts` (~line
226) — confirms which engine ships by default, §2
- `packages/opencode/src/provider/transform.ts` (~line 1468, `maxOutputTokens`) — output-buffer
sizing used in the reserved-token default, §2
- https://github.com/anomalyco/opencode/issues/11314 — "Configurable Context Compaction Threshold"
(closed, not planned) — source of the "hardcoded 75%" community claim, §2/§4
- https://github.com/anomalyco/opencode/issues/11930 — "Configurable compaction threshold and model
(global + per-model)" (closed, not planned) — source of the conflicting "100% threshold" claim,
§2/§4
- https://github.com/anomalyco/opencode/issues/8089 — "Auto-compaction enabled by default, but
context_length_exceeded errors still occur in agent workflows" (closed, not planned) — cited in §3
as a related-but-unconfirmed failure report
- https://github.com/anomalyco/opencode/issues/8140, #16375 — titles only, confirmed via
`WebSearch`, not individually fetched; §4
- This repo: `docs/research/opencode-cli-setup.md` — structural template, and source of the
`limit.context`/`limit.output` per-model config shape referenced throughout
- This repo: `litellm-config.yaml``qwen3.8-27b-local` model block, `max_tokens` comment on
`reasoning_content`, §3
## Confidence summary
| Claim | Confidence |
|---|---|
| `compaction` is a single global top-level config key (`auto`/`prune`/`reserved`/`tail_turns`/`preserve_recent_tokens`) | High |
| No percentage-threshold config key exists | High |
| Trigger formula is `usedTokens >= context reservedBuffer`, not a flat percentage | High |
| "Hardcoded 75%" (issue #11314) and "100% threshold" (issue #11930) are unverified community claims, not confirmed mechanism | High (that they're unverified/conflicting); the actual mechanism per source is definitive |
| `overflow.ts`/V1 path ships by default; newer `core` engine is gated behind `OPENCODE_EXPERIMENTAL_NATIVE_LLM` | High (default path); Medium (exact newer-engine behavior if enabled) |
| Reasoning tokens counted via `usage.total_tokens` in the normal (non-fallback) path | High |
| Reasoning tokens excluded from the fallback sum if `total_tokens` is ever absent | High (source); Medium (whether this repo's llama.cpp/litellm stack ever hits that fallback in practice) |
| No per-model/per-agent compaction threshold override exists; confirmed by rejected feature requests | High |
| Compaction mechanism is provider-agnostic — applies identically to `@ai-sdk/openai-compatible` custom providers | High |
| `limit.context` unset/0 on a custom model silently disables compaction for it | High |