docs(research): confirm lazytainer/omniroute idle-stop conflict (#40)

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zCwaWJQuKgXDUfRPBqDS7
This commit is contained in:
2026-09-05 20:57:24 +02:00
co-authored by Claude-Bot
parent d8736b6dd7
commit 451d5c7b28
@@ -0,0 +1,194 @@
# Research: Why lazytainer's idle-stop on llama-server doesn't fire, and what switch-model.sh should do about it
**Question:** lazytainer is configured on `llama-server` (`docker-compose.yml`
`lazytainer.group.llamaserver.*` labels) but its idle-stop never triggers in
practice — OmniRoute appears to keep the container looking "active" to
lazytainer's packet-threshold detector. Confirm the mechanism, find root
cause, and recommend how the future `scripts/switch-model.sh` (#43, blocked)
should handle GPU-residency swaps between `llama-server` and a new `comfyui`
service given this.
**Answer:** Confirmed. lazytainer's detector is a dumb per-port packet
counter with no traffic classification — it cannot tell OmniRoute's
background provider health-check pings apart from real inference traffic,
and there is no config knob in lazytainer or a per-provider one in OmniRoute
that fixes this. **`switch-model.sh` should bypass lazytainer entirely** for
the swap: drive `docker compose stop`/`up -d` directly on both services,
rather than trying to make lazytainer's idle-stop cooperate.
## Current config (`docker-compose.yml`)
```yaml
labels:
- "lazytainer.group.llamaserver.sleepMethod=stop"
- "lazytainer.group.llamaserver.ports=8080"
- "lazytainer.group.llamaserver.inactiveTimeout=${LAZYTAINER_INACTIVE_TIMEOUT:-900}"
- "lazytainer.group.llamaserver.minPacketThreshold=2"
```
`ports=8080` matches the container's real internal port (`expose: ["8080"]`,
confirmed in the same file) — not a misconfiguration. `minPacketThreshold=2`
is already far *below* lazytainer's own documented default of `30`, i.e. this
deployment already tried loosening the threshold to make idle-stop easier to
reach, not harder.
## How lazytainer's detector actually works (primary source: `vmorganp/Lazytainer`)
Confirmed against the project's README and Go source
(`src/group.go`) on [github.com/vmorganp/Lazytainer](https://github.com/vmorganp/Lazytainer):
- It captures packets with **gopacket/libpcap directly on the configured
`netInterface`** (default `eth0`), applying a BPF filter built from the
group's `ports` list (`"port 8080"` here, per the source's filter-string
construction, e.g. `"port 80 or port 81 or etc."` in the general case).
- The filter matches **every packet to or from the port** — SYN, ACK,
data, FIN, everything. It is not restricted to new-connection SYNs.
- Every `pollRate` seconds (default `30`; not overridden in this repo's
config) it samples a rolling packet counter (`rxHistory`) and compares the
delta against `minPacketThreshold`:
`rxHistory[0]+minPacketThreshold > rxHistory[len(rxHistory)-1]` → treated as
active, `inactiveSeconds` resets to 0.
- `ignoreActiveClients` (default `false`, not set here) only changes whether
an ESTABLISHED-connection count is also checked; it does not add any
content- or source-based filtering.
- **There is no mechanism anywhere in lazytainer to exclude specific traffic
(by source IP, path, header, or request type) from the packet count.** The
README's config table (`ports`, `inactiveTimeout`, `minPacketThreshold`,
`ignoreActiveClients`, `pollRate`, `sleepMethod`, `netInterface`) is
exhaustive — nothing else exists to tune this per-caller.
Consequence: a single TCP connection to port 8080 — a bare connect + one
small HTTP exchange + close — already produces well over `minPacketThreshold=2`
packets purely from the handshake and teardown (SYN, SYN-ACK, ACK, ..., FIN,
ACK), regardless of payload size or purpose. At this threshold, essentially
*any* connection to the port counts as "active" and resets `inactiveTimeout`.
Raising the threshold wouldn't help either — the fix would need to be
"ignore packets from OmniRoute's health-checker," which the tool has no way
to express; it only counts packets on a port, source-blind.
## How OmniRoute actually touches registered providers (primary source: `diegosouzapw/OmniRoute`)
Confirmed against
[`docs/reference/ENVIRONMENT.md`](https://github.com/diegosouzapw/OmniRoute/blob/main/docs/reference/ENVIRONMENT.md)
in the OmniRoute repo:
- OmniRoute runs a **background credential/connection health-check
scheduler** (`src/lib/credentialHealth/scheduler.ts`) on
`CREDENTIAL_HEALTH_CHECK_INTERVAL`, default `300000` ms (5 min), minimum
`10000` ms (10s) — this periodically re-tests each registered provider's
connection, which for a provider like `llama-server` (a plain HTTP base
URL, no API key) means an actual request/connection to
`llama-server:8080`.
- Results are cached for `CREDENTIAL_HEALTH_CACHE_TTL` (default also 5 min).
- **Only one exclusion exists, and it's hardcoded by provider category, not
configurable per-provider**: search providers
(`SEARCH_VALIDATOR_CONFIGS` in
`src/lib/providers/validation/searchProviders.ts`, e.g. `tavily-search`)
are permanently skipped because their validation call is a real billed
upstream query. `llama-server` is an inference provider, not a search
provider — it is not in this exclusion list.
- The only toggle that actually stops the sweep is global:
`OMNIROUTE_DISABLE_CREDENTIAL_HEALTH_CHECK=1`/`true`, which "disable[s]
background periodic testing of provider connections" for **every**
registered provider at once. There is no documented per-provider
disable/pause flag in
[`docs/reference/PROVIDER_REFERENCE.md`](https://github.com/diegosouzapw/OmniRoute/blob/main/docs/reference/PROVIDER_REFERENCE.md) —
the dashboard's `/dashboard/providers` page is described only as where you
"enable, configure, and test each provider," with no documented
independent "pause health checks for this one provider" control.
So: OmniRoute is not the sole cause, but it is a live, recurring cause. Every
5 minutes (at most — could also be triggered ad hoc by dashboard/API use) it
opens a connection to `llama-server:8080` purely to check the provider is
alive, which is exactly the kind of traffic lazytainer's port-level counter
cannot distinguish from real inference calls. With `inactiveTimeout=900`
(15 min) and a health-check every ≤300s, the container practically always
sees qualifying traffic before its idle timer would expire.
## Root cause
Two independent, both-true facts combine to defeat idle-stop:
1. **lazytainer's detector is fundamentally traffic-blind** — it counts raw
packets on a port with no way to exclude any specific caller or traffic
class. This is a property of the tool, not a misconfiguration in this
repo (`ports=8080` is correct; `minPacketThreshold=2` is already at the
permissive end).
2. **OmniRoute periodically pings every registered non-search provider**
(default every ≤5 min) to keep its health/availability status current,
and that ping is indistinguishable, at the packet level, from a real
inference request.
Neither side offers a targeted fix: lazytainer has no allowlist/denylist by
source, and OmniRoute's only "stop pinging" lever
(`OMNIROUTE_DISABLE_CREDENTIAL_HEALTH_CHECK`) is all-or-nothing across every
provider, not scoped to just `llama-server`. Tuning `minPacketThreshold`
higher or lower doesn't change the outcome either way, since the health-check
traffic and real traffic land on the exact same port with no distinguishing
packet-level signature.
## Recommendation for `scripts/switch-model.sh` (#43)
**Bypass lazytainer entirely for the GPU-residency swap.** Drive both
services directly:
```bash
docker compose stop llama-server
docker compose up -d comfyui
# ...and the reverse when swapping back
```
Justification:
- The swap is a **deliberate, scripted, known-in-advance** event — the
script always knows exactly which service should go up and which should
go down. Idle-stop detection exists to handle the case where nobody knows
when a service last had traffic; that's not this case, so routing the
swap through a passive heuristic (lazytainer's idle timer) that this
research shows is already unreliable for `llama-server` adds a point of
failure for no benefit. Direct `docker compose stop`/`up -d` is
deterministic and immune to the packet-counting confound described above.
- Reconfiguring lazytainer's thresholds was considered and rejected: no
threshold value fixes a detector that cannot distinguish OmniRoute's
keepalive traffic from real traffic on the same port (see Root cause).
This is a ceiling in the tool itself, not a tuning problem.
- Pausing OmniRoute's polling for the swap window was also considered.
It's the one lever available (`OMNIROUTE_DISABLE_CREDENTIAL_HEALTH_CHECK`),
but it is global — it would blind OmniRoute's health status for *every*
provider (including `searxng-search`, if registered) for the duration of
the swap, and adds an extra env-toggle-and-restart step to the script for
a problem that direct compose control sidesteps completely. It's worth
flagging for #43's implementation as a *secondary* safety measure — briefly
disabling the sweep (or accepting that OmniRoute may show `llama-server` as
errored/offline for up to `CREDENTIAL_HEALTH_CHECK_INTERVAL` after it's
stopped) — but it should not be the primary mechanism the swap relies on.
- This does **not** require removing the existing `lazytainer.group.llamaserver.*`
labels — they can stay for whatever idle-stop benefit they still provide
between swaps (e.g. genuinely idle periods where nothing, including
OmniRoute, has recently touched the container long enough to matter) while
`switch-model.sh` simply never depends on lazytainer to do the actual
stop/start for a swap.
## Bottom line for #43 (blocked ticket, once unblocked)
- `switch-model.sh` should call `docker compose stop <from-service>` /
`docker compose up -d <to-service>` directly — never rely on lazytainer's
idle-stop to free the GPU as part of a swap.
- No lazytainer config change (threshold, ports, poll rate) is a viable fix;
the detector has no way to exclude OmniRoute's traffic by source.
- Optionally, as a secondary hygiene step, the script may toggle
`OMNIROUTE_DISABLE_CREDENTIAL_HEALTH_CHECK` around the swap (or simply
tolerate a stale "errored" status in OmniRoute's dashboard for up to one
`CREDENTIAL_HEALTH_CHECK_INTERVAL`) to avoid OmniRoute flagging the
just-stopped provider as failed mid-swap — but this is cosmetic/status
hygiene, not what makes the swap itself work.
Sources: [`vmorganp/Lazytainer`](https://github.com/vmorganp/Lazytainer)
(README config table; `src/group.go` packet-capture and threshold-comparison
logic), [`diegosouzapw/OmniRoute` —
`docs/reference/ENVIRONMENT.md`](https://github.com/diegosouzapw/OmniRoute/blob/main/docs/reference/ENVIRONMENT.md)
(credential health-check scheduler env vars), [`diegosouzapw/OmniRoute` —
`docs/reference/PROVIDER_REFERENCE.md`](https://github.com/diegosouzapw/OmniRoute/blob/main/docs/reference/PROVIDER_REFERENCE.md)
(provider dashboard controls), this repo's `docker-compose.yml`
(`lazytainer.group.llamaserver.*` labels, `llama-server`/`omniroute` service
definitions).