Two real-hardware findings from issue #5, both researched and fixed
together (docs/research/rocm-gpu-pin-and-render-group.md):
- GPU_MAX_HW_QUEUES=1 on llama-server and llama-server-fast. Confirmed
on real hardware: either container alone is fine (3% GPU, low
power), only two concurrent HIP contexts pin the R9700 at 100%/
boost-clock (ROCm/ROCm#5706, an MES firmware bug). The env var is
validated on the exact image this stack uses, per-process by design
— applying it to both containers is the correct scope.
- group_add switched from plain names (video/render) to resolved
numeric GIDs (HOST_VIDEO_GID/HOST_RENDER_GID) on all three GPU
services. The "unable to find group render: no matching entries in
group file" error confirmed new since the second GPU service was
added is a known Docker bug (docker/cli#4714): group_add by name
resolves against the container's own /etc/group, not the host's,
and multiple GPU services starting concurrently race on that lookup.
Numeric GIDs skip resolution entirely. scripts/update.sh's existing
comfyui-only GID resolution is generalized to resolve these once for
all three services.
docker compose config -q validated (fails fast with a clear error if
the GIDs aren't resolved yet, passes once they are).
Refs #5
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MrnMEdzeQzqZE5soVEXPCx