docs(research): add image-gen backend survey and omniroute/qwen websearch notes

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zCwaWJQuKgXDUfRPBqDS7
This commit is contained in:
2026-09-05 21:24:23 +02:00
co-authored by Claude-Bot
parent ac3f730f83
commit 7d1ff2f54f
2 changed files with 579 additions and 0 deletions
+249
View File
@@ -0,0 +1,249 @@
# Research: adding local image generation to the stack
**Question:** What's the best way to add local image generation alongside
the existing Qwen3.8-27B / llama.cpp text stack, given a single AMD Radeon
R9700 (32GB VRAM, ROCm/gfx1201 — not CUDA), routed through the OmniRoute
gateway on the `ai-stack` Docker network?
**Answer, short version:** run **ComfyUI** (official AMD-blessed ROCm
Docker path exists, and OmniRoute already has a first-class `comfyui`
provider — no bespoke API wrapper needed) with **FLUX.1 [schnell]**
(Apache-2.0, 4-step, GGUF-quantizable) as the default model, falling back to
**SDXL** for anything schnell's distilled-step license/quality tradeoffs
don't suit. VRAM headroom against the current llama-server footprint is too
tight for both to be resident at once at any real image quality — plan for
**time-sliced use** (llama-server's existing lazytainer stop-on-idle pattern,
mirrored for the image-gen service, or a manual "stop one, start the other"
toggle), not concurrent operation.
## Current VRAM baseline (from this repo)
Per `docker-compose.yml` and `.env.example`, llama-server runs
`Qwen3.8-27B-UD-Q4_K_XL.gguf` (17.6 GB weights) at `--ctx-size 262144` with
`--cache-type-k q8_0 --cache-type-v q8_0`, landing at **~25.6 GB** total
(weights + q8_0 KV cache), leaving **~6 GB** free on the 32GB card — this
matches the math already recorded in
[`docs/research/qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md). Per the auto-memory
note on this repo, real measured VRAM use has run closer to ~75% (~24 GB) in
practice versus the theoretical estimate, which doesn't change the
conclusion below but means the ~6 GB figure is closer to a ceiling than a
comfortable number.
**Implication:** 6 GB is not enough for any current-generation image model at
usable quality (see VRAM table below — even the smallest practical FLUX
quant wants ~7 GB alone, before ComfyUI's own runtime/VAE overhead). Running
image-gen *concurrently* with llama-server resident is not realistic on this
card. The two need to time-share the GPU, not split it.
## Backend evaluation (ROCm support, checked against primary sources)
### ComfyUI — recommended
- **Official AMD ROCm docs exist and are current.** AMD's own ROCm docs site
hosts a dedicated ComfyUI install guide with a prebuilt Docker image path
(recommended) or build-from-source, listing ROCm 7.2.0 and 7.1.0 as
supported versions, explicit `--device=/dev/kfd --device=/dev/dri
--group-add video` flags (same device-passthrough pattern this repo
already uses for llama-server), and template workflows including "SD3.5
Simple". Officially the guide only names AMD Instinct
MI355X/MI325X/MI300X (datacenter cards) as supported platforms.
Source: [ROCm docs — ComfyUI on ROCm installation](https://rocm.docs.amd.com/projects/comfyui/en/docs-26.04/install/comfyui-install.html).
- **The upstream ComfyUI README itself documents AMD support directly**,
including consumer cards: stable ROCm install via
`pip install torch torchvision torchaudio --index-url
https://download.pytorch.org/whl/rocm7.2`, plus an experimental Windows
build explicitly naming **RDNA 3 (RX 7000), RDNA 3.5 (Strix Halo), and
RDNA 4 (RX 9000 series)** — i.e. the same RDNA4 generation as the R9700 —
and `HSA_OVERRIDE_GFX_VERSION` workarounds for older/unlisted cards.
Source: [comfyanonymous/ComfyUI README](https://github.com/comfyanonymous/ComfyUI).
- **AMD has published a specific RDNA4/RX 9000 ComfyUI guide** (separate
from the Instinct-only install page above), confirming RDNA4 consumer
cards are an explicitly supported, first-party-documented target, not just
a community workaround.
Source: [ROCm blog — Getting Started with ComfyUI on AMD Radeon RX 9000 Series GPUs](https://rocm.blogs.amd.com/artificial-intelligence/comfyui-radeon-9000/README.html).
- **gfx1201 (R9700's arch) specifically has active community Docker images**:
`yurisasc/comfyui-rocm-rdna4` targets ROCm 7.1 + PyTorch 2.9.1 with
`HSA_OVERRIDE_GFX_VERSION=12.0.1` / `PYTORCH_ROCM_ARCH=gfx1201` baked in,
and there's a published community patch specifically for R9700 AI Pro +
ComfyUI video-gen speedups, evidence the card is being run today, not just
theoretically compatible.
Sources: [yurisasc/comfyui-rocm-rdna4](https://github.com/yurisasc/comfyui-rocm-rdna4),
[charlie12345/R9700AIProComfyUIPatch](https://github.com/charlie12345/R9700AIProComfyUIPatch).
- **Known gfx1201 caveat:** AMD's own TransformerEngine repo has an open
issue confirming gfx1201 is missing from the FP8 architecture table, so
FP8 kernels silently fall back to FP32 with ~50% throughput loss
(18-22 vs. 35-40 tok/s in the reporter's LLM benchmark) — not a
correctness blocker, but relevant if planning to use FP8-quantized image
models expecting native FP8 speed on this card; GGUF/Q-quants (see below)
avoid this path entirely since they dequantize to bf16/fp16, not fp8.
Source: [ROCm/TransformerEngine issue #520](https://github.com/ROCm/TransformerEngine/issues/520).
- **Actively maintained community Docker packaging** beyond AMD's own image:
`YanWenKun/ComfyUI-Docker` ships parallel `rocm` (PyTorch-build-based,
faster releases) and `rocm7` (AMD-build-based, more comprehensive)
variants, both targeting ROCm 7, with ~1000 commits of ongoing history —
a viable alternative to the official AMD image if it lags behind ComfyUI
releases.
Source: [YanWenKun/ComfyUI-Docker](https://github.com/YanWenKun/ComfyUI-Docker).
### AUTOMATIC1111 / Forge — usable but a step down for this hardware
- ROCm support for A1111/Forge is real but community-patched, not
first-party. The upstream `lllyasviel/stable-diffusion-webui-forge` repo's
own discussion thread on AMD support points users to
`lshqqytiger/stable-diffusion-webui-amdgpu-forge`, a community fork
specifically maintained for AMD, "regarded as the go-to version" for
running FLUX-era models on AMD — i.e. the *mainline* Forge repo does not
claim ROCm support itself; you're expected to run a fork.
Source: [lllyasviel/stable-diffusion-webui-forge discussion #67](https://github.com/lllyasviel/stable-diffusion-webui-forge/discussions/67).
- No first-party AMD vendor documentation (unlike ComfyUI's AMD-authored
ROCm/RDNA4 blog posts above) was found for A1111/Forge specifically.
Given ComfyUI already has an AMD-blessed path plus a first-class OmniRoute
provider (below), there's no reason to take on a community fork's
maintenance risk instead.
### InvokeAI — usable but weaker AMD story for a new-generation card
- InvokeAI documents ROCm support but flags it as second-tier: "AMD GPUs
are only supported on Linux," and "support for newer AMD GPUs is spotty
... you may experience garbled images, black images, or long startup
delays." Its own install docs reference ROCm 5.4.2-era wheels, notably
older than the ROCm 7.x this stack's llama-server image already runs on
gfx1201.
Source: [InvokeAI installation docs (mauwii mirror)](https://mauwii.github.io/InvokeAI/installation/030_INSTALL_CUDA_AND_ROCM/).
- No OpenAI-compatible-images angle either — same drawback as A1111/Forge.
Not recommended as primary given ComfyUI's stronger, more current AMD
documentation trail.
## Model choice: FLUX.1 [schnell] vs FLUX.1 [dev] vs SDXL vs SD3.5
| Model | License | Params | Notes |
|---|---|---|---|
| **FLUX.1 [schnell]** | **Apache-2.0** — fully open, no commercial restriction | 12B | Distilled for 1-4 step inference (fast); Black Forest Labs' own model card states this license directly |
| FLUX.1 [dev] | [FLUX.1-dev Non-Commercial License](https://github.com/black-forest-labs/flux/blob/main/model_licenses/LICENSE-FLUX1-dev) | 12B | Non-commercial for the *model/weights*; generated *outputs* are explicitly usable commercially per BFL's license page. Higher quality than schnell (more steps, non-distilled) but the weights themselves can't be redistributed/used commercially |
| SDXL | CreativeML OpenRAIL++ (permissive, commercial-friendly) | ~3.5B | Older (2023), lower fidelity than FLUX/SD3.5 by current standards, but lowest VRAM footprint and best long-standing tooling maturity |
| SD3.5 (Large/Medium) | Stability AI Community License — free commercial use under $1M annual revenue, else enterprise license required | 8B / 2.5B | Free for this repo's non-commercial homelab use regardless; template already listed in AMD's own ComfyUI-ROCm doc ("SD3.5 Simple") as a first-party example workflow |
Sources: [black-forest-labs/flux model cards](https://github.com/black-forest-labs/flux/blob/main/model_cards/FLUX.1-dev.md),
[black-forest-labs/FLUX.1-schnell on Hugging Face](https://huggingface.co/black-forest-labs/FLUX.1-schnell)
(license: apache-2.0), [FLUX.1-dev LICENSE.md](https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/main/LICENSE.md),
[Stability AI — Introducing Stable Diffusion 3.5](https://stability.ai/news-updates/introducing-stable-diffusion-3-5),
[stabilityai/stable-diffusion-3.5-large LICENSE.md](https://huggingface.co/stabilityai/stable-diffusion-3.5-large/blob/main/LICENSE.md).
**Recommendation: FLUX.1 [schnell].** For a private homelab, license
enforcement isn't the deciding factor by itself, but schnell's Apache-2.0
status removes any future ambiguity if outputs or the setup are ever shared
or repurposed, and its whole design point — good quality in 1-4 sampling
steps — directly addresses the VRAM/time-slicing constraint below (less
GPU-resident time per image than a 20-50 step dev/SDXL/SD3.5 run).
Quantized via `city96/ComfyUI-GGUF` (an actively-referenced, community-
trusted quantization node — its GGUF Q-quants dequantize to bf16/fp16 at
runtime, sidestepping the gfx1201 FP8 dequant bug above entirely), FLUX fits
in a fraction of its fp16 footprint:
| Precision | Approx. VRAM (model only) |
|---|---|
| fp16 (baseline) | ~24 GB |
| fp8 | ~12 GB |
| GGUF Q5_K_S | ~12-15 GB (practical quality floor) |
| GGUF Q4_K_S | ~7 GB (quality starts degrading on hands/text below Q4) |
Source: aggregated VRAM figures from GGUF-quantization write-ups referencing
city96's FLUX GGUF conversions — treat as secondary/community sourced
(no single BFL-published VRAM table was found), consistent across multiple
independent sources.
[city96/ComfyUI-GGUF README](https://github.com/city96/ComfyUI-GGUF/blob/main/README.md),
[city96/FLUX.1-dev-gguf model card](https://huggingface.co/city96/FLUX.1-dev-gguf).
**Fallback pick: SDXL.** If schnell's distilled quality ceiling proves too
low for some use case, SDXL is the safer second choice over FLUX.1 [dev] or
SD3.5 specifically *because* of this card's tight headroom: it's the
smallest of the four by a wide margin, has the longest production track
record on ROCm of any of these models, and its OpenRAIL++ license carries no
revenue-threshold clause to track (unlike SD3.5's Community License) or
non-commercial weight restriction (unlike FLUX.1 [dev]).
## OpenAI-compatible API / OmniRoute integration
This is the best news in this research: **OmniRoute already ships a
first-class, built-in `comfyui` provider** — not a generic "point it at an
OpenAI base URL and hope" integration. Its provider reference documents it
explicitly: *"No API key required. Configure the local ComfyUI base URL
(default: http://localhost:8188)."* OmniRoute's own image-routing feature
set (`/v1/images/generations`, `/v1/images/edits`, `/v1/images/variations`,
with automatic provider fallback) is designed for exactly this pattern:
register ComfyUI as a backend, then any client already calling OmniRoute's
OpenAI-compatible images endpoints reaches it with no extra shim.
Source: [diegosouzapw/OmniRoute PROVIDER_REFERENCE.md](https://github.com/diegosouzapw/OmniRoute/blob/main/docs/reference/PROVIDER_REFERENCE.md),
[diegosouzapw/OmniRoute repo description](https://github.com/diegosouzapw/OmniRoute).
This means **ComfyUI does not need an extra OpenAI-API wrapper project**
the wrapper projects found during this research
([`ComfyUI-OpenAI-Compatible-API`](https://github.com/yeeyou/ComfyUI-OpenAI-Compatible-API))
turned out to be a ComfyUI *custom node* for calling *outbound* to LLM APIs
from within a workflow (the reverse direction), not something this stack
needs — OmniRoute's own native ComfyUI provider is the actual integration
point, one layer up.
**Confidence note:** the provider-reference detail above was fetched via an
automated summarizer against the raw doc rather than manually re-verified
line-by-line; re-check `PROVIDER_REFERENCE.md`'s `comfyui` entry directly
before wiring this up, in case ComfyUI's own `/prompt` API (a
workflow-graph-shaped API, not a simple text-prompt-in/image-out call) needs
a specific default workflow JSON configured on the OmniRoute side to produce
a plain text-to-image call.
## Integration sketch (not a full compose — see caveats above)
- New service in `docker-compose.yml`, e.g. `comfyui`, image
`rocm/comfyui-rocm` (or `yurisasc/comfyui-rocm-rdna4` for a gfx1201-tuned
build) or built from AMD's own ROCm ComfyUI Dockerfile, same
`/dev/kfd` + `/dev/dri` + `group_add: [video, render]` device-passthrough
block already used for `llama-server`, joined to the same `ai-stack`
network so `omniroute` can reach it as `http://comfyui:8188` — no host
port needed (matches the existing llama-server pattern of no published
port, gateway-only access).
- Register it in OmniRoute's dashboard as a `comfyui` provider pointing at
that internal URL, same manual-registration pattern already used for
llama-server and searxng-search per `docs/proxy-key-onboarding.md`.
- **VRAM contention is the real design problem, not networking.** Given the
~6 GB headroom, the two services can't both sit GPU-resident.
Two workable patterns, in order of how well they fit what's already in
this repo:
1. **Mirror the existing lazytainer stop-on-idle pattern** already applied
to `llama-server` (`docker-compose.yml`'s `lazytainer.group.*` labels) —
add an equivalent idle-timeout group for `comfyui`, and rely on the two
services naturally not being hit at the same time for a single-user
homelab. This doesn't *guarantee* mutual exclusion (both could still be
woken concurrently and both try to fit in 6 GB free), so it's a
reasonable-effort fit, not a hard guarantee.
2. **Explicit mutual exclusion**: a small script/compose profile that
stops `llama-server` before starting `comfyui` (and vice versa) rather
than relying on lazytainer's independent idle timers — worth doing if
the reasonable-effort version above causes a visible OOM in practice.
Either way, this is a "pick one, then the other" story, not "run both."
- Given FLUX.1 [schnell]'s 1-4 step design, a cold-start-and-generate cycle
(wake ComfyUI from lazytainer sleep, generate, let it idle back down) is
a reasonably good fit for occasional image requests through the same
gateway that already does this for llama-server.
## Sources consulted
- [ROCm docs — ComfyUI on ROCm installation](https://rocm.docs.amd.com/projects/comfyui/en/docs-26.04/install/comfyui-install.html)
- [ROCm blog — ComfyUI on AMD Radeon RX 9000 Series (RDNA4)](https://rocm.blogs.amd.com/artificial-intelligence/comfyui-radeon-9000/README.html)
- [comfyanonymous/ComfyUI README](https://github.com/comfyanonymous/ComfyUI)
- [YanWenKun/ComfyUI-Docker](https://github.com/YanWenKun/ComfyUI-Docker)
- [yurisasc/comfyui-rocm-rdna4](https://github.com/yurisasc/comfyui-rocm-rdna4)
- [charlie12345/R9700AIProComfyUIPatch](https://github.com/charlie12345/R9700AIProComfyUIPatch)
- [ROCm/TransformerEngine issue #520 (gfx1201 FP8 fallback)](https://github.com/ROCm/TransformerEngine/issues/520)
- [lllyasviel/stable-diffusion-webui-forge discussion #67 (AMD support)](https://github.com/lllyasviel/stable-diffusion-webui-forge/discussions/67)
- [InvokeAI CUDA/ROCm install docs](https://mauwii.github.io/InvokeAI/installation/030_INSTALL_CUDA_AND_ROCM/)
- [black-forest-labs/flux GitHub repo + model cards](https://github.com/black-forest-labs/flux)
- [black-forest-labs/FLUX.1-schnell (Hugging Face, Apache-2.0)](https://huggingface.co/black-forest-labs/FLUX.1-schnell)
- [black-forest-labs/FLUX.1-dev LICENSE.md](https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/main/LICENSE.md)
- [Stability AI — Introducing Stable Diffusion 3.5](https://stability.ai/news-updates/introducing-stable-diffusion-3-5)
- [stabilityai/stable-diffusion-3.5-large LICENSE.md](https://huggingface.co/stabilityai/stable-diffusion-3.5-large/blob/main/LICENSE.md)
- [city96/ComfyUI-GGUF](https://github.com/city96/ComfyUI-GGUF)
- [city96/FLUX.1-dev-gguf](https://huggingface.co/city96/FLUX.1-dev-gguf)
- [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) and its `PROVIDER_REFERENCE.md`
- [yeeyou/ComfyUI-OpenAI-Compatible-API](https://github.com/yeeyou/ComfyUI-OpenAI-Compatible-API) (checked and ruled out — wrong direction)
- This repo: `docker-compose.yml`, `.env.example`, `docs/research/qwen3.8-27b-quant.md`