diff --git a/docs/research/image-generation-model-choice.md b/docs/research/image-generation-model-choice.md new file mode 100644 index 0000000..30f858a --- /dev/null +++ b/docs/research/image-generation-model-choice.md @@ -0,0 +1,220 @@ +# Research: which diffusion model to target with the full ~32GB R9700 + +**Question:** With Qwen/llama-server fully stopped (per issue #39's premise — see +map #38), ComfyUI has the whole ~32GB R9700 (ROCm/gfx1201) to itself instead +of the ~6GB left over during concurrent operation +(per [`image-generation-options.md`](image-generation-options.md)). Given that +headroom, should the build stay on FLUX.1-schnell, or move up to +FLUX.1-dev, SD3.5-large, Qwen-Image, HunyuanImage-3.0, or Krea-2? + +**Answer, short version:** move up to **Qwen-Image at FP8 precision** +(`qwen_image_fp8_e4m3fn.safetensors` diffusion weights + +`qwen_2.5_vl_7b_fp8_scaled.safetensors` text encoder, ~25GB combined). It's +the only one of the five upgrade candidates with **direct, hardware-specific +evidence of running on this exact GPU architecture** (gfx1201/R9700) rather +than a generic "ComfyUI supports ROCm" inference, it carries the cleanest +license of the group (Apache-2.0, no revenue threshold, no non-commercial +clause), and its 20B MMDiT is a real capability step up from schnell's +distilled 12B (notably for text rendering and prompt adherence), while still +fitting with real margin inside 32GB. + +## Why not just re-derive the schnell/dev/SDXL/SD3.5 findings + +`docs/research/image-generation-options.md` already covers, with primary +sources: ComfyUI's ROCm/gfx1201 story (official AMD docs + RDNA4 blog post + +community gfx1201 Docker images), FLUX.1-schnell vs FLUX.1-dev vs SDXL vs +SD3.5 licenses, and FLUX GGUF VRAM figures at the ~6GB-headroom scale. None +of that is repeated here except where the ~32GB ceiling changes the +conclusion. This doc adds: FLUX.1-dev/SD3.5 at the *larger* headroom, plus +three models the prior doc didn't cover at all (Qwen-Image, HunyuanImage-3.0, +Krea-2). + +## Candidate comparison + +| Model | License (primary source) | Params | Stated/typical VRAM | ROCm/gfx1201 evidence | +|---|---|---|---|---| +| FLUX.1-schnell (current) | Apache-2.0 | 12B | GGUF Q4_K_S ~7GB (per prior doc) | Confirmed on gfx1201 (prior doc) | +| **Qwen-Image** | **Apache-2.0** | 20B (20.4B DiT + 8.3B Qwen2.5-VL text encoder) | fp8 ~16GB (diffusion) + ~9.4GB (fp8 text encoder) ≈ 25GB total; bf16 needs 24GB+ and "48GB+" per some quant write-ups | **Direct**: [kyuz0/amd-r9700-comfy](https://github.com/kyuz0/amd-r9700-comfy) ships a pre-validated "Qwen Image 2512 (FP8) & Lightning LoRA (4 steps)" ComfyUI workflow specifically for the R9700 AI Pro (gfx1201), on a ROCm 7 (TheRock nightlies) toolbox | +| FLUX.1-dev | [FLUX.1-dev Non-Commercial License](https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/main/LICENSE.md) — non-commercial weights, outputs usable commercially | 12B | bf16 ~24GB; GGUF Q8 ~12-13GB | Same *family* evidence as schnell (gfx1201 Docker images target FLUX generally), but no R9700-specific FLUX.1-dev report found | +| SD3.5-large | [Stability Community License](https://huggingface.co/stabilityai/stable-diffusion-3.5-large/blob/main/LICENSE.md) — free under $1M annual revenue | 8B | bf16 ~16GB; 4-bit NF4 fits small GPUs | AMD's own ComfyUI-ROCm doc lists an "SD3.5 Simple" template workflow (per prior doc) — vendor-blessed but not R9700-specific | +| HunyuanImage-3.0 | [tencent-hunyuan-community license](https://github.com/Tencent-Hunyuan/HunyuanImage-3.0) | 80B total / 13B active (MoE, 64 experts) | Official repo: "≥ 3 × 80GB" VRAM for the base model, "≥ 8 × 80GB" for -Instruct; ~177GB at fp16 | **None, and actively contraindicated**: setup requires CUDA 12.8 + FlashAttention2/FlashInfer, no AMD/ROCm mention anywhere in the official repo | +| Krea-2 (Turbo) | [Krea 2 Community License](https://www.krea.ai/krea-2-licensing) — free under $1M annual revenue, homelab/personal explicitly covered | 12-13B DiT | No official VRAM figure; community reports (RTX hardware only) cite fp8 ~16GB, GGUF ~12GB | **None found** — released [June 22, 2026 per the HF model card](https://huggingface.co/krea/Krea-2-Turbo); ComfyUI added native support per [blog.comfy.org](https://blog.comfy.org/p/krea-2-open-source-models-are-now), and GGUF quants exist ([molbal/krea2-gguf](https://huggingface.co/molbal/krea2-gguf)), but no R9700/gfx1201-specific report exists yet — too new for that evidence to have accumulated | + +## Per-model detail + +### Qwen-Image — recommended + +- **License**: Apache-2.0, stated directly on the model card, no revenue + threshold, no non-commercial clause, no attribution/naming requirement. + The cleanest license of every model considered in this doc or the prior + one. Source: [Qwen/Qwen-Image on Hugging Face](https://huggingface.co/Qwen/Qwen-Image). +- **Architecture**: 20B-parameter MMDiT (Multimodal Diffusion Transformer) + combined with an 8.3B Qwen2.5-VL text encoder — notably larger and more + capable than FLUX.1-schnell's 12B distilled model, particularly for + multilingual text rendering and instruction-following, per the official + [QwenLM/Qwen-Image GitHub repo](https://github.com/QwenLM/Qwen-Image). +- **ComfyUI support**: native, not a wrapper/custom-node integration — + landed August 2025 per [ComfyUI Wiki's native-support announcement](https://comfyui-wiki.com/en/news/2025-08-05-qwen-image), + with an official FP8 checkpoint (`qwen_image_fp8_e4m3fn.safetensors`) and + FP8-scaled text encoder (`qwen_2.5_vl_7b_fp8_scaled.safetensors`, 9.38GB) + published under [Comfy-Org/Qwen-Image_ComfyUI](https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/tree/main/split_files/text_encoders). +- **VRAM at FP8**: the official Comfy-Org FP8 diffusion checkpoint plus FP8 + text encoder land around ~25GB combined — comfortably inside the 32GB + ceiling with headroom for ComfyUI runtime/VAE overhead, versus 24GB+ (some + sources say 48GB+) for the unquantized bf16 model. Sources: community + VRAM write-ups aggregated via [Comfy-Org/ComfyUI issue #10852 ("Qwen-image in 24GB VRAM and 32GB RAM")](https://github.com/Comfy-Org/ComfyUI/issues/10852) + and the [official ComfyUI Qwen-Image example](https://comfyanonymous.github.io/ComfyUI_examples/qwen_image/) — + treat exact GB figures as secondary/community-sourced, consistent with how + the prior doc flagged FLUX's GGUF VRAM table. +- **Direct R9700/gfx1201 evidence (the deciding factor)**: + [kyuz0/amd-r9700-comfy](https://github.com/kyuz0/amd-r9700-comfy) is a + Fedora-toolbox ROCm 7 (TheRock nightlies) environment built specifically + for the "AMD Radeon 9700 AI PRO (32GB)" and ships a pre-validated + "Qwen Image 2512 (FP8) & Lightning LoRA (4 steps)" ComfyUI workflow + (plus a Qwen-Image-Edit 2511 FP8 workflow) as one of only four workflows + in the whole repo. This is the only model in this comparison with + card-architecture-specific (not just "ComfyUI supports ROCm generically") + validation — everything else relies on family-level or vendor-generic + ROCm claims. A related community discussion + ([pollockjj/ComfyUI-MultiGPU #133](https://github.com/pollockjj/ComfyUI-MultiGPU/discussions/133)) + does flag that "Qwen image edit doesn't always work on AMD HIP/ROCm" in + some configurations — worth testing the specific workflow before assuming + zero friction, but this is a known-quantity, actively-discussed rough edge + rather than a documented hard blocker. +- **4-step Lightning LoRA**: the R9700-validated workflow pairs Qwen-Image + with a "Lightning" LoRA for 4-step inference — the same fast-inference + pattern FLUX.1-schnell uses, so switching models doesn't have to mean + giving up the short GPU-resident-time-per-image property the prior doc + called out as valuable for time-sliced use alongside llama-server (even + though this ticket's premise is llama-server being stopped, that pattern + still helps if the two services are ever run in an overlapping window). + +### FLUX.1-dev — solid alternative, not the pick + +- License permits non-commercial use of the weights; BFL's own license page + states generated outputs are separately usable commercially (already + covered in the prior doc; unchanged here). Source: + [FLUX.1-dev LICENSE.md](https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/main/LICENSE.md). +- At bf16 (~24GB) or GGUF Q8 (~12-13GB), it fits the 32GB ceiling with room + to spare — a legitimate move up from schnell's distilled quality. + Source: [city96/FLUX.1-dev-gguf](https://huggingface.co/city96/FLUX.1-dev-gguf) + and community Q8 VRAM reports. +- Loses to Qwen-Image on two counts: license (non-commercial-weights clause + vs. Apache-2.0 — not a hard blocker for this homelab per the prior doc's + own reasoning, but strictly worse) and hardware evidence (FLUX's gfx1201 + validation is at the *family* level — the prior doc's `yurisasc/comfyui-rocm-rdna4` + and `charlie12345/R9700AIProComfyUIPatch` sources are about running FLUX + models on this card generally, not a FLUX.1-dev-specific report the way + kyuz0's repo is Qwen-Image-specific). + +### SD3.5-large — no longer the best use of the freed headroom + +- License unchanged from the prior doc: Stability Community License, free + under $1M annual revenue (irrelevant threshold for this homelab). Source: + [stabilityai/stable-diffusion-3.5-large LICENSE.md](https://huggingface.co/stabilityai/stable-diffusion-3.5-large/blob/main/LICENSE.md). +- At 8B params (bf16 ~16GB), it's the smallest of the upgrade candidates — + which mattered when 6GB was the ceiling, but with 32GB available there's + no VRAM reason to pick the model the prior doc already flagged as "lower + fidelity than FLUX/SD3.5 by current standards" over Qwen-Image or + FLUX.1-dev. AMD's own ComfyUI-ROCm docs do list it as a first-party + example template ("SD3.5 Simple", per the prior doc), so it remains a fine + fallback if Qwen-Image's FP8 path hits the ROCm rough edge noted above. + +### HunyuanImage-3.0 — ruled out + +- Official repo states VRAM requirements of "≥ 3 × 80GB" for the base model + and "≥ 8 × 80GB" for HunyuanImage-3.0-Instruct — i.e. multi-GPU + datacenter-class NVIDIA clusters, not a single 32GB consumer/workstation + card at any precision. Source: + [Tencent-Hunyuan/HunyuanImage-3.0 GitHub repo](https://github.com/Tencent-Hunyuan/HunyuanImage-3.0). +- Setup instructions require CUDA 12.8, PyTorch 2.8.0 built for CUDA, and + optionally FlashAttention2/FlashInfer for MoE routing speed — no AMD or + ROCm path is mentioned anywhere in the official repo. Even the + MoE-efficient "13B active" framing doesn't help here: the tooling itself + assumes an NVIDIA multi-GPU cluster, and the 80B total parameter set still + has to be resident somewhere. +- 32GB of headroom on one AMD card doesn't move this model into reach at + any precision considered here; it's excluded regardless of how much VRAM + frees up on this specific box. + +### Krea-2 (Turbo) — promising, but unverified on this hardware + +- Verified directly against primary sources per the ticket's instruction + (this is a June 2026 release, past most training cutoffs): the + [Hugging Face model card](https://huggingface.co/krea/Krea-2-Turbo) states + a release date of **June 22, 2026**, a 12-billion-parameter single-stream + diffusion transformer, `torch.bfloat16` as the reference precision, and + the **Krea 2 Community License**. +- License, per [krea.ai/krea-2-licensing](https://www.krea.ai/krea-2-licensing): + non-commercial (including explicitly personal/homelab) use is free; + commercial use is permitted royalty-free for entities under $1M + trailing-12-month revenue (same shape as SD3.5's and matching this + homelab's use case); content-filter and AI-disclosure obligations apply if + deployed publicly; derivative model names must start with "Krea". +- ComfyUI added native support for both open-weight checkpoints (Krea 2 Raw + and Krea 2 Turbo) per [blog.comfy.org's announcement](https://blog.comfy.org/p/krea-2-open-source-models-are-now), + and community GGUF quants already exist + ([molbal/krea2-gguf](https://huggingface.co/molbal/krea2-gguf)), with + reports (RTX hardware only) of fp8 fitting 16GB and GGUF fitting 12GB. +- **No AMD/ROCm mention anywhere** in the model card, and no gfx1201/R9700 + community report was found — unsurprising given the model is roughly + 2.5 months old at the time of this research. Its architecture (a + standard-shaped DiT that ComfyUI loads through its normal diffusion-model + nodes, per the ComfyUI blog post) gives reasonable expectation it will run + on the same ROCm/PyTorch backend already proven for FLUX and Qwen-Image on + this card, but that's an inference, not a verified fact the way + kyuz0's Qwen-Image workflow is. +- **Not the pick today**, precisely because Qwen-Image already offers a + hardware-verified path at a comparable parameter count and VRAM budget. + Worth a follow-up research ticket once R9700/gfx1201-specific Krea-2 + reports exist — the license and ComfyUI support are both already in + place, so the only open question is real-world ROCm behavior. + +## Recommendation + +**Qwen-Image, FP8 precision** (`qwen_image_fp8_e4m3fn.safetensors` + +`qwen_2.5_vl_7b_fp8_scaled.safetensors`, ~25GB combined), optionally paired +with the 4-step Lightning LoRA the R9700-specific validated workflow uses. +It wins on all three axes the ticket asked about: + +1. **License**: Apache-2.0 — no restriction at all, strictly better than + every other candidate including the current FLUX.1-schnell pick. +2. **ROCm/gfx1201 compatibility**: the only candidate with a workflow + pre-validated specifically on this GPU architecture + ([kyuz0/amd-r9700-comfy](https://github.com/kyuz0/amd-r9700-comfy)), + not just "ComfyUI supports ROCm in general." +3. **VRAM at the ~32GB ceiling**: ~25GB at FP8 leaves real margin for + ComfyUI runtime/VAE overhead, without needing the multi-step, + non-distilled FLUX.1-dev's full 24GB bf16 footprint or accepting + SD3.5's lower fidelity ceiling — and it's a genuine capability upgrade + over schnell (20B vs. 12B, non-distilled-quality text rendering) rather + than just a bigger file. + +If the known Qwen-Image/ROCm edit-mode rough edge +([pollockjj/ComfyUI-MultiGPU #133](https://github.com/pollockjj/ComfyUI-MultiGPU/discussions/133)) +turns out to affect plain text-to-image generation too, SD3.5-large (AMD's +own first-party "SD3.5 Simple" ComfyUI-ROCm template workflow) is the +fallback, with FLUX.1-dev as a second option. HunyuanImage-3.0 is excluded +outright regardless of available VRAM (CUDA-only tooling, multi-GPU +datacenter VRAM floor). Krea-2 is worth revisiting once R9700-specific +field reports exist. + +## Sources consulted + +- [docs/research/image-generation-options.md](image-generation-options.md) (this repo — prior findings, not re-derived) +- [Qwen/Qwen-Image (Hugging Face)](https://huggingface.co/Qwen/Qwen-Image) +- [QwenLM/Qwen-Image (GitHub)](https://github.com/QwenLM/Qwen-Image) +- [ComfyUI Wiki — Qwen-Image native support announcement](https://comfyui-wiki.com/en/news/2025-08-05-qwen-image) +- [Comfy-Org/Qwen-Image_ComfyUI (Hugging Face, FP8 checkpoints)](https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/tree/main/split_files/text_encoders) +- [Comfy-Org/ComfyUI issue #10852 — Qwen-Image VRAM](https://github.com/Comfy-Org/ComfyUI/issues/10852) +- [ComfyUI official Qwen-Image example](https://comfyanonymous.github.io/ComfyUI_examples/qwen_image/) +- [kyuz0/amd-r9700-comfy (R9700-specific ROCm ComfyUI toolbox)](https://github.com/kyuz0/amd-r9700-comfy) +- [pollockjj/ComfyUI-MultiGPU discussion #133 (Qwen-Image-Edit ROCm rough edge)](https://github.com/pollockjj/ComfyUI-MultiGPU/discussions/133) +- [black-forest-labs/FLUX.1-dev (Hugging Face) + LICENSE.md](https://huggingface.co/black-forest-labs/FLUX.1-dev) +- [city96/FLUX.1-dev-gguf](https://huggingface.co/city96/FLUX.1-dev-gguf) +- [stabilityai/stable-diffusion-3.5-large (Hugging Face) + LICENSE.md](https://huggingface.co/stabilityai/stable-diffusion-3.5-large) +- [Tencent-Hunyuan/HunyuanImage-3.0 (GitHub)](https://github.com/Tencent-Hunyuan/HunyuanImage-3.0) +- [krea/Krea-2-Turbo (Hugging Face)](https://huggingface.co/krea/Krea-2-Turbo) +- [Krea 2 Community License Agreement (krea.ai)](https://www.krea.ai/krea-2-licensing) +- [blog.comfy.org — Krea 2 open-source models in ComfyUI](https://blog.comfy.org/p/krea-2-open-source-models-are-now) +- [molbal/krea2-gguf (Hugging Face)](https://huggingface.co/molbal/krea2-gguf)