Files
LLM-Server/docs/research/image-generation-model-choice.md
haylanandClaude-Bot ac3f730f83 docs(research): recommend Qwen-Image FP8 for full-VRAM diffusion build
Closes issue #39 — with llama-server stopped and the full ~32GB R9700
available, Qwen-Image (Apache-2.0, 20B MMDiT) at FP8 precision (~25GB)
is the recommended upgrade from FLUX.1-schnell: it has the cleanest
license of the candidates and is the only one with a ComfyUI workflow
pre-validated specifically on this GPU architecture (gfx1201/R9700),
per kyuz0/amd-r9700-comfy. HunyuanImage-3.0 is ruled out (CUDA-only,
multi-GPU datacenter VRAM floor); Krea-2 flagged as promising but too
new for R9700-specific field evidence.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zCwaWJQuKgXDUfRPBqDS7
2026-09-05 21:03:13 +02:00

16 KiB
Raw Permalink Blame History

Research: which diffusion model to target with the full ~32GB R9700

Question: With Qwen/llama-server fully stopped (per issue #39's premise — see map #38), ComfyUI has the whole ~32GB R9700 (ROCm/gfx1201) to itself instead of the ~6GB left over during concurrent operation (per image-generation-options.md). Given that headroom, should the build stay on FLUX.1-schnell, or move up to FLUX.1-dev, SD3.5-large, Qwen-Image, HunyuanImage-3.0, or Krea-2?

Answer, short version: move up to Qwen-Image at FP8 precision (qwen_image_fp8_e4m3fn.safetensors diffusion weights + qwen_2.5_vl_7b_fp8_scaled.safetensors text encoder, ~25GB combined). It's the only one of the five upgrade candidates with direct, hardware-specific evidence of running on this exact GPU architecture (gfx1201/R9700) rather than a generic "ComfyUI supports ROCm" inference, it carries the cleanest license of the group (Apache-2.0, no revenue threshold, no non-commercial clause), and its 20B MMDiT is a real capability step up from schnell's distilled 12B (notably for text rendering and prompt adherence), while still fitting with real margin inside 32GB.

Why not just re-derive the schnell/dev/SDXL/SD3.5 findings

docs/research/image-generation-options.md already covers, with primary sources: ComfyUI's ROCm/gfx1201 story (official AMD docs + RDNA4 blog post + community gfx1201 Docker images), FLUX.1-schnell vs FLUX.1-dev vs SDXL vs SD3.5 licenses, and FLUX GGUF VRAM figures at the ~6GB-headroom scale. None of that is repeated here except where the ~32GB ceiling changes the conclusion. This doc adds: FLUX.1-dev/SD3.5 at the larger headroom, plus three models the prior doc didn't cover at all (Qwen-Image, HunyuanImage-3.0, Krea-2).

Candidate comparison

Model License (primary source) Params Stated/typical VRAM ROCm/gfx1201 evidence
FLUX.1-schnell (current) Apache-2.0 12B GGUF Q4_K_S ~7GB (per prior doc) Confirmed on gfx1201 (prior doc)
Qwen-Image Apache-2.0 20B (20.4B DiT + 8.3B Qwen2.5-VL text encoder) fp8 ~16GB (diffusion) + ~9.4GB (fp8 text encoder) ≈ 25GB total; bf16 needs 24GB+ and "48GB+" per some quant write-ups Direct: kyuz0/amd-r9700-comfy ships a pre-validated "Qwen Image 2512 (FP8) & Lightning LoRA (4 steps)" ComfyUI workflow specifically for the R9700 AI Pro (gfx1201), on a ROCm 7 (TheRock nightlies) toolbox
FLUX.1-dev FLUX.1-dev Non-Commercial License — non-commercial weights, outputs usable commercially 12B bf16 ~24GB; GGUF Q8 ~12-13GB Same family evidence as schnell (gfx1201 Docker images target FLUX generally), but no R9700-specific FLUX.1-dev report found
SD3.5-large Stability Community License — free under $1M annual revenue 8B bf16 ~16GB; 4-bit NF4 fits small GPUs AMD's own ComfyUI-ROCm doc lists an "SD3.5 Simple" template workflow (per prior doc) — vendor-blessed but not R9700-specific
HunyuanImage-3.0 tencent-hunyuan-community license 80B total / 13B active (MoE, 64 experts) Official repo: "≥ 3 × 80GB" VRAM for the base model, "≥ 8 × 80GB" for -Instruct; ~177GB at fp16 None, and actively contraindicated: setup requires CUDA 12.8 + FlashAttention2/FlashInfer, no AMD/ROCm mention anywhere in the official repo
Krea-2 (Turbo) Krea 2 Community License — free under $1M annual revenue, homelab/personal explicitly covered 12-13B DiT No official VRAM figure; community reports (RTX hardware only) cite fp8 ~16GB, GGUF ~12GB None found — released June 22, 2026 per the HF model card; ComfyUI added native support per blog.comfy.org, and GGUF quants exist (molbal/krea2-gguf), but no R9700/gfx1201-specific report exists yet — too new for that evidence to have accumulated

Per-model detail

  • License: Apache-2.0, stated directly on the model card, no revenue threshold, no non-commercial clause, no attribution/naming requirement. The cleanest license of every model considered in this doc or the prior one. Source: Qwen/Qwen-Image on Hugging Face.
  • Architecture: 20B-parameter MMDiT (Multimodal Diffusion Transformer) combined with an 8.3B Qwen2.5-VL text encoder — notably larger and more capable than FLUX.1-schnell's 12B distilled model, particularly for multilingual text rendering and instruction-following, per the official QwenLM/Qwen-Image GitHub repo.
  • ComfyUI support: native, not a wrapper/custom-node integration — landed August 2025 per ComfyUI Wiki's native-support announcement, with an official FP8 checkpoint (qwen_image_fp8_e4m3fn.safetensors) and FP8-scaled text encoder (qwen_2.5_vl_7b_fp8_scaled.safetensors, 9.38GB) published under Comfy-Org/Qwen-Image_ComfyUI.
  • VRAM at FP8: the official Comfy-Org FP8 diffusion checkpoint plus FP8 text encoder land around ~25GB combined — comfortably inside the 32GB ceiling with headroom for ComfyUI runtime/VAE overhead, versus 24GB+ (some sources say 48GB+) for the unquantized bf16 model. Sources: community VRAM write-ups aggregated via Comfy-Org/ComfyUI issue #10852 ("Qwen-image in 24GB VRAM and 32GB RAM") and the official ComfyUI Qwen-Image example — treat exact GB figures as secondary/community-sourced, consistent with how the prior doc flagged FLUX's GGUF VRAM table.
  • Direct R9700/gfx1201 evidence (the deciding factor): kyuz0/amd-r9700-comfy is a Fedora-toolbox ROCm 7 (TheRock nightlies) environment built specifically for the "AMD Radeon 9700 AI PRO (32GB)" and ships a pre-validated "Qwen Image 2512 (FP8) & Lightning LoRA (4 steps)" ComfyUI workflow (plus a Qwen-Image-Edit 2511 FP8 workflow) as one of only four workflows in the whole repo. This is the only model in this comparison with card-architecture-specific (not just "ComfyUI supports ROCm generically") validation — everything else relies on family-level or vendor-generic ROCm claims. A related community discussion (pollockjj/ComfyUI-MultiGPU #133) does flag that "Qwen image edit doesn't always work on AMD HIP/ROCm" in some configurations — worth testing the specific workflow before assuming zero friction, but this is a known-quantity, actively-discussed rough edge rather than a documented hard blocker.
  • 4-step Lightning LoRA: the R9700-validated workflow pairs Qwen-Image with a "Lightning" LoRA for 4-step inference — the same fast-inference pattern FLUX.1-schnell uses, so switching models doesn't have to mean giving up the short GPU-resident-time-per-image property the prior doc called out as valuable for time-sliced use alongside llama-server (even though this ticket's premise is llama-server being stopped, that pattern still helps if the two services are ever run in an overlapping window).

FLUX.1-dev — solid alternative, not the pick

  • License permits non-commercial use of the weights; BFL's own license page states generated outputs are separately usable commercially (already covered in the prior doc; unchanged here). Source: FLUX.1-dev LICENSE.md.
  • At bf16 (~24GB) or GGUF Q8 (~12-13GB), it fits the 32GB ceiling with room to spare — a legitimate move up from schnell's distilled quality. Source: city96/FLUX.1-dev-gguf and community Q8 VRAM reports.
  • Loses to Qwen-Image on two counts: license (non-commercial-weights clause vs. Apache-2.0 — not a hard blocker for this homelab per the prior doc's own reasoning, but strictly worse) and hardware evidence (FLUX's gfx1201 validation is at the family level — the prior doc's yurisasc/comfyui-rocm-rdna4 and charlie12345/R9700AIProComfyUIPatch sources are about running FLUX models on this card generally, not a FLUX.1-dev-specific report the way kyuz0's repo is Qwen-Image-specific).

SD3.5-large — no longer the best use of the freed headroom

  • License unchanged from the prior doc: Stability Community License, free under $1M annual revenue (irrelevant threshold for this homelab). Source: stabilityai/stable-diffusion-3.5-large LICENSE.md.
  • At 8B params (bf16 ~16GB), it's the smallest of the upgrade candidates — which mattered when 6GB was the ceiling, but with 32GB available there's no VRAM reason to pick the model the prior doc already flagged as "lower fidelity than FLUX/SD3.5 by current standards" over Qwen-Image or FLUX.1-dev. AMD's own ComfyUI-ROCm docs do list it as a first-party example template ("SD3.5 Simple", per the prior doc), so it remains a fine fallback if Qwen-Image's FP8 path hits the ROCm rough edge noted above.

HunyuanImage-3.0 — ruled out

  • Official repo states VRAM requirements of "≥ 3 × 80GB" for the base model and "≥ 8 × 80GB" for HunyuanImage-3.0-Instruct — i.e. multi-GPU datacenter-class NVIDIA clusters, not a single 32GB consumer/workstation card at any precision. Source: Tencent-Hunyuan/HunyuanImage-3.0 GitHub repo.
  • Setup instructions require CUDA 12.8, PyTorch 2.8.0 built for CUDA, and optionally FlashAttention2/FlashInfer for MoE routing speed — no AMD or ROCm path is mentioned anywhere in the official repo. Even the MoE-efficient "13B active" framing doesn't help here: the tooling itself assumes an NVIDIA multi-GPU cluster, and the 80B total parameter set still has to be resident somewhere.
  • 32GB of headroom on one AMD card doesn't move this model into reach at any precision considered here; it's excluded regardless of how much VRAM frees up on this specific box.

Krea-2 (Turbo) — promising, but unverified on this hardware

  • Verified directly against primary sources per the ticket's instruction (this is a June 2026 release, past most training cutoffs): the Hugging Face model card states a release date of June 22, 2026, a 12-billion-parameter single-stream diffusion transformer, torch.bfloat16 as the reference precision, and the Krea 2 Community License.
  • License, per krea.ai/krea-2-licensing: non-commercial (including explicitly personal/homelab) use is free; commercial use is permitted royalty-free for entities under $1M trailing-12-month revenue (same shape as SD3.5's and matching this homelab's use case); content-filter and AI-disclosure obligations apply if deployed publicly; derivative model names must start with "Krea".
  • ComfyUI added native support for both open-weight checkpoints (Krea 2 Raw and Krea 2 Turbo) per blog.comfy.org's announcement, and community GGUF quants already exist (molbal/krea2-gguf), with reports (RTX hardware only) of fp8 fitting 16GB and GGUF fitting 12GB.
  • No AMD/ROCm mention anywhere in the model card, and no gfx1201/R9700 community report was found — unsurprising given the model is roughly 2.5 months old at the time of this research. Its architecture (a standard-shaped DiT that ComfyUI loads through its normal diffusion-model nodes, per the ComfyUI blog post) gives reasonable expectation it will run on the same ROCm/PyTorch backend already proven for FLUX and Qwen-Image on this card, but that's an inference, not a verified fact the way kyuz0's Qwen-Image workflow is.
  • Not the pick today, precisely because Qwen-Image already offers a hardware-verified path at a comparable parameter count and VRAM budget. Worth a follow-up research ticket once R9700/gfx1201-specific Krea-2 reports exist — the license and ComfyUI support are both already in place, so the only open question is real-world ROCm behavior.

Recommendation

Qwen-Image, FP8 precision (qwen_image_fp8_e4m3fn.safetensors + qwen_2.5_vl_7b_fp8_scaled.safetensors, ~25GB combined), optionally paired with the 4-step Lightning LoRA the R9700-specific validated workflow uses. It wins on all three axes the ticket asked about:

  1. License: Apache-2.0 — no restriction at all, strictly better than every other candidate including the current FLUX.1-schnell pick.
  2. ROCm/gfx1201 compatibility: the only candidate with a workflow pre-validated specifically on this GPU architecture (kyuz0/amd-r9700-comfy), not just "ComfyUI supports ROCm in general."
  3. VRAM at the ~32GB ceiling: ~25GB at FP8 leaves real margin for ComfyUI runtime/VAE overhead, without needing the multi-step, non-distilled FLUX.1-dev's full 24GB bf16 footprint or accepting SD3.5's lower fidelity ceiling — and it's a genuine capability upgrade over schnell (20B vs. 12B, non-distilled-quality text rendering) rather than just a bigger file.

If the known Qwen-Image/ROCm edit-mode rough edge (pollockjj/ComfyUI-MultiGPU #133) turns out to affect plain text-to-image generation too, SD3.5-large (AMD's own first-party "SD3.5 Simple" ComfyUI-ROCm template workflow) is the fallback, with FLUX.1-dev as a second option. HunyuanImage-3.0 is excluded outright regardless of available VRAM (CUDA-only tooling, multi-GPU datacenter VRAM floor). Krea-2 is worth revisiting once R9700-specific field reports exist.

Sources consulted