Qwen-Image-Edit Locally: ComfyUI and VRAM necessary
Qwen-Image-Edit is Alibaba's open image-editing model: give it a photo and a natural-language instruction, and it modifies the image while preserving everything else. It runs locally in ComfyUI, but it is a 20-billion-parameter model, and the question that blocks everyone is the same: which version should you download, FP8 or GGUF, and how much VRAM should you plan for? This guide follows ComfyUI's official workflow, explains where to find file sizes and VRAM requirements in the source materials, and lists the node errors that occur most often. It does not cover installing ComfyUI itself, which is covered in a dedicated guide.
#Why Qwen-Image-Edit, and what the family contains
The Qwen-Image family covers two use cases. Qwen-Image generates an image from text, with a recognized strength noted in its technical report: rendering text in the image, including entire sentences and non-Latin alphabets. Qwen-Image-Edit starts from the same model to modify an existing image. The GitHub repository QwenLM/Qwen-Image and the Hugging Face page Qwen/Qwen-Image-Edit distinguish two types of editing: semantic editing, which changes the content or style while preserving the subject's identity, and appearance editing, which adds, removes, or replaces an element without affecting the rest. The model can also correct or replace text in an image while preserving the font and layout.
Architecture and licensing are the two key facts to remember before downloading anything. The diffusion model is a multimodal transformer (MMDiT) with approximately 20 billion parameters. The text encoder is not a small CLIP: it is Qwen2.5-VL 7B, a full vision-language model that reads both your prompt and the input image. Both are released under the Apache 2.0 license, which permits commercial use of the weights. Each component is downloaded separately, and this distribution across three files causes most of the node errors described below.
The lineage is evolving quickly. The first version of Qwen-Image-Edit was released in August 2025. A version dated 2509 followed in September 2025: it accepts multiple input images and natively understands ControlNet-style conditions, such as a depth map, edges, or a pose. Version 2511 arrived at the end of 2025 with better subject consistency, while the generation model itself received a 2512 revision. Versions 2.x were announced in 2026. This guide follows the lineage documented step by step by ComfyUI, from the August 2025 version through 2511: the nodes and file organization are the same. Before choosing, read the top of the GitHub README, which lists the versions and their release dates.
#Native, FP8, or GGUF: which version of Qwen-Image-Edit should you download
AI images and videos on your own machine, no subscription and no credits: ComfyUI, Flux, Z-Image and Wan 2.2 with ready-to-load workflows, VRAM tiers, LoRA training and the legal frame.
- Lifetime online access
- PDF + files
- Lifetime updates
Three versions of the same model are circulating, and they are not used the same way.
- Native (bf16, Diffusers format)
- The Qwen/Qwen-Image-Edit repository on Hugging Face, split across several safetensors files and intended for the Diffusers library in Python. In bf16, 20 billion parameters represent approximately 40 GB for the transformer alone, plus the text encoder. This is the route for Python scripts and 48 GB or larger cards, or for partially offloading to RAM. Not the ComfyUI route.
- FP8 repacked by Comfy Org
- The Comfy-Org/Qwen-Image-Edit_ComfyUI repository on Hugging Face provides the weights in a single FP8 file (e4m3fn format), with the text encoder and VAE split into subdirectories. This is the version used by ComfyUI's official workflows. The diffusion model file weighs a little over 20 GB. An approximately 41 GB bf16 variant is available in the same repository for cards with enough space.
- Community GGUF
- GGUF-format conversions, the format for quantized LLMs, published by contributors such as city96, QuantStack, or Unsloth. They go down to Q8_0, Q6_K, Q5_K_M, Q4_K_M, Q3_K_M, or Q2_K, and require the ComfyUI-GGUF extension. This is the most realistic option under 16 GB of VRAM, and the most comfortable on a Mac.
The selection rule is simple. 24 GB card: official FP8, no questions asked. 16 GB card: FP8 works thanks to ComfyUI’s memory offloading, but a Q6_K or Q5_K_M GGUF avoids slowdowns. 12 GB card: Q4_K_M GGUF or lower. Mac Apple Silicon: GGUF, because ComfyUI does not compute in FP8 on Metal and dequantizes these weights when loading them, doubling the memory footprint.
#Required VRAM: where to find the figures and how to interpret them
Neither the GitHub README nor the Hugging Face page for Qwen provides a table showing how much VRAM each card needs. The ComfyUI documentation does not either, for this model: its tutorials list the files to download, the destination folder, and the workflow, and sometimes provide the hardware and timing from their own tests, which should be treated as a reference, not a promise. The reliable, dated, verifiable data is the size of each file displayed on its Hugging Face page. Everything else follows from that.
- The diffusion model
- About 20 GB in FP8 (the qwen_image_edit_fp8_e4m3fn.safetensors file from the Comfy-Org repository, published in August 2025), about 41 GB in bf16. In GGUF, the weight size is calculated as follows: 20 billion parameters multiplied by the number of bits per parameter, divided by eight. An 8.5-bit Q8_0 is about 21 GB, a 4.5- to 5-bit Q4_K_M is about 12 GB, a Q3_K_M is about 9 to 10 GB, and a Q2_K is about 7 GB. Each GGUF repository page displays the exact value for each file.
- The text encoder
- The Comfy-Org repository’s qwen_2.5_vl_7b_fp8_scaled.safetensors file weighs about 9.4 GB. It doesn’t need to remain in VRAM during denoising: ComfyUI loads it there to encode the prompt and image, then unloads it if space runs out. That’s why a 20 GB model and a 9 GB encoder can run on a 24 GB card.
- The VAE
- The qwen_image_vae.safetensors file is about 250 MB. Negligible in memory, but essential and specific to this family: a Flux or SDXL VAE produces colored noise.
- The activations
- The computation itself requires additional space beyond the weights, and more at higher resolutions. The official workflow reduces the input image to about one megapixel, which limits this portion. Double the output resolution and this margin grows.
ComfyUI enables dynamic VRAM management by default, as described in our installation guide: when the weights exceed the card's capacity, part of them remains in RAM and is transferred as needed. The model can then run on a 12 or 16 GB card, but each transfer takes time, and system RAM becomes the real ceiling: with a 20 GB model and a 9 GB encoder, 32 GB of RAM is the minimum and 64 GB is comfortable. The reference table below summarizes the logic without claiming to be a measurement.
- 24 GB and up (RTX 3090, 4090, 5090)
- Full official FP8, FP8 encoder, one-megapixel output image with no offloading. The 41 GB bf16 version remains reserved for professional cards with 48 GB.
- 16 GB (RTX 4080, 5080, 4060 Ti 16 GB)
- FP8 with offloading, or GGUF Q6_K to Q5_K_M models that fit on the card with the encoder offloaded between steps.
- 12 GB (RTX 3060 12 GB, 4070, 5070, 5070 Ti)
- GGUF Q4_K_M or Q3_K_M, FP8 encoder loaded and then unloaded, 32 GB of RAM. Expect generations to be noticeably longer than with SDXL.
- 8 GB
- Possible in Q2_K or Q3 with extensive offloading, at the cost of reduced speed and quality. For occasional editing only.
- Mac Apple Silicon
- Unified memory serves as VRAM. A 12 GB GGUF Q4_K_M plus the 9 GB encoder requires a 32 GB Mac; 48 GB or more for comfortable work.
#Prerequisites
- Up-to-date ComfyUI
- The nodes specific to Qwen-Image, such as TextEncodeQwenImageEdit, arrived in ComfyUI in August 2025, and those for version 2509 in September 2025. An older installation ignores them and displays red nodes. Update first: Desktop updates itself, the portable archive uses the update script, and the manual installation uses git pull.
- A NVIDIA card with at least 12 GB, or a Mac with 32 GB
- Below that, editing remains feasible but becomes an exercise in patience. AMD works on Linux with ROCm and, since January 2026, on Windows with the Desktop application.
- Room on the disk
- FP8: a little over 30 GB for the model, encoder, and VAE trio. Add 10 to 21 GB for each GGUF variant you try.
- A working ComfyUI installation
- First image generated with SDXL or Flux. If that is not the case, start with the installation guide: this guide begins when ComfyUI opens on http://127.0.0.1:8188.
#Steps 1 and 2: download the right files and organize them
The shortest path is ComfyUI's built-in workflow model library: Workflow menu, Browse Models, Image category, then the Qwen-Image-Edit model, or its 2509 or 2511 variant depending on which version you want. ComfyUI detects missing files and offers to download them directly into the correct folder. If you prefer to download manually, for example to control the variant you choose, here's the expected organization as documented in the Qwen-Image-Edit tutorial on docs.comfy.org.
- 01Download the diffusion modelOn Hugging Face, repository Comfy-Org/Qwen-Image-Edit_ComfyUI, folder split_files/diffusion_models. Choose the FP8 file corresponding to the desired version. Check the displayed size on the page before clicking, then compare it with the received file: an interrupted download produces a truncated file and an unreadable error when loading.
- 02Download the text encoder and VAESame repository, split_files/text_encoders and split_files/vae directories. The encoder is shared by Qwen-Image and all versions of Qwen-Image-Edit: a single copy is enough, even if you install multiple model versions. The same applies to the VAE.
- 03Put each file in its own subfolderModel in diffusion_models, encoder in text_encoders, VAE in vae. With ComfyUI Desktop, the models folder is located wherever you chose during installation, as shown in the application settings. A file placed in checkpoints instead of diffusion_models does not appear in the loading node.
- 04Optional: Lightning LoRAThe lightx2v/Qwen-Image-Lightning repository publishes LoRAs that reduce generation to 4 or 8 steps instead of 20, for both Qwen-Image and Qwen-Image-Edit. File in models/loras. Later versions of the editing model have their own Lightning LoRAs: use the one bearing your version's name.
- 05Refresh ComfyUIAfter copying the files, click the interface refresh button or reload the page. The loading nodes reread the folders at that point, not continuously.
#Step 3: load the editing workflow and understand its nodes
Open the Qwen-Image-Edit workflow model from the library. If you downloaded the files manually, ComfyUI may offer to download them again: decline, then select your files in the loading nodes. The graph has about ten nodes, and each has a specific role.
- Load Diffusion Model
- Loads the file from models/diffusion_models. The dropdown should display your Qwen-Image-Edit file; if it is empty, the file is in the wrong location or the interface has not been refreshed.
- Load CLIP, type qwen_image
- Loads the Qwen2.5-VL encoder from models/text_encoders. The type field must be qwen_image. With any other type, the encoder loads without errors, but the prompt is misinterpreted and you get nonsense.
- Load VAE
- The Qwen-Image VAE. Nothing else.
- Load Image, then Scale Image to Total Pixels
- Your image to be edited, resized to about one megapixel. This resizing is not optional: the model was trained at this scale, and a much larger input reduces coherence while increasing VRAM usage.
- TextEncodeQwenImageEdit
- The node specific to this family. It receives the encoder, VAE, image, and your prompt, and produces the conditioning. There are two: one for the positive prompt and one for the negative prompt. For version 2509 and later, the node is called TextEncodeQwenImageEditPlus and accepts up to three images.
- ModelSamplingAuraFlow and CFGNorm
- Two denoising-control nodes, preconfigured in the official workflow, including a shift around 3. Don't remove them: without them, images come out washed out or oversaturated.
- KSampler, VAE Decode, Save Image
- Denoising, decoding, and saving, as in any ComfyUI workflow. The image is saved in the output folder with the complete workflow embedded in the PNG.
Write the prompt as an instruction, not as a description of the final image: “replace the text on the sign with OUVERT, keep the same font,” “change the jacket to red leather, keep everything else,” “remove the person on the left.” The model primarily understands Chinese and English; French often works, but English remains more reliable for precise instructions. Leave the negative prompt empty or very short.
#Step 4: settings and the Lightning LoRA for speed
The official workflow starts with 20 steps, a CFG of 2.5, the euler sampler, and the simple scheduler. These values are a good starting point: this model’s CFG is low by design, and raising it to 7 as with SDXL produces saturated, distorted images. Change the seed first to get variants, then increase the number of steps if the result lacks fine detail.
- 01First edition without LoRARun the workflow as is on a simple image with a short instruction. Watch the terminal: the first run loads 30 GB of weights from disk, which can take several minutes. Later runs reuse what is in memory.
- 02Add Lightning LoRAInsert a LoraLoaderModelOnly node between Load Diffusion Model and ModelSamplingAuraFlow, select the Lightning 4-step LoRA, strength 1. Set KSampler to 4 steps and CFG to 1.0: these LoRAs are trained to work without guidance. The number of steps is divided by five, so denoising time is reduced by the same proportion, at the possible cost of reduced fidelity in fine details and text, which the LoRA repository does not quantify.
- 03Set the seed for comparisonSet KSampler to a fixed seed before comparing two prompts or two model variants. Otherwise, you’ll attribute differences caused by randomness to the settings.
- 04Chain multiple editsTo correct it in multiple passes, reload the output image as a new input. Each pass goes through the VAE again and loses a little detail: limit yourself to two or three, and keep the original.
#Step 5: the GGUF variant for 12–16 GB graphics cards and Macs
ComfyUI doesn't read GGUF natively: you need city96's ComfyUI-GGUF extension, available in ComfyUI-Manager or by cloning it with Git into the custom_nodes folder. It adds dedicated loader nodes, and the organization logic remains the same.
- 01Choosing the quantizationIn the selected GGUF repository—for example, city96/Qwen-Image-Edit-gguf or QuantStack/Qwen-Image-Edit-2509-GGUF—the page lists each file with its size. Choose the largest one that leaves about 2 GB of headroom in your VRAM after accounting for the target resolution: Q6_K or Q5_K_M for 16 GB, Q4_K_M for 12 GB. As with LLMs, Q4_K_M is the usual compromise; below Q3, text and small details suffer.
- 02File awayPlace the diffusion model’s GGUF in models/diffusion_models, or in models/unet, which the extension also reads. Keep the text encoder and VAE from step 1, in safetensors.
- 03Replace the loading nodeIn the official workflow, remove Load Diffusion Model and replace it with Unet Loader (GGUF). Connect its model output wherever the old one went. The rest of the graph remains unchanged: TextEncodeQwenImageEdit, VAE, KSampler—all work identically.
- 04Keep the encoder in safetensorsGGUFs for Qwen2.5-VL 7B exist, but those produced for llama.cpp are designed for dialogue and do not necessarily include the vision component that editing needs to read the input image. With the official qwen_2.5_vl_7b_fp8_scaled.safetensors file, you eliminate one possible source of errors. Switch to a GGUF encoder only if the repository explicitly indicates compatibility with Qwen-Image-Edit in ComfyUI-GGUF.
On Mac, the same procedure applies, with one difference: monitor memory pressure in Activity Monitor. When unified memory overflows, macOS writes to the SSD and each image takes several times longer, without displaying an error.
#Frequent node errors with Qwen-Image-Edit
- Red nodes: TextEncodeQwenImageEdit or TextEncodeQwenImageEditPlus not found
- ComfyUI is too old. These nodes are part of ComfyUI's core, not an extension: no manager will find them. Update ComfyUI and restart.
- Prompt outputs failed validation, value not in list
- The file selected in a loading node does not exist in the expected folder. This happens when you open a downloaded workflow whose filenames differ from yours. Reopen the dropdown menu for each loading node and select your file.
- The qwen_image type does not appear in Load CLIP
- Same cause: a version of ComfyUI from before August 2025. Update it.
- CUDA out of memory at the start of denoising
- The weights exceed the VRAM, and offloading was not enough. In order: reduce the output resolution, switch to a smaller GGUF quantization, close other applications using the GPU, then try the ComfyUI memory options described in our installation guide, such as reserving VRAM or disabling dynamic VRAM.
- Output image made of colored noise or black
- VAE or encoder from another family. Check that Load VAE points to qwen_image_vae.safetensors and Load CLIP to the Qwen2.5-VL encoder with the qwen_image type.
- The image isn't modified, or barely modified
- The prompt is too vague, or CFG dropped to 1 without LoRA Lightning. Raise CFG to 2.5 without LoRA, rewrite the prompt as a direct instruction, and verify that the image is connected to the TextEncodeQwenImageEdit node, not just to VAE Encode.
- The rewritten text is false or unreadable
- Put the desired text in quotation marks in the prompt, specifying that the font must be preserved. The most aggressive quantizations, Q2 and Q3, impair this capability first: switch back to Q4 or FP8 for text editing.
- The version 2509 workflow accepts only one image
- You loaded the old TextEncodeQwenImageEdit node. Replace it with TextEncodeQwenImageEditPlus, which exposes the image1, image2, and image3 inputs.
- Unet Loader (GGUF) missing after installing the extension
- Python dependencies aren't installed, or ComfyUI hasn't been restarted. Run pip install again on the extension's requirements.txt file in the ComfyUI environment, then restart. At startup, the terminal displays the list of loaded extensions and any import errors.
#Go further
- Install ComfyUI and understand its memory options
- Desktop, laptop, or manual setup, model directories, dynamic VRAM, and command-line options. https://quelllm.fr/guide/comfyui-installation-guide-debutant
- Generating images locally: the landscape
- Stable Diffusion, Flux, ComfyUI, Draw Things on Mac: which option for which hardware. https://quelllm.fr/guide/generer-images-en-local-guide
- ControlNet: master composition
- Layout, contours, depth: what version 2509 of Qwen-Image-Edit natively integrates, explained on Stable Diffusion. https://quelllm.fr/guide/controlnet-guide-stable-diffusion
- VRAM: what it is and how much you need
- Read your card’s memory, understand offloading to RAM, and choose a card. https://quelllm.fr/guide/vram-c-est-quoi-combien-pour-ia
- Official sources
- GitHub repository QwenLM/Qwen-Image (version list and dates), Hugging Face page Qwen/Qwen-Image-Edit (Apache 2.0 license, Diffusers example), Comfy-Org/Qwen-Image-Edit_ComfyUI and Comfy-Org/Qwen-Image_ComfyUI repositories (FP8 files, sizes), Qwen-Image and Qwen-Image-Edit tutorials on docs.comfy.org (directory structure, workflows), github.com/city96/ComfyUI-GGUF extension.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.