ControlNet: controlling the composition of your images
ControlNet adds a visual condition to a diffusion model (outlines, depth, pose) that fixes the composition while the prompt controls the style. Its method won the best paper award at ICCV 2023. It runs in ComfyUI, Forge, or AUTOMATIC1111 and exists for SD 1.5, SDXL, SD 3.5 Large, Flux.1, and Z-Image Turbo. Note: a control works only with the model family it was trained for.
A prompt describes content; it does not describe a position. ControlNet conditions image generation on a reference structure, pose, outline, or depth map so the result follows a composition you define. Updated September 28, 2026, this guide explains which control to choose, how to set its strength and application range, how much the files weigh for each model family, and the current state of the official offering, which has changed considerably since SD 1.5.
#The problem it solves
Text is a poor language for geometry. You describe a subject, style, and mood in a prompt, but words can’t lock down the position of a hand, the angle of a building, or the exact silhouette of a product. Rerun with a different seed and the composition changes completely. ControlNet changes the input, not the prompt: you provide an image that encodes the desired structure, the model is conditioned on it at every denoising step, and the prompt governs everything else.
The method comes from a research paper: “Adding Conditional Control to Text-to-Image Diffusion Models,” by Lvmin Zhang, Anyi Rao, and Maneesh Agrawala, published in the ICCV 2023 proceedings and awarded the conference’s best paper prize, the Marr Prize. Its abstract describes the principle: ControlNet locks the diffusion model, reuses its pretrained encoding layers as a foundation, and connects the trainable copy using “zero convolutions”—convolution layers initialized to zero that gradually grow the parameters and prevent harmful noise from disrupting training. The authors also show that training remains robust with small datasets of fewer than 50,000 images, as well as with large datasets of more than one million.
What this changes for you: you add a control to an existing model without retraining it, and each type of control is a separate file. That’s also why you need the right file for the right model family.
#Choose the control type
AI images and videos on your own machine, no subscription and no credits: ComfyUI, Flux, Z-Image and Wan 2.2 with ready-to-load workflows, VRAM tiers, LoRA training and the legal frame.
- Lifetime online access
- PDF + files
- Lifetime updates
| Control | What it captures | What for | ControlNet 1.1 model for SD 1.5 |
|---|---|---|---|
| Contours (canny, softedge, lineart) | Lines and silhouettes | Preserve an exact shape: a product, a logo, a coloring page | control_v11p_sd15_canny, _softedge, _lineart, _lineart_anime |
| Depth | Distance from the camera | Keep the volume and perspective while changing materials and details | control_v11f1p_sd15_depth |
| Prompt | A skeleton of joints | The characters: the posture is fixed, everything else is open | control_v11p_sd15_openpose |
| Segmentation | Regions by category | A scene's layout: sky here, building there | control_v11p_sd15_seg |
| Sketch (scribble) | A rough sketch | Turn a hand-drawn sketch into a rendered image | control_v11p_sd15_scribble |
| Normal | Surface orientation | Relief and details sensitive to lighting | control_v11p_sd15_normalbae |
| Tile | The content of an image region | Add detail by making it larger | control_v11f1e_sd15_tile |
The official ControlNet 1.1 repository lists 14 models, including 11 production-ready and 3 experimental models, all named according to a single rule. For depth, it lists three supported preprocessors: Depth_Midas, Depth_Leres, and Depth_Zoe. A preprocessor is the tool that transforms your reference image into a control map, such as turning a photo into a depth map.
The practical rule: choose the loosest control that still captures what matters to you. Conditioning on outlines in a photograph reproduces the photograph: if that's the desired result, generation was unnecessary. Depth or pose generally provide a better balance between adherence and creativity.
#The settings that matter
Names vary by tool, but three settings appear everywhere. The table gives their names in the three most common environments, based on their documentation.
| Setting | ComfyUI (Apply ControlNet node) | diffusers library | AUTOMATIC1111 extension |
|---|---|---|---|
| Control strength | strength: the higher the value, the more strongly the control influences the image | controlnet_conditioning_scale: the weight assigned to the control | Control Weight: the weight of the control's influence, comparable to emphasizing a word in the prompt |
| Application timing | start_percent and end_percent: 0.2 means control starts at 20% denoising, while 0.8 means it stops at 80% | control_guidance_start and control_guidance_end, as a fraction of the steps (0 to 1) | Start and End: the fraction of the generation where control starts and stops |
| Card accuracy | Preprocessor resolution, set in the preprocessor node | Control image size | Pixel-Perfect mode calculates the best preprocessor resolution automatically |
Strength is the most effective lever: near the maximum, the result is rigid, and lowering the strength fixes most frozen-looking images. Stopping the control before denoising finishes locks in the composition while leaving the model free to render materials and light. The AUTOMATIC1111 extension adds a three-position Control Mode: Balanced, My prompt is more important, and ControlNet is more important, to arbitrate between the prompt and the control.
#Starting values published by the model authors
Rather than making up numbers, start with the ones provided by the publishers. They vary significantly from one family to another, which is a reminder that a setting that works for SD 1.5 cannot simply be carried over.
| Model | Recommended setting | Source |
|---|---|---|
| Flux.1 dev, Union Pro 2.0 (Shakker Labs), canny, or soft edge | Strength 0.7; end at 0.8 | Model card |
| Flux.1 dev, Union Pro 2.0, depth | Strength 0.8; end at 0.8 | Model card |
| Flux.1 dev, Union Pro 2.0, pose | Force 0.9; end at 0.65 | Model card |
| Flux.1 dev, Union Pro 2.0, grayscale | Strength 0.9; ending at 0.8 | Model card |
| Z-Image Turbo, Fun ControlNet Union | control_context_scale between 0.65 and 0.80 | Model card |
| Flux.1 dev, InstantX ControlNet canny | Strength of 0,5 in the official example | diffusers documentation |
| SDXL, two combined controls (Canny and depth) | 0.5 for each in the official example | diffusers documentation |
#Which ControlNet for which family
The landscape has changed significantly since SD 1.5, and many guides are outdated. The table distinguishes what the model lab publishes from what the community publishes, because this affects maintenance and licensing.
| Family | Controls published by the model vendor | Community or third-party controls |
|---|---|---|
| SD 1.5 | ControlNet 1.1 by lllyasviel: 14 models (canny, depth, pose, segmentation, sketch, tile…) | Many derivative models |
| SDXL | Stability AI's Control-LoRA: canny, depth, recolor, and sketch | ControlNet canny and depth from the diffusers team; xinsir's ControlNet Union (one file, multiple modes) |
| SD 3.5 Large | Three ControlNets from Stability AI, published on November 26, 2024: Blur, Canny, and Depth | Few comparable references |
| Flux.1 dev | Black Forest Labs' FLUX.1 Canny and Depth, released in November 2024 in both full-model and LoRA versions | InstantX ControlNet Union (beta version) and Shakker Labs Union Pro 2.0 |
| Z-Image Turbo | Aucun sur la page Hugging Face de l'éditeur (Tongyi-MAI) lors de notre vérification | Alibaba PAI’s Fun ControlNet Union: canny, HED, depth, pose, and MLSD |
Two points deserve mention. First, Black Forest Labs has marked FLUX.1 Depth and FLUX.1 Canny as deprecated for its API, including the LoRA versions: the open weights remain available on Hugging Face, but the publisher no longer maintains them. Second, “Union” models combine several modes in a single file, saving disk space at the cost of variable maturity: InstantX’s ControlNet Union listing for Flux describes it as a first beta version that may not be fully trained.
The license does not automatically follow the base model: each control file has its own, and it varies from one file to another.
| Control | License | Key takeaway |
|---|---|---|
| ControlNet 1.1 for SD 1.5 | OpenRAIL | RAIL-type license with usage restrictions |
| ControlNet canny SDXL from the diffusers team | OpenRAIL++ | Like SDXL: usage restrictions listed in the appendix |
| ControlNet Union SDXL by xinsir | Apache 2.0 | Permissive |
| Stability AI’s Control-LoRA | Stability AI Control-LoRA Community License | Above 1 million active users per month, a Stability AI license is required for commercial use |
| ControlNets SD 3.5 Large | Stability AI Community License | Free for commercial use under 1 million dollars in annual revenue |
| FLUX.1 Canny and Depth [dev] | Non-commercial FluxDev license | The product sheet specifies that generated images may be used for personal, scientific, and commercial purposes according to the license |
| InstantX and Shakker Labs’ Union Pro 2.0 combined, for Flux | Non-commercial FluxDev license | Read before any commercial use |
| Fun ControlNet Union for Z-Image Turbo | Apache 2.0 | Permissive |
#What it costs in files and memory
A ControlNet is an additional network that runs alongside the base model: it adds VRAM usage and time per image. File sizes, published on Hugging Face, provide the first-order estimate: they are added to the weights of the base model and its text encoder.
| Control | File size |
|---|---|
| ControlNet 1.1 for SD 1.5, original version (.pth) | 1.45 GB each |
| SDXL, ControlNet canny from the diffusers team | 5 GB at full precision, 2.5 GB at half precision |
| SDXL, xinsir’s ControlNet Union | 2.51 GB |
| SDXL, Stability AI Control-LoRA | About 738 MB, compared with 4.7 GB for the original ControlNet |
| SD 3.5 Large, ControlNet canny | 8.61 GB |
| Flux.1 dev, InstantX ControlNet Union | 6.6 GB |
| Flux.1 dev, Union Pro 2.0 from Shakker Labs | 4.28 GB |
| Flux.1 dev, FLUX.1 Depth in LoRA (Black Forest Labs) | 1.24 GB, versus 23.8 GB for the full Canny model |
| Z-Image Turbo, Fun ControlNet Union | 3.1 GB |
Two takeaways. First, a control module in LoRA form is a fraction of the size of a full model: Stability AI presents its Control-LoRA models as a way to bring control to a wider range of consumer graphics cards. Second, watch the format: ControlNet 1.1 .pth files are Python pickles, which Hugging Face flags with a pickle-import warning. Prefer .safetensors versions, which cannot execute code.
How much additional VRAM? Vendors do not publish a general figure, and the measurement depends on the model and resolution. If memory is saturated, the remedies should be applied in this order: lower the resolution, reduce the number of simultaneous controls, then switch to a lighter control, such as a Control-LoRA or a half-precision version. For baseline memory, see the VRAM guide.
#Where to use it
ControlNet is not a standalone application: it runs inside an image-generation interface. In ComfyUI, two nodes do the work: Load ControlNet, which reads files from the models/controlnet folder, and Apply ControlNet, which receives the control image and the settings from the table above. The ComfyUI documentation states that the software core does not include all preprocessors: you often need an extension such as ComfyUI ControlNet aux. Remember that an extension is code executed with your permissions: install only extensions you know. For Z-Image Turbo, place the control file in a different folder, models/model_patches.
In the classic web interface, AUTOMATIC1111's sd-webui-controlnet extension adds a panel below the prompt. Its latest release, version 1.1.454, dates from July 2024, as does its latest commit: it still works with SD 1.5 and SDXL, but will not keep up with newer models. Forge integrates ControlNet directly, according to its README, without using this extension.
#Combine multiple controls at once
Nothing prevents you from stacking two controls, such as a pose control for the character’s posture and depth for the background. The diffusers documentation gives the rule: for good results, mask the conditioning inputs so they do not overlap, and experiment with different strengths to adjust the weight of each control. Its SDXL example combines canny and depth at 0.5 each. In ComfyUI, chain Apply ControlNet nodes. The same principles apply to both: stack only controls that address different aspects of the image, and lower each control’s strength because the effects add up.
#When control is ignored
- The control does not appear in the list
- The file is in the wrong folder, or its configuration file doesn't have the same name as the file itself. In ComfyUI, the folder is models/controlnet.
- The output ignores the reference
- A mismatch between the base model family and the control family, a preprocessor that produced an empty map, or a strength that is too low. Display the control map itself before suspecting anything else.
- The images are static and lifeless
- Strength too high, or control applied throughout the denoising process. Lower the strength and stop the control before the end.
- Disappointing result with a Union model
- The authors of InstantX's Union write that even a fully trained Union model can be worse than specialized models, such as the pose model. Try a dedicated ControlNet, and write a detailed prompt, as Shakker Labs' model card recommends.
- Memory saturated
- Lower the resolution, remove one control, or use a lighter version of the control.
- ComfyUI: install the node-based interface
- AUTOMATIC1111: the classic web interface
- Overview of local image generation
- The AI image toolkit: ControlNet workflows and settings by graphics card
- Source: the official ControlNet repository
- Source: the original research paper, ICCV 2023
- Source: Stability AI’s Control-LoRA models for SDXL
- Source: FLUX.1 Tools and the deprecation of Depth and Canny
#FAQ
What is ControlNet for?+
Do you need a powerful graphics card for ControlNet?+
Why does my output ignore the control image?+
Which control for a character's pose?+
Does ControlNet exist for SDXL, Flux, and Z-Image?+
Can you use multiple ControlNets at once?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.