Intermediate 13 minImage

ControlNet: controlling the composition of your images

Direct response

ControlNet adds a visual condition to a diffusion model (outlines, depth, pose) that fixes the composition while the prompt controls the style. Its method won the best paper award at ICCV 2023. It runs in ComfyUI, Forge, or AUTOMATIC1111 and exists for SD 1.5, SDXL, SD 3.5 Large, Flux.1, and Z-Image Turbo. Note: a control works only with the model family it was trained for.

A prompt describes content; it does not describe a position. ControlNet conditions image generation on a reference structure, pose, outline, or depth map so the result follows a composition you define. Updated September 28, 2026, this guide explains which control to choose, how to set its strength and application range, how much the files weigh for each model family, and the current state of the official offering, which has changed considerably since SD 1.5.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#The problem it solves

Text is a poor language for geometry. You describe a subject, style, and mood in a prompt, but words can’t lock down the position of a hand, the angle of a building, or the exact silhouette of a product. Rerun with a different seed and the composition changes completely. ControlNet changes the input, not the prompt: you provide an image that encodes the desired structure, the model is conditioned on it at every denoising step, and the prompt governs everything else.

The method comes from a research paper: “Adding Conditional Control to Text-to-Image Diffusion Models,” by Lvmin Zhang, Anyi Rao, and Maneesh Agrawala, published in the ICCV 2023 proceedings and awarded the conference’s best paper prize, the Marr Prize. Its abstract describes the principle: ControlNet locks the diffusion model, reuses its pretrained encoding layers as a foundation, and connects the trainable copy using “zero convolutions”—convolution layers initialized to zero that gradually grow the parameters and prevent harmful noise from disrupting training. The authors also show that training remains robust with small datasets of fewer than 50,000 images, as well as with large datasets of more than one million.

What this changes for you: you add a control to an existing model without retraining it, and each type of control is a separate file. That’s also why you need the right file for the right model family.

#Choose the control type

The Local Image AI Kit

AI images and videos on your own machine, no subscription and no credits: ComfyUI, Flux, Z-Image and Wan 2.2 with ready-to-load workflows, VRAM tiers, LoRA training and the legal frame.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
What each control captures, using the SD 1.5 models from ControlNet 1.1
ControlWhat it capturesWhat forControlNet 1.1 model for SD 1.5
Contours (canny, softedge, lineart)Lines and silhouettesPreserve an exact shape: a product, a logo, a coloring pagecontrol_v11p_sd15_canny, _softedge, _lineart, _lineart_anime
DepthDistance from the cameraKeep the volume and perspective while changing materials and detailscontrol_v11f1p_sd15_depth
PromptA skeleton of jointsThe characters: the posture is fixed, everything else is opencontrol_v11p_sd15_openpose
SegmentationRegions by categoryA scene's layout: sky here, building therecontrol_v11p_sd15_seg
Sketch (scribble)A rough sketchTurn a hand-drawn sketch into a rendered imagecontrol_v11p_sd15_scribble
NormalSurface orientationRelief and details sensitive to lightingcontrol_v11p_sd15_normalbae
TileThe content of an image regionAdd detail by making it largercontrol_v11f1e_sd15_tile

The official ControlNet 1.1 repository lists 14 models, including 11 production-ready and 3 experimental models, all named according to a single rule. For depth, it lists three supported preprocessors: Depth_Midas, Depth_Leres, and Depth_Zoe. A preprocessor is the tool that transforms your reference image into a control map, such as turning a photo into a depth map.

The practical rule: choose the loosest control that still captures what matters to you. Conditioning on outlines in a photograph reproduces the photograph: if that's the desired result, generation was unnecessary. Depth or pose generally provide a better balance between adherence and creativity.

#The settings that matter

Names vary by tool, but three settings appear everywhere. The table gives their names in the three most common environments, based on their documentation.

ControlNet settings by tool
SettingComfyUI (Apply ControlNet node)diffusers libraryAUTOMATIC1111 extension
Control strengthstrength: the higher the value, the more strongly the control influences the imagecontrolnet_conditioning_scale: the weight assigned to the controlControl Weight: the weight of the control's influence, comparable to emphasizing a word in the prompt
Application timingstart_percent and end_percent: 0.2 means control starts at 20% denoising, while 0.8 means it stops at 80%control_guidance_start and control_guidance_end, as a fraction of the steps (0 to 1)Start and End: the fraction of the generation where control starts and stops
Card accuracyPreprocessor resolution, set in the preprocessor nodeControl image sizePixel-Perfect mode calculates the best preprocessor resolution automatically

Strength is the most effective lever: near the maximum, the result is rigid, and lowering the strength fixes most frozen-looking images. Stopping the control before denoising finishes locks in the composition while leaving the model free to render materials and light. The AUTOMATIC1111 extension adds a three-position Control Mode: Balanced, My prompt is more important, and ControlNet is more important, to arbitrate between the prompt and the control.

#Starting values published by the model authors

Rather than making up numbers, start with the ones provided by the publishers. They vary significantly from one family to another, which is a reminder that a setting that works for SD 1.5 cannot simply be carried over.

Recommended values, based on model cards and the diffusers documentation
ModelRecommended settingSource
Flux.1 dev, Union Pro 2.0 (Shakker Labs), canny, or soft edgeStrength 0.7; end at 0.8Model card
Flux.1 dev, Union Pro 2.0, depthStrength 0.8; end at 0.8Model card
Flux.1 dev, Union Pro 2.0, poseForce 0.9; end at 0.65Model card
Flux.1 dev, Union Pro 2.0, grayscaleStrength 0.9; ending at 0.8Model card
Z-Image Turbo, Fun ControlNet Unioncontrol_context_scale between 0.65 and 0.80Model card
Flux.1 dev, InstantX ControlNet cannyStrength of 0,5 in the official examplediffusers documentation
SDXL, two combined controls (Canny and depth)0.5 for each in the official examplediffusers documentation
!
The control model must match the base model
A ControlNet is trained for a specific architecture. A file intended for SD 1.5 does not work with SDXL or Flux. The usual symptom is an output that completely ignores the control or degenerates into noise. Always check the family listed in the file's details before downloading it.

#Which ControlNet for which family

The landscape has changed significantly since SD 1.5, and many guides are outdated. The table distinguishes what the model lab publishes from what the community publishes, because this affects maintenance and licensing.

State of the offering by model family, as of writing
FamilyControls published by the model vendorCommunity or third-party controls
SD 1.5ControlNet 1.1 by lllyasviel: 14 models (canny, depth, pose, segmentation, sketch, tile…)Many derivative models
SDXLStability AI's Control-LoRA: canny, depth, recolor, and sketchControlNet canny and depth from the diffusers team; xinsir's ControlNet Union (one file, multiple modes)
SD 3.5 LargeThree ControlNets from Stability AI, published on November 26, 2024: Blur, Canny, and DepthFew comparable references
Flux.1 devBlack Forest Labs' FLUX.1 Canny and Depth, released in November 2024 in both full-model and LoRA versionsInstantX ControlNet Union (beta version) and Shakker Labs Union Pro 2.0
Z-Image TurboAucun sur la page Hugging Face de l'éditeur (Tongyi-MAI) lors de notre vérificationAlibaba PAI’s Fun ControlNet Union: canny, HED, depth, pose, and MLSD

Two points deserve mention. First, Black Forest Labs has marked FLUX.1 Depth and FLUX.1 Canny as deprecated for its API, including the LoRA versions: the open weights remain available on Hugging Face, but the publisher no longer maintains them. Second, “Union” models combine several modes in a single file, saving disk space at the cost of variable maturity: InstantX’s ControlNet Union listing for Flux describes it as a first beta version that may not be fully trained.

The license does not automatically follow the base model: each control file has its own, and it varies from one file to another.

Control licenses, based on the metadata of the Hugging Face repositories and the cited licenses
ControlLicenseKey takeaway
ControlNet 1.1 for SD 1.5OpenRAILRAIL-type license with usage restrictions
ControlNet canny SDXL from the diffusers teamOpenRAIL++Like SDXL: usage restrictions listed in the appendix
ControlNet Union SDXL by xinsirApache 2.0Permissive
Stability AI’s Control-LoRAStability AI Control-LoRA Community LicenseAbove 1 million active users per month, a Stability AI license is required for commercial use
ControlNets SD 3.5 LargeStability AI Community LicenseFree for commercial use under 1 million dollars in annual revenue
FLUX.1 Canny and Depth [dev]Non-commercial FluxDev licenseThe product sheet specifies that generated images may be used for personal, scientific, and commercial purposes according to the license
InstantX and Shakker Labs’ Union Pro 2.0 combined, for FluxNon-commercial FluxDev licenseRead before any commercial use
Fun ControlNet Union for Z-Image TurboApache 2.0Permissive

#What it costs in files and memory

A ControlNet is an additional network that runs alongside the base model: it adds VRAM usage and time per image. File sizes, published on Hugging Face, provide the first-order estimate: they are added to the weights of the base model and its text encoder.

Control file sizes, according to Hugging Face and Stability AI
ControlFile size
ControlNet 1.1 for SD 1.5, original version (.pth)1.45 GB each
SDXL, ControlNet canny from the diffusers team5 GB at full precision, 2.5 GB at half precision
SDXL, xinsir’s ControlNet Union2.51 GB
SDXL, Stability AI Control-LoRAAbout 738 MB, compared with 4.7 GB for the original ControlNet
SD 3.5 Large, ControlNet canny8.61 GB
Flux.1 dev, InstantX ControlNet Union6.6 GB
Flux.1 dev, Union Pro 2.0 from Shakker Labs4.28 GB
Flux.1 dev, FLUX.1 Depth in LoRA (Black Forest Labs)1.24 GB, versus 23.8 GB for the full Canny model
Z-Image Turbo, Fun ControlNet Union3.1 GB

Two takeaways. First, a control module in LoRA form is a fraction of the size of a full model: Stability AI presents its Control-LoRA models as a way to bring control to a wider range of consumer graphics cards. Second, watch the format: ControlNet 1.1 .pth files are Python pickles, which Hugging Face flags with a pickle-import warning. Prefer .safetensors versions, which cannot execute code.

How much additional VRAM? Vendors do not publish a general figure, and the measurement depends on the model and resolution. If memory is saturated, the remedies should be applied in this order: lower the resolution, reduce the number of simultaneous controls, then switch to a lighter control, such as a Control-LoRA or a half-precision version. For baseline memory, see the VRAM guide.

#Where to use it

ControlNet is not a standalone application: it runs inside an image-generation interface. In ComfyUI, two nodes do the work: Load ControlNet, which reads files from the models/controlnet folder, and Apply ControlNet, which receives the control image and the settings from the table above. The ComfyUI documentation states that the software core does not include all preprocessors: you often need an extension such as ComfyUI ControlNet aux. Remember that an extension is code executed with your permissions: install only extensions you know. For Z-Image Turbo, place the control file in a different folder, models/model_patches.

In the classic web interface, AUTOMATIC1111's sd-webui-controlnet extension adds a panel below the prompt. Its latest release, version 1.1.454, dates from July 2024, as does its latest commit: it still works with SD 1.5 and SDXL, but will not keep up with newer models. Forge integrates ControlNet directly, according to its README, without using this extension.

#Combine multiple controls at once

Nothing prevents you from stacking two controls, such as a pose control for the character’s posture and depth for the background. The diffusers documentation gives the rule: for good results, mask the conditioning inputs so they do not overlap, and experiment with different strengths to adjust the weight of each control. Its SDXL example combines canny and depth at 0.5 each. In ComfyUI, chain Apply ControlNet nodes. The same principles apply to both: stack only controls that address different aspects of the image, and lower each control’s strength because the effects add up.

#When control is ignored

The control does not appear in the list
The file is in the wrong folder, or its configuration file doesn't have the same name as the file itself. In ComfyUI, the folder is models/controlnet.
The output ignores the reference
A mismatch between the base model family and the control family, a preprocessor that produced an empty map, or a strength that is too low. Display the control map itself before suspecting anything else.
The images are static and lifeless
Strength too high, or control applied throughout the denoising process. Lower the strength and stop the control before the end.
Disappointing result with a Union model
The authors of InstantX's Union write that even a fully trained Union model can be worse than specialized models, such as the pose model. Try a dedicated ControlNet, and write a detailed prompt, as Shakker Labs' model card recommends.
Memory saturated
Lower the resolution, remove one control, or use a lighter version of the control.

#FAQ

FAQ
What is ControlNet for?+
Condition image generation on a structural reference—edges, depth, or a pose skeleton—so the image follows that structure while the prompt controls the subject and style. The method, published in the ICCV 2023 proceedings, freezes the diffusion model and trains a connected copy using zero-initialized convolutions: the original model is not retrained.
Do you need a powerful graphics card for ControlNet?+
It adds a second network to the base model, which requires memory and time. Publishers do not provide a general VRAM overhead figure. On SD 1.5, a control weighs 1.45 GB in the original version; on SDXL, a Stability AI Control-LoRA weighs approximately 738 MB versus 4.7 GB. Start with a single lightweight control, then scale up.
Why does my output ignore the control image?+
Most often, the control model does not match the base model family (SD 1.5, SDXL, Flux), the preprocessor produced a blank map, or the strength is too low. First, display the control map itself. With newer models, the application range also matters: if control stops too early, it has no further effect.
Which control for a character's pose?+
Pose conditioning, which fixes the joint skeleton while leaving clothing, appearance, and setting free. On Flux with Shakker Labs’ Union Pro 2.0 model, the documentation recommends a strength of 0.9 and stopping at 0.65 denoising for pose. Edges would lock the entire silhouette, which is rarely the goal for a character.
Does ControlNet exist for SDXL, Flux, and Z-Image?+
Yes, but the source varies. SDXL: Stability AI's Control-LoRA and third-party models. SD 3.5 Large: three ControlNets from Stability AI. Flux.1: FLUX.1 Canny and Depth from Black Forest Labs, deprecated for the API, and community Union models. Z-Image Turbo: Alibaba PAI's Fun ControlNet Union, with a recommended value of 0.65 to 0.80.
Can you use multiple ControlNets at once?+
Yes. The diffusers documentation recommends masking the conditioning inputs so they don’t overlap and adjusting the strength of each; its SDXL example uses 0.5 for canny and 0.5 for depth. In ComfyUI, you chain Apply ControlNet nodes. Combine only complementary controls, such as pose and setting.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.