Intermediate 11 minImage

Generate images locally: Ollama, Stable Diffusion

The search “ollama image generation” keeps coming up, and the answer fits in one sentence: Ollama doesn’t generate images; it serves text and vision models. But generating images locally, without the cloud or a subscription, is entirely possible—with other tools. This guide provides a verified overview of what Ollama can and cannot do, then lists the real options (Stable Diffusion, ComfyUI, Draw Things on Mac), the VRAM actually required, the simplest installation for your system, and what to expect in terms of quality compared with Midjourney or DALL-E.

By Mohamed Meguedmi·Update 2026-09-08·Tested on Windows, macOS, and Linux

#Ollama and images: what it can and cannot do

Let's start by clearing up the most common confusion behind the “ollama image generation” query. Ollama is a runtime for large language models: it downloads and runs models in GGUF format, exposes an OpenAI-compatible API at http://localhost:11434, and manages loading and unloading in memory. Its domain is text—and, since multimodal models emerged, reading images as input.

The distinction that matters is between understanding an image and creating one. Ollama does the former through vision models such as Qwen VL, Gemma, or Llama Vision: you send a photo, and the model describes it, performs OCR on it, or answers questions about it. That’s text generation from an image. Image generation — producing a PNG file from a prompt — relies on an entirely different architecture (diffusion models, not language transformers), which Ollama does not run.

!
Don’t look for an image model on ollama.com
There is no “ollama run stable-diffusion”. The Ollama library contains only LLMs and vision-language models. Any tutorial claiming to generate images directly with the ollama run se command is using the wrong tool. To create images, you need a dedicated diffusion engine.

In summary: keep Ollama for text and image analysis, and add a diffusion tool alongside it for creation. Both work very well on the same machine and can even share the GPU as long as you do not use them at the same time.


#Why generate your images locally

The Local Image AI Kit

Your first image came out of your machine. The AI Images kit builds what comes next: a text→image pipeline you control (ch. 5), fitting it in your video memory and making it faster (ch. 7), then training your own style (ch. 11).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Generating locally takes a little more effort than opening Midjourney, but the reasons to do it are concrete and quickly add up once you go beyond occasional use.

Zero usage cost
once the hardware and models are in place, you can generate as many images as you want with no credits or subscription. No monthly quota, no per-image billing.
Total privacy
your prompts and images never leave the machine. Decisive for unreleased product visuals, customer content under an NDA, or sensitive data.
No server-side censorship
Cloud service content filters are often broad and opaque. Locally, it's your machine, your rules—with the legal framework still yours.
Control and reproducibility
Fixed seed, exact settings, selected models and LoRAs: you reproduce an image pixel for pixel, which cloud services don’t guarantee from one update to the next.
Offline
once the models are downloaded, everything works offline. Convenient when traveling or on an isolated machine.

#The real local options for image generation

Since Ollama is out of the running for creation, here are the serious tools that actually generate images locally in 2026. They all rely on diffusion models (the Stable Diffusion family, SDXL, or newer architectures such as FLUX); what varies by tool is the interface and target audience.

AUTOMATIC1111 (Stable Diffusion WebUI)
The classic web interface, feature-rich and well documented. txt2img / img2img tabs, countless extensions, and a huge community. A little dated but still the reference for learning.
ComfyUI
an interface based on a node graph: you wire the pipeline (model → sampler → VAE → output) visually. Steeper learning curve, but complete control and the best support for recent models such as FLUX.
Fooocus
designed for Midjourney-style simplicity: a prompt field, presets, and the tool handles the rest. Ideal for getting started without understanding how diffusion works internally.
Draw Things (macOS / iOS)
Native Apple Silicon application, free and optimized for Metal and unified memory. By far the simplest approach on Mac: everything is integrated, including model downloads.
SD.Next
a fork of AUTOMATIC1111 that is more actively maintained, with better multi-backend support (CUDA, ROCm, Intel). A good alternative if you are using a non-NVIDIA GPU.
→
Which tool for whom
Beginner on Windows/Linux → Fooocus. On Mac → Draw Things. Want to learn and tinker → AUTOMATIC1111. Want maximum control and the latest models → ComfyUI. The models (.safetensors) are shared among these tools, so there is nothing stopping you from trying several.

#VRAM required by diffusion model

Image generation is VRAM-intensive, but less linearly than LLMs: what matters is the chosen diffusion model and target resolution. Here are realistic guidelines for comfortable generation in 2026.

Stable Diffusion 1.5 — 4 GB of VRAM
the lightest model. Runs even on a modest card, with native 512×512 images. Dated but fast and highly permissive on hardware.
SDXL — 8 to 12 GB of VRAM
the quality/accessibility standard. Native 1024×1024 images with a good level of detail. A RTX 3060 12GB is more than enough; a RTX 4070 is comfortable.
FLUX (dev / schnell) — 12 to 24 GB of VRAM
Recent-generation models, with significantly better photorealistic quality and prompt adherence. Quantized versions (GGUF, fp8) reduce the footprint, but 16 to 24 GB remains the real sweet spot.
Mac Apple Silicon — unified memory
An M4 Pro 24–48 GB shares its memory between the CPU and GPU: SDXL runs smoothly, and FLUX does too with a little patience. Slower than a high-end RTX, but quiet and energy-efficient.
i
Possible without a dedicated GPU, but slow
Generation on pure CPU works (AUTOMATIC1111's --use-cpu mode, or Draw Things on an entry-level Mac), but expect several minutes per image compared with a few seconds on a GPU. For regular use, a card with at least 8 GB of VRAM changes everything.

#Simplest installation for your OS

The shortest path differs by machine. Here’s the least painful option for each system, without aiming for completeness.

#macOS (Apple Silicon): Draw Things

On Mac, there is no need to touch Python or a command line. Draw Things is a native app downloadable from the Mac App Store: it handles model installation, Metal optimization, and the interface in a single package.

  1. 01
    1. Install the application
    Search for “Draw Things” on the Mac App Store and install it. It's free.
  2. 02
    2. Download a model
    At first launch, the app offers to download a base model (SDXL is a good starting point). The download is handled through the interface.
  3. 03
    3. Generate
    Enter a prompt, leave the default settings, and click Generate. The first image appears within a few dozen seconds, depending on the chip.

#Windows / Linux with NVIDIA GPU: Fooocus or ComfyUI

On a PC with a NVIDIA card, Fooocus is the fastest to get up and running. It downloads its default model on first launch and exposes almost no settings to configure.

Terminal — install Fooocus
# Prérequis : Python 3.10+ et un GPU NVIDIA avec pilotes récents
git clone https://github.com/lllyasviel/Fooocus.git
cd Fooocus
python -m venv venv
source venv/bin/activate    # Windows : venv\Scripts\activate
pip install -r requirements_versions.txt
python entry_with_update.py

At launch, Fooocus automatically downloads an SDXL model, then opens a web interface in the browser (by default at http://127.0.0.1:7865). You enter a prompt, generate—that's it.

If you prefer node-based control and the latest models, ComfyUI launches in a similar way:

Terminal — install ComfyUI
git clone https://github.com/comfyanonymous/ComfyUI.git
cd ComfyUI
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt
# Placez vos modèles .safetensors dans models/checkpoints/
python main.py
# Interface : http://127.0.0.1:8188
!
AMD GPU on Windows: more complicated
ROCm support on Windows remains uneven. On AMD GPUs, prefer Linux with ROCm, or a DirectML/SD.Next variant on Windows. Expect some configuration—the same kind of setup required to run an LLM on an AMD GPU.

#Your first results in 30 minutes

Once the tool is installed, here’s how to quickly get good images instead of diving into advanced settings from the start.

  1. 01
    1. Start with an SDXL model
    This is the best quality / VRAM / speed tradeoff for getting started. Save FLUX for later, when you want more realism and know your GPU's limits.
  2. 02
    2. Write a descriptive prompt
    Diffusion models like concrete descriptions: subject, style, framing, and lighting. For example, “a sleeping red fox in a birch forest, golden morning light, photography, shallow depth of field.”
  3. 03
    3. Leave the default settings
    Steps around 25–30, guidance (CFG) around 7 for SDXL. Don't change anything else for your first attempts: the tool's default values are reasonable.
  4. 04
    4. Set a seed to iterate
    An identical seed plus a slightly modified prompt lets you see the effect of one word. That is the real workflow: change one detail at a time.
  5. 05
    5. Switch to img2img for refinement
    Use a generated image as a base and regenerate it with partial denoising to make corrections without starting from scratch.
→
The negative prompt does half the work
Most tools offer a negative prompt field: list what you do not want there ("blur, deformed hands, text, watermark, low resolution"). With SD 1.5 and SDXL, this is often what separates a failed image from a clean one.

#Expected quality compared with Midjourney and DALL-E

Let's be direct: for raw “one-click” quality, Midjourney is still ahead. Its images are aesthetically pleasing by default, with no tuning required. DALL-E, integrated into ChatGPT, excels at following a complex prompt and rendering text in the image. That's the price of massive models trained and tuned by dedicated teams and served from the cloud.

But the gap is no longer what you might expect. Local FLUX produces photorealistic images that rival Midjourney for many subjects, and a well-prompted SDXL is more than enough for most use cases. Above all, local deployment wins on dimensions the cloud cannot offer.

Fine-grained control
ControlNet (enforcing a pose, depth, or contours), precise inpainting, specialized LoRAs: local models offer a level of control that Midjourney and DALL-E do not expose.
Customization
Thousands of community models and LoRAs (styles, characters, aesthetics) can be downloaded freely. You adapt the model to your exact needs.
Volume and cost
Generating a thousand variants costs only electricity. In the cloud, that means the same amount of credits consumed.
Initial effort
That is the real tradeoff: Midjourney produces a beautiful result immediately, while local generation requires learning how to prompt and tune. The return on that effort is control and free usage.

The right framing: local tools do not replace Midjourney for a one-off attractive image with no fuss. They surpass it as soon as you need privacy, volume, reproducibility, or precise control—and their state-of-the-art quality with FLUX is no longer marginal.


#Common troubleshooting

“CUDA out of memory”
the model or resolution exceeds your VRAM. Lower the resolution, enable the tool's “medvram” / “lowvram” mode, or switch to a lighter model (SDXL → SD 1.5, or quantized FLUX).
Blurry or distorted images
increase the number of steps slightly, adjust the CFG, and above all strengthen the negative prompt. With SDXL, make sure you’re generating at 1024×1024 rather than 512.
Very slow generation
confirm that the GPU is actually being used (rather than the CPU by default). On NVIDIA, check the drivers and the CUDA-enabled PyTorch version; the Fooocus/ComfyUI installer normally handles this.
Misshapen hands and faces
A classic weakness of diffusion models. Use a face-restoration module, inpainting on the affected area, or a model/LoRA known to handle anatomy better.

#Go further

Image generation naturally fits alongside the rest of a local AI stack. These site guides build on what you just set up:

Analyze images rather than create them
“Local multimodal vision LLM: analyze images with Ollama” covers the other half of the image topic—reading and OCR with a vision model.
Understanding VRAM and GPU selection
The memory guidelines in this guide follow the same logic as “Choosing Your Quantization (Q4, Q5, Q8, FP16),” which is useful for sizing your card.
The rest of the text stack
“Install Ollama in 5 minutes” remains the foundation for the language component, running alongside your diffusion engine on the same machine.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.