Generate images locally: Ollama, Stable Diffusion
The search “ollama image generation” keeps coming up, and the answer fits in one sentence: Ollama doesn’t generate images; it serves text and vision models. But generating images locally, without the cloud or a subscription, is entirely possible—with other tools. This guide provides a verified overview of what Ollama can and cannot do, then lists the real options (Stable Diffusion, ComfyUI, Draw Things on Mac), the VRAM actually required, the simplest installation for your system, and what to expect in terms of quality compared with Midjourney or DALL-E.
#Ollama and images: what it can and cannot do
Let's start by clearing up the most common confusion behind the “ollama image generation” query. Ollama is a runtime for large language models: it downloads and runs models in GGUF format, exposes an OpenAI-compatible API at http://localhost:11434, and manages loading and unloading in memory. Its domain is text—and, since multimodal models emerged, reading images as input.
The distinction that matters is between understanding an image and creating one. Ollama does the former through vision models such as Qwen VL, Gemma, or Llama Vision: you send a photo, and the model describes it, performs OCR on it, or answers questions about it. That’s text generation from an image. Image generation — producing a PNG file from a prompt — relies on an entirely different architecture (diffusion models, not language transformers), which Ollama does not run.
In summary: keep Ollama for text and image analysis, and add a diffusion tool alongside it for creation. Both work very well on the same machine and can even share the GPU as long as you do not use them at the same time.
#Why generate your images locally
Your first image came out of your machine. The AI Images kit builds what comes next: a text→image pipeline you control (ch. 5), fitting it in your video memory and making it faster (ch. 7), then training your own style (ch. 11).
- Lifetime online access
- PDF + files
- Lifetime updates
Generating locally takes a little more effort than opening Midjourney, but the reasons to do it are concrete and quickly add up once you go beyond occasional use.
- Zero usage cost
- once the hardware and models are in place, you can generate as many images as you want with no credits or subscription. No monthly quota, no per-image billing.
- Total privacy
- your prompts and images never leave the machine. Decisive for unreleased product visuals, customer content under an NDA, or sensitive data.
- No server-side censorship
- Cloud service content filters are often broad and opaque. Locally, it's your machine, your rules—with the legal framework still yours.
- Control and reproducibility
- Fixed seed, exact settings, selected models and LoRAs: you reproduce an image pixel for pixel, which cloud services don’t guarantee from one update to the next.
- Offline
- once the models are downloaded, everything works offline. Convenient when traveling or on an isolated machine.
#The real local options for image generation
Since Ollama is out of the running for creation, here are the serious tools that actually generate images locally in 2026. They all rely on diffusion models (the Stable Diffusion family, SDXL, or newer architectures such as FLUX); what varies by tool is the interface and target audience.
- AUTOMATIC1111 (Stable Diffusion WebUI)
- The classic web interface, feature-rich and well documented. txt2img / img2img tabs, countless extensions, and a huge community. A little dated but still the reference for learning.
- ComfyUI
- an interface based on a node graph: you wire the pipeline (model → sampler → VAE → output) visually. Steeper learning curve, but complete control and the best support for recent models such as FLUX.
- Fooocus
- designed for Midjourney-style simplicity: a prompt field, presets, and the tool handles the rest. Ideal for getting started without understanding how diffusion works internally.
- Draw Things (macOS / iOS)
- Native Apple Silicon application, free and optimized for Metal and unified memory. By far the simplest approach on Mac: everything is integrated, including model downloads.
- SD.Next
- a fork of AUTOMATIC1111 that is more actively maintained, with better multi-backend support (CUDA, ROCm, Intel). A good alternative if you are using a non-NVIDIA GPU.
#VRAM required by diffusion model
Image generation is VRAM-intensive, but less linearly than LLMs: what matters is the chosen diffusion model and target resolution. Here are realistic guidelines for comfortable generation in 2026.
- Stable Diffusion 1.5 — 4 GB of VRAM
- the lightest model. Runs even on a modest card, with native 512×512 images. Dated but fast and highly permissive on hardware.
- SDXL — 8 to 12 GB of VRAM
- the quality/accessibility standard. Native 1024×1024 images with a good level of detail. A RTX 3060 12GB is more than enough; a RTX 4070 is comfortable.
- FLUX (dev / schnell) — 12 to 24 GB of VRAM
- Recent-generation models, with significantly better photorealistic quality and prompt adherence. Quantized versions (GGUF, fp8) reduce the footprint, but 16 to 24 GB remains the real sweet spot.
- Mac Apple Silicon — unified memory
- An M4 Pro 24–48 GB shares its memory between the CPU and GPU: SDXL runs smoothly, and FLUX does too with a little patience. Slower than a high-end RTX, but quiet and energy-efficient.
#Simplest installation for your OS
The shortest path differs by machine. Here’s the least painful option for each system, without aiming for completeness.
#macOS (Apple Silicon): Draw Things
On Mac, there is no need to touch Python or a command line. Draw Things is a native app downloadable from the Mac App Store: it handles model installation, Metal optimization, and the interface in a single package.
- 011. Install the applicationSearch for “Draw Things” on the Mac App Store and install it. It's free.
- 022. Download a modelAt first launch, the app offers to download a base model (SDXL is a good starting point). The download is handled through the interface.
- 033. GenerateEnter a prompt, leave the default settings, and click Generate. The first image appears within a few dozen seconds, depending on the chip.
#Windows / Linux with NVIDIA GPU: Fooocus or ComfyUI
On a PC with a NVIDIA card, Fooocus is the fastest to get up and running. It downloads its default model on first launch and exposes almost no settings to configure.
At launch, Fooocus automatically downloads an SDXL model, then opens a web interface in the browser (by default at http://127.0.0.1:7865). You enter a prompt, generate—that's it.
If you prefer node-based control and the latest models, ComfyUI launches in a similar way:
#Your first results in 30 minutes
Once the tool is installed, here’s how to quickly get good images instead of diving into advanced settings from the start.
- 011. Start with an SDXL modelThis is the best quality / VRAM / speed tradeoff for getting started. Save FLUX for later, when you want more realism and know your GPU's limits.
- 022. Write a descriptive promptDiffusion models like concrete descriptions: subject, style, framing, and lighting. For example, “a sleeping red fox in a birch forest, golden morning light, photography, shallow depth of field.”
- 033. Leave the default settingsSteps around 25–30, guidance (CFG) around 7 for SDXL. Don't change anything else for your first attempts: the tool's default values are reasonable.
- 044. Set a seed to iterateAn identical seed plus a slightly modified prompt lets you see the effect of one word. That is the real workflow: change one detail at a time.
- 055. Switch to img2img for refinementUse a generated image as a base and regenerate it with partial denoising to make corrections without starting from scratch.
#Expected quality compared with Midjourney and DALL-E
Let's be direct: for raw “one-click” quality, Midjourney is still ahead. Its images are aesthetically pleasing by default, with no tuning required. DALL-E, integrated into ChatGPT, excels at following a complex prompt and rendering text in the image. That's the price of massive models trained and tuned by dedicated teams and served from the cloud.
But the gap is no longer what you might expect. Local FLUX produces photorealistic images that rival Midjourney for many subjects, and a well-prompted SDXL is more than enough for most use cases. Above all, local deployment wins on dimensions the cloud cannot offer.
- Fine-grained control
- ControlNet (enforcing a pose, depth, or contours), precise inpainting, specialized LoRAs: local models offer a level of control that Midjourney and DALL-E do not expose.
- Customization
- Thousands of community models and LoRAs (styles, characters, aesthetics) can be downloaded freely. You adapt the model to your exact needs.
- Volume and cost
- Generating a thousand variants costs only electricity. In the cloud, that means the same amount of credits consumed.
- Initial effort
- That is the real tradeoff: Midjourney produces a beautiful result immediately, while local generation requires learning how to prompt and tune. The return on that effort is control and free usage.
The right framing: local tools do not replace Midjourney for a one-off attractive image with no fuss. They surpass it as soon as you need privacy, volume, reproducibility, or precise control—and their state-of-the-art quality with FLUX is no longer marginal.
#Common troubleshooting
- “CUDA out of memory”
- the model or resolution exceeds your VRAM. Lower the resolution, enable the tool's “medvram” / “lowvram” mode, or switch to a lighter model (SDXL → SD 1.5, or quantized FLUX).
- Blurry or distorted images
- increase the number of steps slightly, adjust the CFG, and above all strengthen the negative prompt. With SDXL, make sure you’re generating at 1024×1024 rather than 512.
- Very slow generation
- confirm that the GPU is actually being used (rather than the CPU by default). On NVIDIA, check the drivers and the CUDA-enabled PyTorch version; the Fooocus/ComfyUI installer normally handles this.
- Misshapen hands and faces
- A classic weakness of diffusion models. Use a face-restoration module, inpainting on the affected area, or a model/LoRA known to handle anatomy better.
#Go further
Image generation naturally fits alongside the rest of a local AI stack. These site guides build on what you just set up:
- Analyze images rather than create them
- “Local multimodal vision LLM: analyze images with Ollama” covers the other half of the image topic—reading and OCR with a vision model.
- Understanding VRAM and GPU selection
- The memory guidelines in this guide follow the same logic as “Choosing Your Quantization (Q4, Q5, Q8, FP16),” which is useful for sizing your card.
- The rest of the text stack
- “Install Ollama in 5 minutes” remains the foundation for the language component, running alongside your diffusion engine on the same machine.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.