Real-ESRGAN: upscale your images in local
Real-ESRGAN is a free model (BSD-3-Clause) that enlarges an image by ×4 by reconstructing plausible detail. The simplest option is the portable ncnn-vulkan executable, with no CUDA or PyTorch, for Intel, AMD, or NVIDIA cards. Choose the model based on the image: realesrgan-x4plus for photos, and the anime variant for illustrations. The project has not released a version since September 2022, and the added detail is invented: never use it for text or evidence.
Real-ESRGAN enlarges an image by reconstructing plausible detail instead of stretching pixels. It is a small, fast, open-source model that has become the component added at the end of a local image-generation pipeline or used separately to restore old photos. This guide, current as of September 28, 2026, explains which model to choose, how to run it without installing PyTorch, what it invents, which errors to expect, and how the project compares with recent alternatives.
#What an AI upscaler does
Conventional upscaling interpolates: it calculates intermediate pixels from neighboring pixels, producing a larger, blurrier image. A machine-learning upscaler does something else. It was trained to recover a sharp image from a degraded version, and it reconstructs what most resembles the original: textures, edges, grain. The difference in one sentence: interpolation doesn't create information; the model invents it. That's exactly what it's asked to do, which is why you need to know where not to use it.
Real-ESRGAN comes from a paper by Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan (Tencent ARC Lab), submitted to arXiv in July 2021 and presented at ICCV Workshops: « Training Real-World Blind Super-Resolution with Pure Synthetic Data ». Its abstract describes the idea: a model trained exclusively on synthetic data, with a model of “higher-order” degradations that better simulates real-world defects (blur, noise, compression), and a U-Net discriminator. “Blind” means the model doesn't need to know what degradation the image underwent. The repository has more than 36,000 stars.
#The models, and which one to choose
AI images and videos on your own machine, no subscription and no credits: ComfyUI, Flux, Z-Image and Wan 2.2 with ready-to-load workflows, VRAM tiers, LoRA training and the legal frame.
- Lifetime online access
- PDF + files
- Lifetime updates
| Model | Scale | Project description | Use on |
|---|---|---|---|
| RealESRGAN_x4plus | ×4 | ×4 model for general images | Photos and generated images: the default choice |
| RealESRGAN_x2plus | ×2 | ×2 model for general images | When ×4 is too much, without going through resizing again |
| RealESRNet_x4plus | ×4 | ×4 model with MSE loss, effects of excessive smoothing | Rarely: the result is smooth and bland |
| realesr-general-x4v3 | ×4, usable as ×1, ×2, ×3 | Small model, uses much less GPU memory and time, with less denoising capability | Modest machines, batches of images; the -dn option controls denoising |
| RealESRGAN_x4plus_anime_6B | ×4 | Optimized for anime images, smaller network (6 RRDB blocks) | Illustrations, boards, 2D drawings |
| realesr-animevideov3 | ×4, usable as ×1, ×2, ×3 | XS-size model for animated videos | Animation sequences |
The project's catalog separates models for photos from models for illustrations, and this is the choice that has the greatest impact on the result. The anime variant is presented as optimized for this type of image, with a smaller network, and the project compares it with waifu2x. For a drawing, try it first: a photo model may texture areas of solid color that should remain smooth, so check the result on your image.
#Three ways to run it
The README describes three paths: an online demo, the portable ncnn executable, and the Python script. For regular local use, only the latter two matter.
| Path | What you need | Advantages | Limitations |
|---|---|---|---|
| Portable executable (ncnn-vulkan) | An Intel, AMD, or NVIDIA GPU compatible with Vulkan. Archives for Windows, Linux, and macOS. No CUDA, no PyTorch. | Nothing to install; the models are included | Does not support all Python script options, such as --outscale, and may create inconsistencies between blocks |
| Python script (inference_realesrgan.py) | Python 3.7 or later, PyTorch 1.7 or later, basicsr, plus facexlib and gfpgan for faces | Arbitrary scaling, faces, tiling, alpha-channel images, grayscale, or 16-bit images | Fragile dependencies; see troubleshooting |
#The exact commands
The portable executable is used from the command line. The help text lists the options: -i for input (file or folder), -o for output, -s for scale (2, 3, or 4), -n for the model, -t for tile size, -g to select the GPU, and -f for output format.
The Python script provides the --outscale option: the model always performs a ×4 upscale, then the program resizes the result to the desired scale using an inexpensive resize operation. There is also --tile to split the image and save memory, and --fp32 for full precision, which becomes mandatory on a processor (see below).
In ComfyUI, upscaling models go in ComfyUI/models/upscale_models, and a loader node followed by one upscaling node per model applies them. For a desktop interface, the README lists several projects that bundle Real-ESRGAN, including Upscayl, free software for Linux, macOS, and Windows, with around 50,000 stars and an update in the past few days.
#Use it effectively
- 01Choose the model based on the image typePhoto or illustration: this is the setting that changes the result the most, well before the upscaling factor. Always add -n to the command.
- 02Don’t stack passesUpscaling twice in a row by ×4 produces plastic, not detail. One pass at the right factor is enough; if ×4 is too much, use the ×2 model or the scaling option.
- 03Clean up before scaling upLight denoising beforehand prevents the model from treating noise as detail to reconstruct. With realesr-general-x4v3, the -dn option controls denoising strength to prevent an overly smooth result.
- 04Reserve face_enhance for real facesThe face_enhance option relies on GFPGAN and is intended only for real faces: the project's FAQ advises against using it on animation, where it needlessly consumes GPU memory.
- 05Compare at 100%Judge an upscale at the pixel level, not from a thumbnail. Many spectacular results at small sizes look smooth and artificial up close.
#What size should you target: calculate before launching
A ×4 enlargement is useful only if the final size serves a specific purpose. Here is the necessary factor for each destination, based on a simple calculation: pixels = centimeters ÷ 2.54 × resolution in dots per inch, then factor = target long side ÷ starting long side.
| Usage | Large target side | Required factor |
|---|---|---|
| 4K screen, full screen | 3,840 px | ×3.75: one ×4 pass is sufficient |
| A4 printing at 300 dpi | 3,508 px | ×3,4: a ×4 pass, or --outscale 3,4 |
| A3 printing at 300 dpi | 4,961 px | ×4.8: exceeds a ×4 pass; accept a slightly smaller draw rather than stacking two passes |
| Web page or thumbnail | 1,600 to 2,000 px | ×1.6 to ×2: the ×2 model or --outscale is enough |
Keep the principle in mind: target the useful size, not the maximum size. A 1,024 × 1,024 image enlarged ×4 becomes 4,096 × 4,096, or about 16.8 million pixels—sixteen times more than the original: that is sixteen times as much data to store, load, and edit, for a gain that is often invisible on screen. With the Python script’s --outscale option, you choose the exact scale instead of over-enlarging and then reducing.
#In an image-generation pipeline
For local generation, upscaling is the final step: generate at a resolution the card can support, then upscale, which uses less memory than generating directly at a large size. The native sizes of image models—for example, 512 × 512 for SD 1.5 or 1024 × 1024 for SDXL—are detailed in the Stable Diffusion guide. Then, ×2 upscaling on a 1024 × 1024 SDXL image produces 2048 × 2048 without running it through the diffusion model again.
- Stable Diffusion: versions, native resolutions, and licenses
- Generating images locally: the landscape
#What it makes up
- The text
- Enlarged small print turns into plausible but false characters. For a document, you need optical character recognition, not a magnifier.
- Faces
- An enlarged blurry face becomes a sharp face that is not exactly the same person. No identification use is acceptable.
- The fine details
- Fabric patterns, foliage, hair: the model produces a credible texture, not the original texture.
- The drawbacks
- A clearly visible compression artifact may sometimes be enlarged with the application instead of being erased.
#Troubleshooting: common errors
| Symptom | Cause | Patch |
|---|---|---|
| Error “slow_conv2d_cpu not implemented for Half” | Real-ESRGAN uses half precision (fp16) by default, which some operators do not support on the CPU | Add the --fp32 option to the command |
| ModuleNotFoundError: torchvision.transforms.functional_tensor | basicsr imports a module that torchvision removed starting with version 0.17 (August 28, 2024 report; a fix is proposed in a merge request in the BasicSR repository) | Install a compatible version of torchvision, or use the portable ncnn executable, which does not depend on it |
| Visible blocks or seams in the image | The executable splits the image into tiles and then reassembles them, which the README flags as a source of inconsistencies | Set the tile size with -t (0 for automatic): larger tiles reduce seams when memory allows |
| Black image on output | The ncnn repository TODO notes that some PCs produce black images | Try the other route (Python or ncnn) or another card with -g |
| Distorted faces in an illustration | face_enhance option enabled on non-photographic content | Remove it: it is designed for real faces |
#Where the project stands in 2026
You have to be honest about its age. The latest released version, v0.3.0, dates from September 20, 2022, and the last commit to the main repository, from April 2, 2024, removes a continuous integration configuration file. The ncnn implementation, whose README states an MIT license, has not been updated since May 2024. That does not make the model bad: it does its job, it is small, and it runs on modest cards. But you cannot count on fixes for recent incompatibilities, such as the one with torchvision.
Among the alternatives, ComfyUI's README lists SeedVR2 and SUPIR among its supported models. SeedVR2 presents itself as one-step video restoration through adversarial diffusion post-training, while SUPIR targets photorealistic restoration of “in the wild” images. These are diffusion-based approaches: we do not have a reliable numerical comparison with Real-ESRGAN, and the risk of invented detail is at least as high. Read SUPIR's license, which GitHub classifies as “Other,” before commercial use. For most everyday photos and illustrations, Real-ESRGAN remains the best balance of simplicity, speed, and hardware requirements.
- ControlNet: control composition upfront
- The AI Images Kit
- Source: the official Real-ESRGAN repository
- Source: the research paper on arXiv
- Source: the ncnn-vulkan implementation
#FAQ
Is Real-ESRGAN free?+
Do you need a graphics card?+
Can scanned text be enlarged?+
Why is my illustration grainy after enlargement?+
Can you use it with video?+
Is Real-ESRGAN still maintained?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.