Beginner 11 minImage

Real-ESRGAN: upscale your images in local

Direct response

Real-ESRGAN is a free model (BSD-3-Clause) that enlarges an image by ×4 by reconstructing plausible detail. The simplest option is the portable ncnn-vulkan executable, with no CUDA or PyTorch, for Intel, AMD, or NVIDIA cards. Choose the model based on the image: realesrgan-x4plus for photos, and the anime variant for illustrations. The project has not released a version since September 2022, and the added detail is invented: never use it for text or evidence.

Real-ESRGAN enlarges an image by reconstructing plausible detail instead of stretching pixels. It is a small, fast, open-source model that has become the component added at the end of a local image-generation pipeline or used separately to restore old photos. This guide, current as of September 28, 2026, explains which model to choose, how to run it without installing PyTorch, what it invents, which errors to expect, and how the project compares with recent alternatives.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#What an AI upscaler does

Conventional upscaling interpolates: it calculates intermediate pixels from neighboring pixels, producing a larger, blurrier image. A machine-learning upscaler does something else. It was trained to recover a sharp image from a degraded version, and it reconstructs what most resembles the original: textures, edges, grain. The difference in one sentence: interpolation doesn't create information; the model invents it. That's exactly what it's asked to do, which is why you need to know where not to use it.

Real-ESRGAN comes from a paper by Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan (Tencent ARC Lab), submitted to arXiv in July 2021 and presented at ICCV Workshops: « Training Real-World Blind Super-Resolution with Pure Synthetic Data ». Its abstract describes the idea: a model trained exclusively on synthetic data, with a model of “higher-order” degradations that better simulates real-world defects (blur, noise, compression), and a U-Net discriminator. “Blind” means the model doesn't need to know what degradation the image underwent. The repository has more than 36,000 stars.

#The models, and which one to choose

The Local Image AI Kit

AI images and videos on your own machine, no subscription and no credits: ComfyUI, Flux, Z-Image and Wan 2.2 with ready-to-load workflows, VRAM tiers, LoRA training and the legal frame.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Models from the official repository, according to its Model Zoo
ModelScaleProject descriptionUse on
RealESRGAN_x4plus×4×4 model for general imagesPhotos and generated images: the default choice
RealESRGAN_x2plus×2×2 model for general imagesWhen ×4 is too much, without going through resizing again
RealESRNet_x4plus×4×4 model with MSE loss, effects of excessive smoothingRarely: the result is smooth and bland
realesr-general-x4v3×4, usable as ×1, ×2, ×3Small model, uses much less GPU memory and time, with less denoising capabilityModest machines, batches of images; the -dn option controls denoising
RealESRGAN_x4plus_anime_6B×4Optimized for anime images, smaller network (6 RRDB blocks)Illustrations, boards, 2D drawings
realesr-animevideov3×4, usable as ×1, ×2, ×3XS-size model for animated videosAnimation sequences

The project's catalog separates models for photos from models for illustrations, and this is the choice that has the greatest impact on the result. The anime variant is presented as optimized for this type of image, with a smaller network, and the project compares it with waifu2x. For a drawing, try it first: a photo model may texture areas of solid color that should remain smooth, so check the result on your image.

!
The portable program's default model isn't the one you'd expect
The main README lists realesrgan-x4plus as the portable executable’s default model, but the program’s help text defaults to realesr-animevideov3, the animation model. On a photo, this model produces an unsuitable result. Always specify the model with the -n option instead of relying on the default.

#Three ways to run it

The README describes three paths: an online demo, the portable ncnn executable, and the Python script. For regular local use, only the latter two matter.

The two local paths, according to the README
PathWhat you needAdvantagesLimitations
Portable executable (ncnn-vulkan)An Intel, AMD, or NVIDIA GPU compatible with Vulkan. Archives for Windows, Linux, and macOS. No CUDA, no PyTorch.Nothing to install; the models are includedDoes not support all Python script options, such as --outscale, and may create inconsistencies between blocks
Python script (inference_realesrgan.py)Python 3.7 or later, PyTorch 1.7 or later, basicsr, plus facexlib and gfpgan for facesArbitrary scaling, faces, tiling, alpha-channel images, grayscale, or 16-bit imagesFragile dependencies; see troubleshooting

#The exact commands

The portable executable is used from the command line. The help text lists the options: -i for input (file or folder), -o for output, -s for scale (2, 3, or 4), -n for the model, -t for tile size, -g to select the GPU, and -f for output format.

Terminal
# Exécutable portable : photo agrandie ×4 avec le modèle généraliste
./realesrgan-ncnn-vulkan.exe -i input.jpg -o output.png -n realesrgan-x4plus

# Script Python : échelle arbitraire (ici ×3,5) avec amélioration des visages
python inference_realesrgan.py -n RealESRGAN_x4plus -i inputs --outscale 3.5 --face_enhance

The Python script provides the --outscale option: the model always performs a ×4 upscale, then the program resizes the result to the desired scale using an inexpensive resize operation. There is also --tile to split the image and save memory, and --fp32 for full precision, which becomes mandatory on a processor (see below).

In ComfyUI, upscaling models go in ComfyUI/models/upscale_models, and a loader node followed by one upscaling node per model applies them. For a desktop interface, the README lists several projects that bundle Real-ESRGAN, including Upscayl, free software for Linux, macOS, and Windows, with around 50,000 stars and an update in the past few days.

#Use it effectively

  1. 01
    Choose the model based on the image type
    Photo or illustration: this is the setting that changes the result the most, well before the upscaling factor. Always add -n to the command.
  2. 02
    Don’t stack passes
    Upscaling twice in a row by ×4 produces plastic, not detail. One pass at the right factor is enough; if ×4 is too much, use the ×2 model or the scaling option.
  3. 03
    Clean up before scaling up
    Light denoising beforehand prevents the model from treating noise as detail to reconstruct. With realesr-general-x4v3, the -dn option controls denoising strength to prevent an overly smooth result.
  4. 04
    Reserve face_enhance for real faces
    The face_enhance option relies on GFPGAN and is intended only for real faces: the project's FAQ advises against using it on animation, where it needlessly consumes GPU memory.
  5. 05
    Compare at 100%
    Judge an upscale at the pixel level, not from a thumbnail. Many spectacular results at small sizes look smooth and artificial up close.

#What size should you target: calculate before launching

A ×4 enlargement is useful only if the final size serves a specific purpose. Here is the necessary factor for each destination, based on a simple calculation: pixels = centimeters ÷ 2.54 × resolution in dots per inch, then factor = target long side ÷ starting long side.

Required factor from an image whose longer side is 1,024 pixels (calculation, printing at 300 dpi)
UsageLarge target sideRequired factor
4K screen, full screen3,840 px×3.75: one ×4 pass is sufficient
A4 printing at 300 dpi3,508 px×3,4: a ×4 pass, or --outscale 3,4
A3 printing at 300 dpi4,961 px×4.8: exceeds a ×4 pass; accept a slightly smaller draw rather than stacking two passes
Web page or thumbnail1,600 to 2,000 px×1.6 to ×2: the ×2 model or --outscale is enough

Keep the principle in mind: target the useful size, not the maximum size. A 1,024 × 1,024 image enlarged ×4 becomes 4,096 × 4,096, or about 16.8 million pixels—sixteen times more than the original: that is sixteen times as much data to store, load, and edit, for a gain that is often invisible on screen. With the Python script’s --outscale option, you choose the exact scale instead of over-enlarging and then reducing.

#In an image-generation pipeline

For local generation, upscaling is the final step: generate at a resolution the card can support, then upscale, which uses less memory than generating directly at a large size. The native sizes of image models—for example, 512 × 512 for SD 1.5 or 1024 × 1024 for SDXL—are detailed in the Stable Diffusion guide. Then, ×2 upscaling on a 1024 × 1024 SDXL image produces 2048 × 2048 without running it through the diffusion model again.

#What it makes up

The text
Enlarged small print turns into plausible but false characters. For a document, you need optical character recognition, not a magnifier.
Faces
An enlarged blurry face becomes a sharp face that is not exactly the same person. No identification use is acceptable.
The fine details
Fabric patterns, foliage, hair: the model produces a credible texture, not the original texture.
The drawbacks
A clearly visible compression artifact may sometimes be enlarged with the application instead of being erased.
!
An enlarged image is not proof
In a legal, medical, or identification context, the detail added by a model does not exist in the original. Upscaling improves readability; it adds no information that can be used as evidence.

#Troubleshooting: common errors

Known issues, according to the project FAQ and bug reports
SymptomCausePatch
Error “slow_conv2d_cpu not implemented for Half”Real-ESRGAN uses half precision (fp16) by default, which some operators do not support on the CPUAdd the --fp32 option to the command
ModuleNotFoundError: torchvision.transforms.functional_tensorbasicsr imports a module that torchvision removed starting with version 0.17 (August 28, 2024 report; a fix is proposed in a merge request in the BasicSR repository)Install a compatible version of torchvision, or use the portable ncnn executable, which does not depend on it
Visible blocks or seams in the imageThe executable splits the image into tiles and then reassembles them, which the README flags as a source of inconsistenciesSet the tile size with -t (0 for automatic): larger tiles reduce seams when memory allows
Black image on outputThe ncnn repository TODO notes that some PCs produce black imagesTry the other route (Python or ncnn) or another card with -g
Distorted faces in an illustrationface_enhance option enabled on non-photographic contentRemove it: it is designed for real faces

#Where the project stands in 2026

You have to be honest about its age. The latest released version, v0.3.0, dates from September 20, 2022, and the last commit to the main repository, from April 2, 2024, removes a continuous integration configuration file. The ncnn implementation, whose README states an MIT license, has not been updated since May 2024. That does not make the model bad: it does its job, it is small, and it runs on modest cards. But you cannot count on fixes for recent incompatibilities, such as the one with torchvision.

Among the alternatives, ComfyUI's README lists SeedVR2 and SUPIR among its supported models. SeedVR2 presents itself as one-step video restoration through adversarial diffusion post-training, while SUPIR targets photorealistic restoration of “in the wild” images. These are diffusion-based approaches: we do not have a reliable numerical comparison with Real-ESRGAN, and the risk of invented detail is at least as high. Read SUPIR's license, which GitHub classifies as “Other,” before commercial use. For most everyday photos and illustrations, Real-ESRGAN remains the best balance of simplicity, speed, and hardware requirements.

#FAQ

FAQ
Is Real-ESRGAN free?+
Yes. The repository is licensed under BSD-3-Clause, the ncnn-vulkan implementation is licensed under MIT, and execution is local, with no usage fees or account. However, check the rights to the images you upscale, as well as the license of each third-party model or desktop application that bundles Real-ESRGAN, such as Upscayl, which is licensed under AGPL-3.0.
Do you need a graphics card?+
Not necessarily. The Python script runs on the CPU with the --fp32 option, but much more slowly. The ncnn executable requires a Vulkan-compatible GPU—Intel, AMD, or NVIDIA—and requires neither CUDA nor PyTorch. The small realesr-general-x4v3 model is designed to use much less GPU memory and time, making it suitable for modest machines.
Can scanned text be enlarged?+
Technically yes, and that is not what you should do: the reconstructed characters are plausible and sometimes wrong. For a scanned document, use optical character recognition, which reads the text instead of redrawing it. The magnifier can be used before OCR for visual comfort only, never to validate the text that was read.
Why is my illustration grainy after enlargement?+
You probably used the model intended for photos. For drawings, use RealESRGAN_x4plus_anime_6B, optimized for anime images with a smaller network. With the portable executable, also watch out for the default model: specify it with -n, because the help text shows realesr-animevideov3 as the default.
Can you use it with video?+
Yes, image by image. For animation, the project provides the realesr-animevideov3 model and a video script. With the ncnn executable, the documented method is to extract the frames with ffmpeg, process them, then reassemble them. Processed independently, successive frames may flicker: check the result during playback, not only when paused.
Is Real-ESRGAN still maintained?+
Very little. Its latest version, v0.3.0, dates from September 2022, and its latest commit from April 2024 only affects continuous integration. The model remains usable, but dependency incompatibilities such as torchvision are not fixed by the project. To avoid these issues, the ncnn portable executable is the most stable option.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.