Intermediate 11 minOther tools

Text Generation WebUI (oobabooga)

Direct response

oobabooga is the pseudonym of the developer of Text Generation WebUI, now renamed TextGen: an open-source interface (AGPL-3.0, approximately 47,000 GitHub stars) for running language models locally, without telemetry. It can be downloaded as a portable version, like a desktop application, for GGUF models; the full installation adds other engines, LoRA training, and extensions. Official source: the oobabooga/textgen GitHub repository.

Many Text Generation WebUI tutorials describe a version from two years ago: engines that have disappeared, replaced extensions, and changed commands. This guide revisits the subject based on the current repository: the new name, the three installation methods, the engines actually supported, the extensions, the API, and the security settings to know before exposing the interface to others.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#oobabooga, Text Generation WebUI, TextGen: what are we talking about?

oobabooga is the GitHub pseudonym of the project’s developer. The tool, long known as Text Generation WebUI, is now presented under the name TextGen: the old repository address redirects to github.com/oobabooga/textgen. It is an open-source application under the AGPL-3.0 license that brings together chat, free-form generation, vision, tool calling, the OpenAI-compatible API, and LoRA training. The repository describes it as completely offline and private, with no telemetry, external resources, or remote update requests.

The project is active: the latest version we recorded, 4.9, dates from May 2026. To our knowledge, there is no official site other than this GitHub repository: beware of third-party download sites and forks claiming to replace it, because a tampered executable runs with your permissions.

i
Why old tutorials lead you astray
The current README lists five engines: llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3, and TensorRT-LLM. ExLlamaV2, the AWQ and GPTQ formats, and vLLM, often mentioned in older guides (including the first version of this one), are no longer listed. Always check a tutorial's date.

#1. Installation: portable, one-click, or manual

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The repository offers three options. The portable version is the simplest: it includes all dependencies, starts with a double-click, and loads only GGUF models (llama.cpp engine). It is available for Linux, Windows, and macOS, with CUDA, Vulkan, ROCm, and CPU-only variants. The full installation, required for other engines, training, image generation, and extensions, downloads PyTorch and requires about 10 GB of disk space.

Choose your installation mode
ModeWhat you getWhat you need to know
Portable version (desktop application)GGUF via llama.cpp, all-inclusiveThe simplest option; no extensions or other engines
One-click installer (start_ scripts)Complete installation: ExLlamaV3, Transformers, training, extensionsApproximately 10 GB; GPU manufacturer choice during installation
Manual with venv or CondaFull control over the Python environmentFor developers; reserved for special cases
  1. 01
    Download the portable version
    From the repository's releases page, choose the archive for your system and GPU. For NVIDIA, use the cuda13.1 build if nvidia-smi reports a CUDA version greater than or equal to 13.1; otherwise, use cuda12.4. AMD and Intel use Vulkan; AMD can also use ROCm, and a CPU-only version is available.
  2. 02
    Extract and run
    Extract the archive, then double-click textgen: a window opens.
  3. 03
    Remove quarantine on macOS
    Run xattr -cr followed by the path to the extracted folder before the first launch, as indicated in the release notes.
  4. 04
    Add a model
    Place a GGUF file in the user_data/models folder; the interface detects it automatically.
  5. 05
    Load and chat
    Select the model in the models tab, then open the chat.

Updating a portable build doesn’t mean losing your models or settings: download the new archive, extract it, and replace its user_data folder with the one from your old installation. Since version 4.0, you can also place user_data one level above the installation folders; it will then be detected automatically.

For the complete installation, clone the repository and run your system’s script. The script asks for your GPU manufacturer, installs the dependencies in a local folder, and then you open http://127.0.0.1:7860 in your browser.

Complete installation with the one-click installer (Linux; macOS and Windows: start_macos.sh, start_windows.bat)
git clone https://github.com/oobabooga/textgen
cd textgen
./start_linux.sh

Everything is stored in a local installer_files folder. To reinstall from scratch, delete this folder and rerun the script. To apply options permanently, write them to user_data/CMD_FLAGS.txt, for example --api to enable the API. None of these scripts require administrator privileges.

#2. Download and choose a model

The repository's instructions fit in three sentences: download a GGUF file from Hugging Face, place it in user_data/models, and the interface will detect it. Multi-file models (16-bit Transformers, EXL3) go in a subdirectory of the same folder and require the full installation, not the portable version.

The only practical criterion is memory. The site's Q4 quantization reference points, weights only: approximately 2 GB for 3B, approximately 5 GB for 7B to 8B, approximately 9 GB for 14B, approximately 19 to 20 GB for 32B, and approximately 40 GB for 70B. Add the context cache and some headroom for the system. The site's VRAM calculator gives you the exact total for your context.

#Context and cache: the two options that save memory

Two launch options directly affect memory. --ctx-size sets the context size in tokens: a value of 0 means automatic with llama.cpp, provided all layers are placed on the GPU (--gpu-layers=-1), while the default is 8,192 for other engines. --cache-type selects the context-cache format; with llama.cpp, valid values are fp16, q8_0, and q4_0. A quantized cache uses less memory than an fp16 cache, at the cost of slightly lower precision: try q8_0 first if memory is tight before reducing the model size.

The repository also points to two community tools to help frame this choice: a memory calculator for GGUF models and a list of recommended quantizations. Check their results with the site's calculator, then monitor the card's actual utilization during a long conversation.

#3. Choose an engine (loader)

TextGen loads a model with an engine called a loader. By default, it detects the loader; you can force one with the --loader option. The accepted values today are Transformers, llama.cpp, ExLlamav3_HF, ExLlamav3, and TensorRT-LLM, and you can switch engines and models without restarting.

The engines in the current README
EngineModel formatUsage
llama.cppGGUFThe default choice: CPU and GPU, the only engine in the portable version
ik_llama.cppGGUFA variant of llama.cpp, listed as an additional backend
ExLlamaV3EXL3GPU quantization, in the complete installation
Transformers16-bit models, safetensorsVersatile and memory-hungry; also used for training
TensorRT-LLMModels compiled for NVIDIAFor NVIDIA GPUs, full installation

If you’re just getting started, stick with llama.cpp and the GGUF format: it’s the shortest path, the portable version, and the one that other tools (Ollama, LM Studio) use as well. Other engines are justified by a specific need: a particular format, fine-tuning, or making maximum use of a NVIDIA card.

#4. Chat, instruct, notebook, vision

Instruct
Instruction-following mode, such as ChatGPT. Prompts are automatically formatted using Jinja2 templates.
Chat-instruct and chat
For conversing with custom characters.
Notebook
A free-generation tab outside the chat turns: you write, and the model continues.
Vision and files
You can attach images to messages, along with text files, PDFs, and .docx files, to discuss their contents.
Edition and branches
Edit a message, navigate between its versions, and open a branch at any point in the conversation.

Models can also call tools during the conversation: web search, page retrieval, and calculations. Each tool is a simple Python file, and MCP servers are supported. This feature interacts with context: an agent that calls tools quickly fills a short context window, so consider increasing the context size with the --ctx-size option.

#5. Extensions: what really exists

The extensions work only with the full installation. The repository's extensions folder includes coqui_tts, silero_tts (speech synthesis), whisper_stt (voice input), sd_api_pictures and send_pictures (images), google_translate, superbooga and superboogav2, ngrok, gallery, and perplexity_colors. The README adds that community extensions also exist.

→
No need for an “openai” extension anymore
OpenAI- and Anthropic-compatible APIs are now built in: just add --api. Tutorials that tell you to enable an openai extension describe the old behavior.

#6. API mode

Add --api to your options to launch the API. According to the repository wiki, the default port is 5000, which can be changed with --api-port, and the API remains local unless you explicitly ask to expose it. The wiki specifies that it is offline, does not connect to OpenAI, and creates no logs. The endpoints cover chats, completions, and Anthropic-format messages, with tool calling.

Request to the local API (example from the official wiki)
curl http://127.0.0.1:5000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Hello!"}],
    "temperature": 0.6
  }'

You can connect any OpenAI client—a chat frontend, a code agent, or LangChain—by pointing it at this address. Interactive documentation for all endpoints is served at the same address, under /docs.

#Open the interface to others: options to know about

By default, the interface is reachable only from your machine. Three options change this. --listen makes the interface accessible from your local network. --share creates a public address, which should be avoided except for a one-off test. And --multi-user does not save chat histories: the repository describes it as suitable for small teams you trust, not as an enterprise solution.

For the API, add --api-key with a key of your choice as soon as you listen on the network: without a key, any device that reaches the port can query your model. For real team use (accounts, roles, logging), put a dedicated frontend in front of the engine instead.

#TextGen compared with Ollama and LM Studio

Which tool for which need
NeedTextGenOllama or LM Studio
Get started quickly with a GGUFPortable version, simple tooOllama: one command; LM Studio: guided interface
Test multiple engines and formatsYes: that's its strengthFocused on GGUF
Fine-tune a LoRA in the interfaceYes, with the full installationNo
Serve a teamPoor fit (limited multi-user support)Ollama behind Open WebUI or LibreChat
Smooth everyday useMore settings, more stepsEasier to maintain

In summary: TextGen is the tool for exploration and fine-tuning, not for a beginner's daily use. If you want to talk to a model first, start with Ollama or LM Studio; return to TextGen when you need an engine, format, or training capability that the others do not offer.

Frequently asked questions
oobabooga and Text Generation WebUI: are they the same thing?+
Yes. oobabooga is the developer's pseudonym, and Text Generation WebUI is the project's historical name, now presented as TextGen. The old GitHub address redirects to github.com/oobabooga/textgen. It is the only official source we know of: avoid third-party download sites and verify the address before running anything.
How do you install Text Generation WebUI in 2026?+
The simplest option is the portable version, downloadable from the repository’s releases: it’s ready to use but limited to GGUF models. For other engines, training, or extensions, clone the repository and run start_windows.bat, start_linux.sh, or start_macos.sh, allowing for about 10 GB of disk space.
Which engines and formats are supported?+
The current README lists llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3, and TensorRT-LLM. The GGUF format runs through llama.cpp, the only engine in the portable version. ExLlamaV2, AWQ, GPTQ, and vLLM, mentioned in older tutorials, are no longer included: check the documentation before following a guide that mentions them.
Is it really offline and telemetry-free?+
The repository claims to operate 100% offline and privately, without telemetry, external resources, or update requests. The wiki specifies that the API does not connect to OpenAI and does not create a log. Web search features, however, contact the Internet by nature: disable them for strictly local use.
How do you enable the API to connect another tool?+
Add --api to the command line or to user_data/CMD_FLAGS.txt. The default port is 5000, and can be changed with --api-port; the address is http://127.0.0.1:5000/v1. If you listen on the network, add --api-key to require a key; otherwise, any device that can reach the port can use it.
Should I prefer TextGen to Ollama?+
Not to start with: Ollama and LM Studio are simpler for everyday use, with fewer options to understand. TextGen makes sense when you want to test multiple engines and formats, train a LoRA in the interface, or fine-tune generation. A common approach is to keep Ollama for everyday use and TextGen for exploration.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.