Text Generation WebUI (oobabooga)
oobabooga is the pseudonym of the developer of Text Generation WebUI, now renamed TextGen: an open-source interface (AGPL-3.0, approximately 47,000 GitHub stars) for running language models locally, without telemetry. It can be downloaded as a portable version, like a desktop application, for GGUF models; the full installation adds other engines, LoRA training, and extensions. Official source: the oobabooga/textgen GitHub repository.
Many Text Generation WebUI tutorials describe a version from two years ago: engines that have disappeared, replaced extensions, and changed commands. This guide revisits the subject based on the current repository: the new name, the three installation methods, the engines actually supported, the extensions, the API, and the security settings to know before exposing the interface to others.
#oobabooga, Text Generation WebUI, TextGen: what are we talking about?
oobabooga is the GitHub pseudonym of the project’s developer. The tool, long known as Text Generation WebUI, is now presented under the name TextGen: the old repository address redirects to github.com/oobabooga/textgen. It is an open-source application under the AGPL-3.0 license that brings together chat, free-form generation, vision, tool calling, the OpenAI-compatible API, and LoRA training. The repository describes it as completely offline and private, with no telemetry, external resources, or remote update requests.
The project is active: the latest version we recorded, 4.9, dates from May 2026. To our knowledge, there is no official site other than this GitHub repository: beware of third-party download sites and forks claiming to replace it, because a tampered executable runs with your permissions.
#1. Installation: portable, one-click, or manual
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The repository offers three options. The portable version is the simplest: it includes all dependencies, starts with a double-click, and loads only GGUF models (llama.cpp engine). It is available for Linux, Windows, and macOS, with CUDA, Vulkan, ROCm, and CPU-only variants. The full installation, required for other engines, training, image generation, and extensions, downloads PyTorch and requires about 10 GB of disk space.
| Mode | What you get | What you need to know |
|---|---|---|
| Portable version (desktop application) | GGUF via llama.cpp, all-inclusive | The simplest option; no extensions or other engines |
| One-click installer (start_ scripts) | Complete installation: ExLlamaV3, Transformers, training, extensions | Approximately 10 GB; GPU manufacturer choice during installation |
| Manual with venv or Conda | Full control over the Python environment | For developers; reserved for special cases |
- 01Download the portable versionFrom the repository's releases page, choose the archive for your system and GPU. For NVIDIA, use the cuda13.1 build if nvidia-smi reports a CUDA version greater than or equal to 13.1; otherwise, use cuda12.4. AMD and Intel use Vulkan; AMD can also use ROCm, and a CPU-only version is available.
- 02Extract and runExtract the archive, then double-click textgen: a window opens.
- 03Remove quarantine on macOSRun xattr -cr followed by the path to the extracted folder before the first launch, as indicated in the release notes.
- 04Add a modelPlace a GGUF file in the user_data/models folder; the interface detects it automatically.
- 05Load and chatSelect the model in the models tab, then open the chat.
Updating a portable build doesn’t mean losing your models or settings: download the new archive, extract it, and replace its user_data folder with the one from your old installation. Since version 4.0, you can also place user_data one level above the installation folders; it will then be detected automatically.
For the complete installation, clone the repository and run your system’s script. The script asks for your GPU manufacturer, installs the dependencies in a local folder, and then you open http://127.0.0.1:7860 in your browser.
Everything is stored in a local installer_files folder. To reinstall from scratch, delete this folder and rerun the script. To apply options permanently, write them to user_data/CMD_FLAGS.txt, for example --api to enable the API. None of these scripts require administrator privileges.
#2. Download and choose a model
The repository's instructions fit in three sentences: download a GGUF file from Hugging Face, place it in user_data/models, and the interface will detect it. Multi-file models (16-bit Transformers, EXL3) go in a subdirectory of the same folder and require the full installation, not the portable version.
The only practical criterion is memory. The site's Q4 quantization reference points, weights only: approximately 2 GB for 3B, approximately 5 GB for 7B to 8B, approximately 9 GB for 14B, approximately 19 to 20 GB for 32B, and approximately 40 GB for 70B. Add the context cache and some headroom for the system. The site's VRAM calculator gives you the exact total for your context.
#Context and cache: the two options that save memory
Two launch options directly affect memory. --ctx-size sets the context size in tokens: a value of 0 means automatic with llama.cpp, provided all layers are placed on the GPU (--gpu-layers=-1), while the default is 8,192 for other engines. --cache-type selects the context-cache format; with llama.cpp, valid values are fp16, q8_0, and q4_0. A quantized cache uses less memory than an fp16 cache, at the cost of slightly lower precision: try q8_0 first if memory is tight before reducing the model size.
The repository also points to two community tools to help frame this choice: a memory calculator for GGUF models and a list of recommended quantizations. Check their results with the site's calculator, then monitor the card's actual utilization during a long conversation.
#3. Choose an engine (loader)
TextGen loads a model with an engine called a loader. By default, it detects the loader; you can force one with the --loader option. The accepted values today are Transformers, llama.cpp, ExLlamav3_HF, ExLlamav3, and TensorRT-LLM, and you can switch engines and models without restarting.
| Engine | Model format | Usage |
|---|---|---|
| llama.cpp | GGUF | The default choice: CPU and GPU, the only engine in the portable version |
| ik_llama.cpp | GGUF | A variant of llama.cpp, listed as an additional backend |
| ExLlamaV3 | EXL3 | GPU quantization, in the complete installation |
| Transformers | 16-bit models, safetensors | Versatile and memory-hungry; also used for training |
| TensorRT-LLM | Models compiled for NVIDIA | For NVIDIA GPUs, full installation |
If you’re just getting started, stick with llama.cpp and the GGUF format: it’s the shortest path, the portable version, and the one that other tools (Ollama, LM Studio) use as well. Other engines are justified by a specific need: a particular format, fine-tuning, or making maximum use of a NVIDIA card.
#4. Chat, instruct, notebook, vision
- Instruct
- Instruction-following mode, such as ChatGPT. Prompts are automatically formatted using Jinja2 templates.
- Chat-instruct and chat
- For conversing with custom characters.
- Notebook
- A free-generation tab outside the chat turns: you write, and the model continues.
- Vision and files
- You can attach images to messages, along with text files, PDFs, and .docx files, to discuss their contents.
- Edition and branches
- Edit a message, navigate between its versions, and open a branch at any point in the conversation.
Models can also call tools during the conversation: web search, page retrieval, and calculations. Each tool is a simple Python file, and MCP servers are supported. This feature interacts with context: an agent that calls tools quickly fills a short context window, so consider increasing the context size with the --ctx-size option.
#5. Extensions: what really exists
The extensions work only with the full installation. The repository's extensions folder includes coqui_tts, silero_tts (speech synthesis), whisper_stt (voice input), sd_api_pictures and send_pictures (images), google_translate, superbooga and superboogav2, ngrok, gallery, and perplexity_colors. The README adds that community extensions also exist.
#6. API mode
Add --api to your options to launch the API. According to the repository wiki, the default port is 5000, which can be changed with --api-port, and the API remains local unless you explicitly ask to expose it. The wiki specifies that it is offline, does not connect to OpenAI, and creates no logs. The endpoints cover chats, completions, and Anthropic-format messages, with tool calling.
You can connect any OpenAI client—a chat frontend, a code agent, or LangChain—by pointing it at this address. Interactive documentation for all endpoints is served at the same address, under /docs.
#Open the interface to others: options to know about
By default, the interface is reachable only from your machine. Three options change this. --listen makes the interface accessible from your local network. --share creates a public address, which should be avoided except for a one-off test. And --multi-user does not save chat histories: the repository describes it as suitable for small teams you trust, not as an enterprise solution.
For the API, add --api-key with a key of your choice as soon as you listen on the network: without a key, any device that reaches the port can query your model. For real team use (accounts, roles, logging), put a dedicated frontend in front of the engine instead.
#TextGen compared with Ollama and LM Studio
| Need | TextGen | Ollama or LM Studio |
|---|---|---|
| Get started quickly with a GGUF | Portable version, simple too | Ollama: one command; LM Studio: guided interface |
| Test multiple engines and formats | Yes: that's its strength | Focused on GGUF |
| Fine-tune a LoRA in the interface | Yes, with the full installation | No |
| Serve a team | Poor fit (limited multi-user support) | Ollama behind Open WebUI or LibreChat |
| Smooth everyday use | More settings, more steps | Easier to maintain |
In summary: TextGen is the tool for exploration and fine-tuning, not for a beginner's daily use. If you want to talk to a model first, start with Ollama or LM Studio; return to TextGen when you need an engine, format, or training capability that the others do not offer.
- Ollama vs. LM Studio vs. Jan vs. GPT4All
- Source: TextGen repository (oobabooga)
- Source: version 4.9 notes
- Source: wiki, OpenAI-compatible API
oobabooga and Text Generation WebUI: are they the same thing?+
How do you install Text Generation WebUI in 2026?+
Which engines and formats are supported?+
Is it really offline and telemetry-free?+
How do you enable the API to connect another tool?+
Should I prefer TextGen to Ollama?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.