Intermediate 12 minInterfaces

SillyTavern: the character interface for LLMs local

SillyTavern is a web interface designed for role-playing and interactive writing with an LLM. It does not run a model itself: it connects to a local backend such as KoboldCpp or Ollama and adds everything missing from standard chat—character cards, persistent world memory, fine-grained prompt management, and extensions. This guide covers installation, connecting to your backend, creating characters, and the settings that make a difference in role-play.

By Léa B.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why SillyTavern instead of a regular chat

An interface like Open WebUI or LM Studio treats each conversation as an assistant/user exchange. SillyTavern addresses a different need: embodying a consistent character across hundreds of messages in a world that remembers its own rules. It’s the go-to tool for the local LLM role-playing community, and it provides mechanics that raw chat doesn’t.

Character cards
A standardized format (PNG with embedded metadata) that describes persona, style, dialogue examples, and greeting message. Importable and shareable as a single file.
World knowledge
The lorebook injects information into the prompt only when a keyword appears, keeping the universe consistent without overwhelming the context.
Prompt control
You can see and modify every building block sent to the model: system prompt, instruction format, injection order. Nothing is hidden.
Backend of your choice
SillyTavern is a pure front end. The same character runs on KoboldCpp, Ollama, llama.cpp, or a remote API without rewriting anything.
i
SillyTavern doesn’t replace your inference server
It’s an interface layer. You still need a backend that loads the model and runs generation. SillyTavern simply orchestrates the prompt and displays the result.

#Prerequisites and backend choice

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

SillyTavern runs on Node.js (version 18 or later) under Windows, macOS, and Linux. It is lightweight on the resource side—the backend is what uses VRAM. Two local backends cover nearly all use cases: KoboldCpp, geared toward GGUF and role-playing, and Ollama, which is more general-purpose.

KoboldCpp
A single binary that runs GGUFs with an API SillyTavern supports natively. Fine-grained memory settings, AMD ROCm support. The RP community's default choice.
Ollama
The daemon is listening on http://localhost:11434. Useful if you already use it for something else; SillyTavern connects to it through its compatible endpoint.
llama.cpp (llama-server)
OpenAI-compatible HTTP server for maximum control over layer offloading.
VRAM
In Q4_K_M: a 7B model fits in ~5 GB, a 14B model in ~9 GB, and a 32B model in ~19 GB. For smooth roleplay, target at least a 12B–14B model if your card supports it.
→
Longer context = more VRAM
Role-playing relies on long contexts (8k, 16k, 32k tokens). The KV cache grows with the context and consumes VRAM in addition to the weights. Leave some headroom, or quantize the KV cache to fit a larger context on the same card.

#Install SillyTavern

The recommended method is a Git clone: it makes updates trivial (a git pull), and SillyTavern evolves quickly. Make sure Node.js 18+ and Git are installed.

  1. 01
    Clone the repository
    Fetch the release branch, which is stable and tested. Avoid the staging branch unless you want early access to new features at the cost of a few bugs.
  2. 02
    Run the startup script
    start.sh on Linux/macOS, Start.bat on Windows. On the first launch, npm installs the dependencies automatically.
  3. 03
    Open the interface
    SillyTavern serves its interface at http://localhost:8000. Open this address in your browser; no desktop application to install.
  4. 04
    Update later
    Run git pull in the directory, then rerun the script. Your characters and conversations are stored separately in data/ and aren’t overwritten.
Terminal (Linux/macOS)
# Cloner la branche stable
git clone https://github.com/SillyTavern/SillyTavern -b release
cd SillyTavern

# Premier lancement : installe les dépendances puis démarre
./start.sh

# Interface accessible sur http://localhost:8000
PowerShell (Windows)
git clone https://github.com/SillyTavern/SillyTavern -b release
cd SillyTavern

# Double-cliquez Start.bat, ou en ligne de commande :
.\Start.bat
i
Network access and security
By default, SillyTavern listens only on localhost. If you want to access it from another device on the network, enable listen in config.yaml and configure authentication (basic auth). Never expose the interface directly to the Internet without protection.

#Connect the local backend

Once SillyTavern is open, everything happens in the API connection tab (the plug-shaped icon at the top). The principle is simple: start the backend on one side and enter its address in SillyTavern on the other.

#With KoboldCpp

Start KoboldCpp with your GGUF model; it exposes its API on port 5001 by default. In SillyTavern, select the « Text Completion » API type, then the « KoboldCpp » backend, and enter the URL. Click Connect: the name of the loaded model appears if everything is working.

Terminal
# Lancer KoboldCpp avec un modèle et 8k de contexte
./koboldcpp --model mon-modele-rp.Q4_K_M.gguf --contextsize 8192

# API disponible sur http://localhost:5001
# Dans SillyTavern : API = Text Completion > KoboldCpp
# URL = http://localhost:5001

#With Ollama

Ollama is already running as a daemon on port 11434. In SillyTavern, select “Text Completion,” then the “Ollama” backend, enter the address, and choose the model from the automatically retrieved drop-down list.

Terminal
# Vérifier qu'Ollama tourne et lister les modèles
ollama list

# Le daemon écoute sur http://localhost:11434
# Dans SillyTavern : API = Text Completion > Ollama
# URL = http://localhost:11434
!
The right instruction template is decisive
Each model family expects a specific format (ChatML for Qwen 3.5/3.8 and gpt-oss, Gemma for Gemma 4, Mistral for Mistral Small…). If the responses are incoherent, cut off, or full of visible tags, it's almost always because the wrong template was selected. Set it in the Advanced Formatting tab so it matches the loaded model.

#Create and import character cards

The character card is the heart of SillyTavern. It is a PNG file whose image serves as the avatar and whose embedded metadata describes the character. You can create one from scratch or import a card shared by the community.

#The fields that matter

Description
The character’s foundation: appearance, traits, and history. This is injected into the context continuously, so keep it dense and concrete rather than verbose.
Personality
A summary of the temperament. Useful for guiding the tone without rewriting the entire description.
First message
The welcome message that sets the scene. It sets the expected tone, writing style, and format—the model tends to imitate its form.
Example dialogues
Examples of lines that show the model how the character speaks. Very effective for locking in a particular voice.
Scenario
The scene context, separate from the character description. Useful for reusing the same character in multiple situations.
  1. 01
    Open the characters panel
    The character icon in the top bar opens the list. The “+” button creates a blank card.
  2. 02
    Fill in the description
    Describe the character densely. Many authors use a structured format (trait lists) rather than a paragraph, which models follow well.
  3. 03
    Write the first message
    Polish it: it’s your strongest style lever. A narrative welcome message in the third person pushes the model toward that format.
  4. 04
    Import an existing map
    Drag a card PNG into the list, or import it. Cards are shared as simple image files on community hubs.
→
The character’s token budget
Everything the card contains takes up space in the context with every message. A 2,000-token description in an 8k context consumes a quarter of the budget before the conversation history is even included. Be concise: quality comes before length.

#The persistent world: the lorebook

The lorebook (or World Info) solves the central problem of long sessions: how to keep a coherent world without constantly injecting everything. The principle is conditional injection based on keywords.

You create entries, each associated with one or more keywords. When one of those words appears in recent messages, the corresponding entry is injected into the context just before generation. The rest of the time, it takes up no space. This lets you describe dozens of locations, minor characters, and rules without ever saturating the prompt.

Entry
A block of text (the lore to inject) and its trigger keywords. Example: keyword “Valmont” → description of the city of Valmont.
Constant vs. selective
An entry can always be active (fundamental rules of the universe) or triggered only by keyword (contextual details).
Injection depth
You choose where the input is inserted in the prompt, which affects the weight the model gives it.
Character-specific or global
A lorebook can accompany a specific card or apply to all your conversations as a shared universe.
i
Lorebook ≠ RAG
Keyword injection is simpler and more predictable than vector search: it triggers on a text match, not semantic similarity. For a structured fictional universe, it's often more reliable than traditional RAG, and it remains entirely under your control.

#Tune generation for interpretation

Sampling parameters determine the balance between consistency and creativity. For role-playing, you want more variety than for factual assistance, without becoming incoherent. These settings are in the generation settings tab (the sliders icon).

Temperature
The creativity/stability slider. Around 0.7–0.9 provides a good balance for roleplay. Too high (> 1.2) becomes incoherent; too low makes the character repetitive and flat.
Min-P
A modern sampler that cuts unlikely tokens in proportion to the most likely ones. A value of 0.05-0.1 cleans up the output and lets you raise the temperature without going off the rails.
Repetition penalty
Penalizes repeating the same tokens. Useful against loops, but if set too high, it forces the model into artificial phrasing. Keep it moderate.
Response length
The maximum number of tokens generated per response. 200-400 tokens for consistent narrative replies without endless monologues.
Context size
It must match what your backend has loaded. There's no point requesting 16k from SillyTavern if KoboldCpp was started with only 8k.
→
Start with a preset before tweaking
SillyTavern provides ready-to-use sampling presets. Load one suited to RP, play out a scene, then adjust one parameter at a time. Changing five sliders at once makes diagnosis impossible.

#Extensions that change the experience

SillyTavern is extensible. Some extensions are included out of the box; others are installed through their repository URL. These are the ones that truly transform the experience.

Vector Storage (memory summary)
Vectorizes the history and reinjects relevant passages from old messages, extending memory beyond the context window. Useful for very long conversations.
Summarize
Generates and maintains an ongoing summary of the conversation, reinjected into the prompt. The character remembers major events even after hundreds of messages.
Text-to-Speech
Give characters a voice through a local TTS engine. Greater immersion for those who play using voice.
Image generation
Connects to a local image backend such as Stable Diffusion to illustrate scenes and characters on the fly from the chat.
Expressions
Displays the avatar with an emotion matching the message's tone, using a sprite set for the character.
i
Each extension costs tokens or VRAM
Summarize and Vector Storage consume context; image generation and TTS require their own models and therefore additional VRAM or CPU beyond your LLM. Enable them based on your hardware budget, not all at once.

#Which local models excel at interpretation

Not all LLMs are equal for role-playing. “Instruct” models aligned for assistance tend to break character, moralize, or refuse fictional scenarios. The community favors specialized fine-tunes, often built on solid base models and then retrained for storytelling and in-character dialogue.

Usable size
Below 4B, character consistency over time suffers. In 2026, a recent base model such as Qwen 3.5 9B (~6.6 GB, 256k context) or Gemma 4 12B (~7.6 GB, Apache 2.0) is the sweet spot for most midrange graphics cards; one step up, Mistral Small 24B (~14 GB, very comfortable in French) fits on 16 GB and delivers more narrative nuance.
RP fine-tunes
Look for models explicitly trained for roleplay or fiction (often labeled “RP,” “storytelling,” or “uncensored” on Hugging Face). The best fine-tunes of 2026 are built on these recent bases (Qwen 3.5, Mistral Small 24B, Gemma 4) rather than the older Llama 2 or Mistral 7B models from 2023; they follow card formats better.
Long context
Favor a model that performs well at 16k+ tokens without degrading: long sections require it. Check the model’s stated native context length.
GGUF format
For KoboldCpp, stick with quantized GGUFs. Q4_K_M is the recommended compromise; move up to Q5_K_M or Q8_0 if your VRAM allows it and you want more precision.
→
Test a card with two or three models
The same character can seem dull on one model and brilliant on another. Before concluding that your card is bad, try it with a few different RP models: the difference in output is often spectacular.

#Troubleshooting common issues

SillyTavern won't connect
Verify that the backend is running properly and that the port is correct (5001 for KoboldCpp, 11434 for Ollama). A browser test at the API URL confirms that it responds.
Responses with visible tags
Almost always an unsuitable instruction template. Match the format (ChatML, Gemma, Mistral…) to the model's format in Advanced Formatting.
The character breaks character
Strengthen the description and dialogue examples, lower the temperature slightly, and make sure your system prompt does not conflict with the card.
Repetitions and loops
Slightly increase the repetition penalty and add a Min-P. If it persists, the context may be saturated: enable Summarize or reduce the map size.
Truncated responses
Response length is too low, or the backend was launched with a smaller context than SillyTavern requires. Align the two values.

#Go further

SillyTavern is just a layer: final quality depends mostly on your backend and model. These guides help you build a solid foundation.

KoboldCpp: installation and getting started
The reference backend for local roleplay: GGUF, memory settings, and an API that SillyTavern supports natively.
What is Ollama, and how does it work?
The general-purpose alternative if you prefer a single daemon on port 11434 for all your use cases.
Quantize the KV cache
To support longer RP contexts on the same graphics card without blowing through VRAM.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.