SillyTavern: the character interface for LLMs local
SillyTavern is a web interface designed for role-playing and interactive writing with an LLM. It does not run a model itself: it connects to a local backend such as KoboldCpp or Ollama and adds everything missing from standard chat—character cards, persistent world memory, fine-grained prompt management, and extensions. This guide covers installation, connecting to your backend, creating characters, and the settings that make a difference in role-play.
#Why SillyTavern instead of a regular chat
An interface like Open WebUI or LM Studio treats each conversation as an assistant/user exchange. SillyTavern addresses a different need: embodying a consistent character across hundreds of messages in a world that remembers its own rules. It’s the go-to tool for the local LLM role-playing community, and it provides mechanics that raw chat doesn’t.
- Character cards
- A standardized format (PNG with embedded metadata) that describes persona, style, dialogue examples, and greeting message. Importable and shareable as a single file.
- World knowledge
- The lorebook injects information into the prompt only when a keyword appears, keeping the universe consistent without overwhelming the context.
- Prompt control
- You can see and modify every building block sent to the model: system prompt, instruction format, injection order. Nothing is hidden.
- Backend of your choice
- SillyTavern is a pure front end. The same character runs on KoboldCpp, Ollama, llama.cpp, or a remote API without rewriting anything.
#Prerequisites and backend choice
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
SillyTavern runs on Node.js (version 18 or later) under Windows, macOS, and Linux. It is lightweight on the resource side—the backend is what uses VRAM. Two local backends cover nearly all use cases: KoboldCpp, geared toward GGUF and role-playing, and Ollama, which is more general-purpose.
- KoboldCpp
- A single binary that runs GGUFs with an API SillyTavern supports natively. Fine-grained memory settings, AMD ROCm support. The RP community's default choice.
- Ollama
- The daemon is listening on http://localhost:11434. Useful if you already use it for something else; SillyTavern connects to it through its compatible endpoint.
- llama.cpp (llama-server)
- OpenAI-compatible HTTP server for maximum control over layer offloading.
- VRAM
- In Q4_K_M: a 7B model fits in ~5 GB, a 14B model in ~9 GB, and a 32B model in ~19 GB. For smooth roleplay, target at least a 12B–14B model if your card supports it.
#Install SillyTavern
The recommended method is a Git clone: it makes updates trivial (a git pull), and SillyTavern evolves quickly. Make sure Node.js 18+ and Git are installed.
- 01Clone the repositoryFetch the release branch, which is stable and tested. Avoid the staging branch unless you want early access to new features at the cost of a few bugs.
- 02Run the startup scriptstart.sh on Linux/macOS, Start.bat on Windows. On the first launch, npm installs the dependencies automatically.
- 03Open the interfaceSillyTavern serves its interface at http://localhost:8000. Open this address in your browser; no desktop application to install.
- 04Update laterRun git pull in the directory, then rerun the script. Your characters and conversations are stored separately in data/ and aren’t overwritten.
#Connect the local backend
Once SillyTavern is open, everything happens in the API connection tab (the plug-shaped icon at the top). The principle is simple: start the backend on one side and enter its address in SillyTavern on the other.
#With KoboldCpp
Start KoboldCpp with your GGUF model; it exposes its API on port 5001 by default. In SillyTavern, select the « Text Completion » API type, then the « KoboldCpp » backend, and enter the URL. Click Connect: the name of the loaded model appears if everything is working.
#With Ollama
Ollama is already running as a daemon on port 11434. In SillyTavern, select “Text Completion,” then the “Ollama” backend, enter the address, and choose the model from the automatically retrieved drop-down list.
#Create and import character cards
The character card is the heart of SillyTavern. It is a PNG file whose image serves as the avatar and whose embedded metadata describes the character. You can create one from scratch or import a card shared by the community.
#The fields that matter
- Description
- The character’s foundation: appearance, traits, and history. This is injected into the context continuously, so keep it dense and concrete rather than verbose.
- Personality
- A summary of the temperament. Useful for guiding the tone without rewriting the entire description.
- First message
- The welcome message that sets the scene. It sets the expected tone, writing style, and format—the model tends to imitate its form.
- Example dialogues
- Examples of lines that show the model how the character speaks. Very effective for locking in a particular voice.
- Scenario
- The scene context, separate from the character description. Useful for reusing the same character in multiple situations.
- 01Open the characters panelThe character icon in the top bar opens the list. The “+” button creates a blank card.
- 02Fill in the descriptionDescribe the character densely. Many authors use a structured format (trait lists) rather than a paragraph, which models follow well.
- 03Write the first messagePolish it: it’s your strongest style lever. A narrative welcome message in the third person pushes the model toward that format.
- 04Import an existing mapDrag a card PNG into the list, or import it. Cards are shared as simple image files on community hubs.
#The persistent world: the lorebook
The lorebook (or World Info) solves the central problem of long sessions: how to keep a coherent world without constantly injecting everything. The principle is conditional injection based on keywords.
You create entries, each associated with one or more keywords. When one of those words appears in recent messages, the corresponding entry is injected into the context just before generation. The rest of the time, it takes up no space. This lets you describe dozens of locations, minor characters, and rules without ever saturating the prompt.
- Entry
- A block of text (the lore to inject) and its trigger keywords. Example: keyword “Valmont” → description of the city of Valmont.
- Constant vs. selective
- An entry can always be active (fundamental rules of the universe) or triggered only by keyword (contextual details).
- Injection depth
- You choose where the input is inserted in the prompt, which affects the weight the model gives it.
- Character-specific or global
- A lorebook can accompany a specific card or apply to all your conversations as a shared universe.
#Tune generation for interpretation
Sampling parameters determine the balance between consistency and creativity. For role-playing, you want more variety than for factual assistance, without becoming incoherent. These settings are in the generation settings tab (the sliders icon).
- Temperature
- The creativity/stability slider. Around 0.7–0.9 provides a good balance for roleplay. Too high (> 1.2) becomes incoherent; too low makes the character repetitive and flat.
- Min-P
- A modern sampler that cuts unlikely tokens in proportion to the most likely ones. A value of 0.05-0.1 cleans up the output and lets you raise the temperature without going off the rails.
- Repetition penalty
- Penalizes repeating the same tokens. Useful against loops, but if set too high, it forces the model into artificial phrasing. Keep it moderate.
- Response length
- The maximum number of tokens generated per response. 200-400 tokens for consistent narrative replies without endless monologues.
- Context size
- It must match what your backend has loaded. There's no point requesting 16k from SillyTavern if KoboldCpp was started with only 8k.
#Extensions that change the experience
SillyTavern is extensible. Some extensions are included out of the box; others are installed through their repository URL. These are the ones that truly transform the experience.
- Vector Storage (memory summary)
- Vectorizes the history and reinjects relevant passages from old messages, extending memory beyond the context window. Useful for very long conversations.
- Summarize
- Generates and maintains an ongoing summary of the conversation, reinjected into the prompt. The character remembers major events even after hundreds of messages.
- Text-to-Speech
- Give characters a voice through a local TTS engine. Greater immersion for those who play using voice.
- Image generation
- Connects to a local image backend such as Stable Diffusion to illustrate scenes and characters on the fly from the chat.
- Expressions
- Displays the avatar with an emotion matching the message's tone, using a sprite set for the character.
#Which local models excel at interpretation
Not all LLMs are equal for role-playing. “Instruct” models aligned for assistance tend to break character, moralize, or refuse fictional scenarios. The community favors specialized fine-tunes, often built on solid base models and then retrained for storytelling and in-character dialogue.
- Usable size
- Below 4B, character consistency over time suffers. In 2026, a recent base model such as Qwen 3.5 9B (~6.6 GB, 256k context) or Gemma 4 12B (~7.6 GB, Apache 2.0) is the sweet spot for most midrange graphics cards; one step up, Mistral Small 24B (~14 GB, very comfortable in French) fits on 16 GB and delivers more narrative nuance.
- RP fine-tunes
- Look for models explicitly trained for roleplay or fiction (often labeled “RP,” “storytelling,” or “uncensored” on Hugging Face). The best fine-tunes of 2026 are built on these recent bases (Qwen 3.5, Mistral Small 24B, Gemma 4) rather than the older Llama 2 or Mistral 7B models from 2023; they follow card formats better.
- Long context
- Favor a model that performs well at 16k+ tokens without degrading: long sections require it. Check the model’s stated native context length.
- GGUF format
- For KoboldCpp, stick with quantized GGUFs. Q4_K_M is the recommended compromise; move up to Q5_K_M or Q8_0 if your VRAM allows it and you want more precision.
#Troubleshooting common issues
- SillyTavern won't connect
- Verify that the backend is running properly and that the port is correct (5001 for KoboldCpp, 11434 for Ollama). A browser test at the API URL confirms that it responds.
- Responses with visible tags
- Almost always an unsuitable instruction template. Match the format (ChatML, Gemma, Mistral…) to the model's format in Advanced Formatting.
- The character breaks character
- Strengthen the description and dialogue examples, lower the temperature slightly, and make sure your system prompt does not conflict with the card.
- Repetitions and loops
- Slightly increase the repetition penalty and add a Min-P. If it persists, the context may be saturated: enable Summarize or reduce the map size.
- Truncated responses
- Response length is too low, or the backend was launched with a smaller context than SillyTavern requires. Align the two values.
#Go further
SillyTavern is just a layer: final quality depends mostly on your backend and model. These guides help you build a solid foundation.
- KoboldCpp: installation and getting started
- The reference backend for local roleplay: GGUF, memory settings, and an API that SillyTavern supports natively.
- What is Ollama, and how does it work?
- The general-purpose alternative if you prefer a single daemon on port 11434 for all your use cases.
- Quantize the KV cache
- To support longer RP contexts on the same graphics card without blowing through VRAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.