Intermediate 13 minCreative

Writing a novel with a local LLM: method and tools

Writing a book with AI doesn’t mean asking it to “write me a novel” and copying the result—that produces lukewarm mush. The real use is as a writing partner: a model that helps you structure your work, unblock a scene, and keep your characters consistent across hundreds of pages. This guide shows how to do that with a local LLM, where your manuscript never leaves your machine and no style filter trims down your prose.

By Léa B.·Update 2026-09-16·Tested on Windows, macOS, and Linux

#Why write a book with a local AI instead of in the cloud

A novel represents months of work and a text you do not want scraped to train a model, or rejected because a scene is violent, dark, or sexually explicit. Mainstream cloud services apply content filters and usage policies that, for adult fiction or hard-boiled crime fiction, regularly kill the momentum at the worst possible moment.

An LLM running on your own machine solves both problems at once. The text stays on your disk: no draft and no spoiler from your plot is sent to a third party. And you choose the model—including versions without style restrictions that write the scene you request, without lectures or refusals. A significant bonus for a long project: no monthly subscription or token counter running while you refine chapter 12 for the tenth time.

Total privacy
Your manuscript, notes, and plot remain local. Nothing is used to train a third-party model.
No style filter
Adult fiction, horror, raw thriller: the model writes whatever the scene calls for without moral restrictions.
Fixed cost
Once the hardware has paid for itself, you can generate as many drafts as you want without usage-based billing.
Available offline
Write on the train or in a cabin without a network, without depending on a connection.

#Prerequisites: hardware and stack

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Creative writing is less demanding than coding: you don't need the largest model on the market, but one that handles French well and preserves the tone. A 12 GB GPU (RTX 3060 12GB, RTX 4070) already runs a 12B–14B in Q4_K_M comfortably, which is the right tier for prose. With 16 GB (RTX 4080) you can move up to a 24B, and 24 GB (RTX 4090) or a Mac M4 Pro with unified memory opens the door to 32B models, which are noticeably more nuanced in long scenes.

3B (~2 GB VRAM)
Quick troubleshooting and reformulations. Too short to maintain a tone across a chapter.
7B (~5 GB)
Usable for brainstorming and short scenes on a small 8 GB setup.
14B (~9 GB)
The sweet spot for a 12 GB GPU: decent prose in French and good context memory.
32B (~19 GB)
The best local-quality compromise for a serious novel, starting at 24 GB of VRAM or a Mac with unified memory.

On the software side, the simplest stack remains Ollama as a daemon (it listens on http://localhost:11434) paired with an interface. Open WebUI or LM Studio work for chatting as you go; for a real writing project, an interface designed for characters and persistent context like SillyTavern is a game changer (see “Going further”).

Terminal — install Ollama and run a creative model
# Linux / macOS
curl -fsSL https://ollama.com/install.sh | sh

# Un modèle solide en français, quantifié Q4_K_M
ollama pull mistral-small

# Vérifier qu'il répond
ollama run mistral-small "Écris la première phrase d'un polar qui se passe à Marseille."

#Which local models write best in French

Not all models perform equally well in literary French. Many, trained mostly on English, produce correct but flat prose, with turns of phrase calqued from English (« She took a deep breath »). For fiction, favor models known for multilingual performance and variants fine-tuned for writing.

Mistral Small (24B)
French origins, excellent idiomatic French. The default first choice for a novel; it maintains the tone well over long passages.
Gemma (12B–27B)
Good prose, fairly flexible style; the 27B is very capable in dialogue if you have the VRAM.
Qwen (14B–32B)
Highly capable, with a comfortable long context; French is good but sometimes less “natural” than Mistral.
Creative fine-tuned variants
On Ollama and Hugging Face, “writer” / “storyteller” derivatives or uncensored models write scenes that base models refuse. Test them according to your genre.
→
Test before committing
Give the same test passage to two or three models (a tense dialogue scene, a descriptive paragraph) and compare them by ear. The “best” prose is subjective: the model that matches your voice matters more than the one that wins the benchmarks.
i
Context: the key battleground
For a novel, context-window size matters as much as prose quality: it determines how much text the model “sees” at once. Aim for at least 16k–32k usable context tokens, and keep VRAM free for that context in addition to the model weights.

#The chapter-by-chapter method that holds up

A beginner's mistake is asking for the entire novel. An LLM has no big-picture view; it excels at short, well-defined tasks. The method that works proceeds from general to specific, in layers, while always keeping control of the structure. You remain the author; the model is a screenwriter and first-draft writer to whom you dictate the direction.

  1. 01
    1. The pitch and theme
    State the subject, genre, and central conflict in two sentences. Ask the model for three or four hook variations to find the angle, then lock in yours. It's your compass for everything that follows.
  2. 02
    2. The plan by acts, then chapters
    Ask for a 3-act breakdown, then develop each act into chapters, with a one-sentence summary (“what happens”) and a one-sentence stake (“what changes”) for each. You get a one-page synopsis that you can rewrite to your liking.
  3. 03
    3. The chapter beat sheet
    Before writing a chapter, break its summary into 5 to 8 “beats”: the scene's micro-steps. The model works much better on a precise beat (“Léa discovers the letter and understands the lie”) than on “write chapter 4.”
  4. 04
    4. The first draft, beat by beat
    Generate prose one beat at a time, providing the current beat and a reminder of the characters present. Correct or restart each passage before moving to the next: you're stitching the chapter together, not dumping it all at once.
  5. 05
    5. The review pass
    Once the chapter has been written by hand and on the computer, ask the model for a targeted review: “spot repetitions,” “tighten the dialogue,” “flag timeline inconsistencies.” Never accept a blind full rewrite.
!
The model has no taste—you do
An LLM quickly produces bland, “acceptable” text. If you approve everything it suggests, your novel will have a generic voice. Treat each output as an assistant's draft: keep the idea, but rewrite the wording in your own words. The voice is yours.

#Character sheets and world bible

A novel's consistency rests on a “bible”: a reference document that you and the model consult continuously. Without one, your heroine will have green eyes in chapter 2 and blue eyes in chapter 18, and the local café will change its name. Create this document during the outline stage and enrich it as you go.

Character profile
Name, age, appearance in 3 traits, voice/verbal tics, goal, inner wound, arc. Half a page per main character.
World bible
Recurring locations, world rules (for SF/fantasy), event timeline, important objects.
Tone glossary
Two or three example sentences capturing “the book’s voice” that you paste at the top of the prompt to set the style for each session.
Continuity log
One line per chapter: what each character knows at that point in the story. Prevents revelations from appearing too early or twice.
example — fiche-personnage.md
# Léa Vasseur
- Âge : 34 ans, journaliste d'investigation à Marseille
- Apparence : cheveux courts châtains, cicatrice au sourcil gauche, toujours en veste militaire
- Voix : phrases courtes, ironie sèche, jure quand elle est acculée
- Objectif : prouver l'innocence de son frère
- Blessure : culpabilité de ne pas l'avoir cru à l'époque
- Arc : de la méfiance systématique vers la capacité à faire confiance
- Sait au ch.4 : que la lettre existe. Ignore encore : qui l'a écrite.

In practice, you paste the relevant sheet or sheets at the top of your prompt before writing a scene (“Here are the characters present… now write the next beat”). It’s handcrafted RAG: you inject the useful context at the right time instead of hoping the model will “remember.”

#Maintain consistency across 300 pages

Here is the real technical challenge. A 300-page novel is about 120,000 words, or roughly 180,000 tokens—far beyond the context window of any local model. The model cannot have “read” your entire book. Consistency therefore comes not from its memory, but from your discipline in giving it the right context for each scene.

Cascading summaries
Keep a one-sentence summary for each chapter. Before writing chapter N, provide summaries of the previous chapters instead of their full text: ten lines are enough to establish the action.
Sliding window
Give the model the full text of the previous 1 to 2 chapters (for tone and transitions) plus summaries of everything else. You preserve immediate continuity without blowing up the context.
Prompt Bible at a Glance
The profiles of the characters present in the scene, every time. That's what prevents inconsistencies in the details.
Targeted checks
Periodically, have the model review a chapter with a specific question: “Is there a contradiction with this plot summary?” It catches factual errors well when the question is framed properly.
i
Summarization is compression
Your summary log is a compressed index of your novel. The more precise it is (who knows what, where each subplot stands), the more consistent the model remains. Update it at the end of every chapter, never afterward: those five minutes save you hours of catch-up.

#Creative settings: temperature and sampling

Sampling parameters radically change the output. For fiction, you want variety and surprise without drifting into incoherence. Temperature is the main control: low (0.3–0.6) produces restrained, predictable prose, useful for technical passages or summaries; high (0.8–1.1) unleashes creativity at the cost of a higher risk of digressions.

Temperature 0.8–1.0
The right range for narrative prose: lively without becoming incoherent. Drop to 0.5 for summaries and reference sheets.
top-p 0.9–0.95
Maintain good vocabulary diversity while cutting truly improbable choices.
Repetition penalty ~1.1
To counter the tendency of LLMs to reuse the same phrasing. Don’t set it too high, or the style becomes artificial.
Response length
Limit generation to the length of a beat or scene; very long outputs drift and lose the thread.
Ollama — generation options for fiction
{
  "model": "mistral-small",
  "prompt": "[fiches + beat]",
  "options": {
    "temperature": 0.9,
    "top_p": 0.92,
    "repeat_penalty": 1.1,
    "num_ctx": 16384
  }
}

#Common pitfalls and troubleshooting

Recognizable “AI” prose
Tics such as “a shiver ran down his spine” and “no matter what awaited him.” Explicitly ban them in the prompt and manually rewrite the signature phrases.
The model summarizes instead of writing
It delivers a scene in three sentences. Explicitly ask it to “show, don’t tell,” scene by scene, with dialogue and sensory details.
Losing the thread midway through the chapter
A sign that the context is overflowing. Reduce the scene length, increase num_ctx if the VRAM allows it, or start over with a clean summary.
Refusal or hedging
A base model limits a harsh scene. Switch to a creative fine-tuned variant, or rephrase the prompt while clearly framing the fictional context.
All characters sound the same
Inject each character’s “voice” sheet and ask for differentiated dialogue; read it aloud to spot uniformity.

#Go further

To turn this method into a comfortable writing workshop, three guides on the site build on this one. SillyTavern provides an interface designed for persistent characters and long context, ideal for managing profiles and scenes without manually pasting everything together. The guide to temperature and sampling explains how to precisely tune your model's creativity. And if you're starting from scratch with hardware and installation, the local LLM installation guide sets up the Ollama stack cleanly before you start your novel.


Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.