Writing a novel with a local LLM: method and tools
Writing a book with AI doesn’t mean asking it to “write me a novel” and copying the result—that produces lukewarm mush. The real use is as a writing partner: a model that helps you structure your work, unblock a scene, and keep your characters consistent across hundreds of pages. This guide shows how to do that with a local LLM, where your manuscript never leaves your machine and no style filter trims down your prose.
#Why write a book with a local AI instead of in the cloud
A novel represents months of work and a text you do not want scraped to train a model, or rejected because a scene is violent, dark, or sexually explicit. Mainstream cloud services apply content filters and usage policies that, for adult fiction or hard-boiled crime fiction, regularly kill the momentum at the worst possible moment.
An LLM running on your own machine solves both problems at once. The text stays on your disk: no draft and no spoiler from your plot is sent to a third party. And you choose the model—including versions without style restrictions that write the scene you request, without lectures or refusals. A significant bonus for a long project: no monthly subscription or token counter running while you refine chapter 12 for the tenth time.
- Total privacy
- Your manuscript, notes, and plot remain local. Nothing is used to train a third-party model.
- No style filter
- Adult fiction, horror, raw thriller: the model writes whatever the scene calls for without moral restrictions.
- Fixed cost
- Once the hardware has paid for itself, you can generate as many drafts as you want without usage-based billing.
- Available offline
- Write on the train or in a cabin without a network, without depending on a connection.
#Prerequisites: hardware and stack
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Creative writing is less demanding than coding: you don't need the largest model on the market, but one that handles French well and preserves the tone. A 12 GB GPU (RTX 3060 12GB, RTX 4070) already runs a 12B–14B in Q4_K_M comfortably, which is the right tier for prose. With 16 GB (RTX 4080) you can move up to a 24B, and 24 GB (RTX 4090) or a Mac M4 Pro with unified memory opens the door to 32B models, which are noticeably more nuanced in long scenes.
- 3B (~2 GB VRAM)
- Quick troubleshooting and reformulations. Too short to maintain a tone across a chapter.
- 7B (~5 GB)
- Usable for brainstorming and short scenes on a small 8 GB setup.
- 14B (~9 GB)
- The sweet spot for a 12 GB GPU: decent prose in French and good context memory.
- 32B (~19 GB)
- The best local-quality compromise for a serious novel, starting at 24 GB of VRAM or a Mac with unified memory.
On the software side, the simplest stack remains Ollama as a daemon (it listens on http://localhost:11434) paired with an interface. Open WebUI or LM Studio work for chatting as you go; for a real writing project, an interface designed for characters and persistent context like SillyTavern is a game changer (see “Going further”).
#Which local models write best in French
Not all models perform equally well in literary French. Many, trained mostly on English, produce correct but flat prose, with turns of phrase calqued from English (« She took a deep breath »). For fiction, favor models known for multilingual performance and variants fine-tuned for writing.
- Mistral Small (24B)
- French origins, excellent idiomatic French. The default first choice for a novel; it maintains the tone well over long passages.
- Gemma (12B–27B)
- Good prose, fairly flexible style; the 27B is very capable in dialogue if you have the VRAM.
- Qwen (14B–32B)
- Highly capable, with a comfortable long context; French is good but sometimes less “natural” than Mistral.
- Creative fine-tuned variants
- On Ollama and Hugging Face, “writer” / “storyteller” derivatives or uncensored models write scenes that base models refuse. Test them according to your genre.
#The chapter-by-chapter method that holds up
A beginner's mistake is asking for the entire novel. An LLM has no big-picture view; it excels at short, well-defined tasks. The method that works proceeds from general to specific, in layers, while always keeping control of the structure. You remain the author; the model is a screenwriter and first-draft writer to whom you dictate the direction.
- 011. The pitch and themeState the subject, genre, and central conflict in two sentences. Ask the model for three or four hook variations to find the angle, then lock in yours. It's your compass for everything that follows.
- 022. The plan by acts, then chaptersAsk for a 3-act breakdown, then develop each act into chapters, with a one-sentence summary (“what happens”) and a one-sentence stake (“what changes”) for each. You get a one-page synopsis that you can rewrite to your liking.
- 033. The chapter beat sheetBefore writing a chapter, break its summary into 5 to 8 “beats”: the scene's micro-steps. The model works much better on a precise beat (“Léa discovers the letter and understands the lie”) than on “write chapter 4.”
- 044. The first draft, beat by beatGenerate prose one beat at a time, providing the current beat and a reminder of the characters present. Correct or restart each passage before moving to the next: you're stitching the chapter together, not dumping it all at once.
- 055. The review passOnce the chapter has been written by hand and on the computer, ask the model for a targeted review: “spot repetitions,” “tighten the dialogue,” “flag timeline inconsistencies.” Never accept a blind full rewrite.
#Character sheets and world bible
A novel's consistency rests on a “bible”: a reference document that you and the model consult continuously. Without one, your heroine will have green eyes in chapter 2 and blue eyes in chapter 18, and the local café will change its name. Create this document during the outline stage and enrich it as you go.
- Character profile
- Name, age, appearance in 3 traits, voice/verbal tics, goal, inner wound, arc. Half a page per main character.
- World bible
- Recurring locations, world rules (for SF/fantasy), event timeline, important objects.
- Tone glossary
- Two or three example sentences capturing “the book’s voice” that you paste at the top of the prompt to set the style for each session.
- Continuity log
- One line per chapter: what each character knows at that point in the story. Prevents revelations from appearing too early or twice.
In practice, you paste the relevant sheet or sheets at the top of your prompt before writing a scene (“Here are the characters present… now write the next beat”). It’s handcrafted RAG: you inject the useful context at the right time instead of hoping the model will “remember.”
#Maintain consistency across 300 pages
Here is the real technical challenge. A 300-page novel is about 120,000 words, or roughly 180,000 tokens—far beyond the context window of any local model. The model cannot have “read” your entire book. Consistency therefore comes not from its memory, but from your discipline in giving it the right context for each scene.
- Cascading summaries
- Keep a one-sentence summary for each chapter. Before writing chapter N, provide summaries of the previous chapters instead of their full text: ten lines are enough to establish the action.
- Sliding window
- Give the model the full text of the previous 1 to 2 chapters (for tone and transitions) plus summaries of everything else. You preserve immediate continuity without blowing up the context.
- Prompt Bible at a Glance
- The profiles of the characters present in the scene, every time. That's what prevents inconsistencies in the details.
- Targeted checks
- Periodically, have the model review a chapter with a specific question: “Is there a contradiction with this plot summary?” It catches factual errors well when the question is framed properly.
#Creative settings: temperature and sampling
Sampling parameters radically change the output. For fiction, you want variety and surprise without drifting into incoherence. Temperature is the main control: low (0.3–0.6) produces restrained, predictable prose, useful for technical passages or summaries; high (0.8–1.1) unleashes creativity at the cost of a higher risk of digressions.
- Temperature 0.8–1.0
- The right range for narrative prose: lively without becoming incoherent. Drop to 0.5 for summaries and reference sheets.
- top-p 0.9–0.95
- Maintain good vocabulary diversity while cutting truly improbable choices.
- Repetition penalty ~1.1
- To counter the tendency of LLMs to reuse the same phrasing. Don’t set it too high, or the style becomes artificial.
- Response length
- Limit generation to the length of a beat or scene; very long outputs drift and lose the thread.
#Common pitfalls and troubleshooting
- Recognizable “AI” prose
- Tics such as “a shiver ran down his spine” and “no matter what awaited him.” Explicitly ban them in the prompt and manually rewrite the signature phrases.
- The model summarizes instead of writing
- It delivers a scene in three sentences. Explicitly ask it to “show, don’t tell,” scene by scene, with dialogue and sensory details.
- Losing the thread midway through the chapter
- A sign that the context is overflowing. Reduce the scene length, increase num_ctx if the VRAM allows it, or start over with a clean summary.
- Refusal or hedging
- A base model limits a harsh scene. Switch to a creative fine-tuned variant, or rephrase the prompt while clearly framing the fictional context.
- All characters sound the same
- Inject each character’s “voice” sheet and ask for differentiated dialogue; read it aloud to spot uniformity.
#Go further
To turn this method into a comfortable writing workshop, three guides on the site build on this one. SillyTavern provides an interface designed for persistent characters and long context, ideal for managing profiles and scenes without manually pasting everything together. The guide to temperature and sampling explains how to precisely tune your model's creativity. And if you're starting from scratch with hardware and installation, the local LLM installation guide sets up the Ollama stack cleanly before you start your novel.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.