Reviewing and revising code with a local LLM (pre-review commit)
Code review by a local LLM gives you a first automated reviewer—bugs, edge cases, obvious vulnerabilities—without ever sending your proprietary code to the cloud. This guide shows how to connect a Ollama model to your Git diff, set up a pre-commit hook that comments on your changes before every commit, and, most importantly, where its capabilities stop short of a real human review.
#Why do code review locally
Pasting a diff into ChatGPT for review is convenient—until that diff contains an API key, a competitor’s business logic, or code covered by an NDA. Reviewing code with a local LLM solves the problem at its root: the model runs on your machine, and the code never leaves port 11434.
- Proprietary code
- Proprietary algorithms, business logic, architecture secrets: nothing is sent to a third party that could log it or train on it.
- NDAs and confidentiality clauses
- Many client contracts explicitly prohibit sending source code to an external service. Local is often the only compliant choice.
- Zero recurring cost
- No per-token billing. You can review every commit and every branch without watching a counter.
- Works offline
- On a train, at an air-gapped site, behind a closed corporate proxy: the review remains available.
#What an LLM detects well (and what it misses)
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Before wiring anything up, you need to set your expectations. A local coding model is good at local errors that are readable in the diff, but weak at anything that requires knowledge of the rest of the system.
- Good: local bugs
- Off-by-one, reversed condition, uninitialized variable, unclosed resource, missing error handling.
- Good: obvious vulnerabilities
- SQL injection through concatenation, hard-coded secret, unvalidated path, unsafe deserialization, basic XSS.
- Good: readability
- Confusing naming, overly long function, dead code, visible duplication in the diff.
- Low: cross-file reasoning
- It only sees the diff. A broken API contract elsewhere, a global invariant, or a remote side effect escapes it.
- Low: business intent
- It doesn't know what the code is supposed to do. It flags plausible issues, not necessarily correct ones.
- Risk: false positives and hallucinations
- It may invent a nonexistent vulnerability or suggest a fix that breaks behavior. Everything must be verified.
#Prerequisites
- Ollama installed
- The daemon listens on http://localhost:11434 by default. If you haven't already, see the Ollama installation guide.
- A Git repository
- The review relies on git diff, so it requires a version-controlled project with changes to review.
- A code model
- A code-focused model downloaded locally (see the next section for choosing based on your VRAM).
- Recommended GPU
- Optional but convenient: a RTX 3060 12 GB is enough for an 8–9B model in Q4. CPU-only works, but more slowly.
#Which local models are best for code review
For code review, choose a code-specialized model rather than a general-purpose one: it understands syntax, idioms, and language-specific pitfalls better. Choose the size based on your VRAM, using Q4_K_M (the best quality/memory tradeoff). The tags below are the 2026 generation verified on the Ollama library.
- qwen3.5:9b — ~7 GB VRAM
- The 2026 entry point. Fast, with 256k context, it runs on a RTX 3060 12 GB or an entry-level M-series Mac. Good for reviewing small diffs.
- devstral:24b — ~14 GB VRAM
- The sweet spot for most workstations. A specialist in code and agentic editing (Mistral AI, Apache 2.0), with better reasoning on subtle bugs, fits on a RTX 4080 16 GB.
- qwen3-coder:30b — ~19 GB VRAM
- Code MoE (30B, 3B active), 256k context, very fast. Significantly better quality for cross-function reasoning. Requires a RTX 4090 24 GB or a Mac with ample unified memory.
- Alternatives
- gpt-oss:20b (OpenAI open-weight, very fast) and glm-4.7-flash (MIT MoE, strong in agent mode) are good options; mistral-small (24B, good in French) works as a general-purpose fallback if you only have one model available.
#Manual review in one command
Before automating, start with a manual review on demand. The idea is to send the diff of your uncommitted changes to the model through the Ollama API and read its response in the terminal. This is the basic building block of the hook we’ll add next.
Make the script executable (chmod +x review.sh), stage your changes with git add, then run ./review.sh. You will get a list of comments that you can ignore or follow as you see fit. Nothing is blocking at this stage—it is assisted review, not a gatekeeper.
#Set up a pre-commit hook with Ollama, step by step
The next step: trigger this review automatically on every git commit via a pre-commit hook. Two approaches—a native Git hook (zero dependencies) or the pre-commit framework. We’ll detail the native hook, which is simpler to understand and audit.
- 011. Create the hook fileGit hooks live in .git/hooks/. Create .git/hooks/pre-commit (with no extension). Git runs it automatically before finalizing each commit; a nonzero exit code cancels the commit.
- 022. Write the review scriptThe hook retrieves the staged diff, sends it to Ollama, and displays the response. Choose either an informational mode (it never cancels the commit) or a blocking mode based on a severity keyword emitted by the model.
- 033. Make the hook executablechmod +x .git/hooks/pre-commit — without this, Git silently ignores it.
- 044. TestRun git add on a file with an intentional bug, then run git commit. The hook must display the model's note before committing.
- 055. Share with the team (optional)Hooks in .git/hooks/ are not versioned. To share them, version a .githooks/ directory and point to it with git config core.hooksPath .githooks.
#Polish the review prompt
The quality of LLM code review depends mainly on the prompt. A poorly scoped model buries the signal under useless style comments. Three principles make the output actionable.
- Restrict the scope
- Explicitly ask to ignore style and report only bugs, vulnerabilities, and edge cases. Otherwise, you get ten cosmetic comments per diff.
- Enforce a format
- A strict format (- file:line — issue) makes the output scannable and parseable if you want to use it later.
- Request an explicit verdict
- A final line such as 'VERDICT: OK/REVIEW' provides a simple binary signal to test in a blocking hook.
- Provide the language and context
- Specify the language and, if useful, the project's conventions. The model adapts its checks (for example, memory management in C and promises in JS).
#Honest limitations compared with human review
Let's be clear about what a local LLM reviewer does not do, to avoid a false sense of security—the worst outcome would be committing less carefully because you think you're covered.
- Vision limited to the diff
- It doesn't know the rest of the repository. A change that breaks a caller in another file goes unnoticed. Integration tests remain essential.
- No business-domain understanding
- It doesn't know whether the code does what the ticket asks. It validates the form, not the intent. A human who knows the product remains irreplaceable.
- False positives and hallucinations
- It may invent a vulnerability or an incorrect fix. Every comment must be verified before you act—never fix anything blindly.
- Does not replace linters or tests
- A linter, a type checker, and a test suite catch classes of errors deterministically. The LLM complements them; it does not replace them.
- Depends on the model and prompt
- A poorly prompted 7B misses things that a well-scaffolded 32B would catch. Quality is neither guaranteed nor reproducible down to the token.
#Go further
To explore model selection, IDE integration, or required memory in more depth, these related guides on the site build on this one:
- Best local LLM for coding in 2026
- The detailed comparison of Devstral, Qwen3-Coder, and alternatives, with VRAM and speed for each model.
- Free local Copilot in VS Code
- Go beyond CLI review: chat and refactoring in the IDE with Cline, Tabby, and CodeGeeX.
- Choose your quantization (Q4, Q5, Q8, FP16)
- Understand why Q4_K_M is recommended and how to fit a 32B model on a 16 GB GPU.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.