Intermediate 11 minFinance

Accounting: extraction of factures

Direct response

To extract invoices locally, serve a vision model (Qwen 3.5 9B, 6.6 GB, or Gemma 4) with Ollama, enforce a JSON schema on its response, then validate in code: subtotal plus VAT equals total including VAT, SIRET, dates. Invoices that fail go for human review. Since September 2026, structured electronic invoices arrive without AI: this pipeline is for simple PDFs and scans.

Entering invoices by hand is slow and error-prone, but sending accounting documents to an online service raises a privacy issue. A local vision model reads the page, a schema enforces the output format, and deterministic checks filter out errors. This guide builds the complete pipeline, from sorting PDFs to import comptable, and explains what e-invoicing changes starting in September 2026.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What we automate, and how reliably

Given a supplier invoice PDF, the goal is to obtain structured JSON (supplier, number, date, amounts excluding tax, VAT and including tax, line items) that can be batch-imported into accounting software. A vision model served by Ollama reads the page image directly, without going through a separate OCR step. The rule that makes the setup reliable can be stated in one sentence: the model reads, the code checks. No extraction is ever imported without passing arithmetic and identifier checks, and anything that fails goes into a human review queue.

This guide focuses on a concrete case: a firm or small business receiving a few hundred invoices per month on a workstation or small local server. No accuracy figure is promised here: accuracy depends on your suppliers, scan quality, and the model. The measurement method appears below, using a sample of your own invoices.

#Electronic invoicing: what changes in 2026 and 2027

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Before building a reading pipeline, you need to look at what the reform does to incoming invoices. Since September 1, 2026, according to impots.gouv.fr, large companies and mid-sized companies must issue their invoices through an approved platform, and all companies must be able to receive electronic invoices. For SMEs, small businesses, and micro-enterprises, the issuance requirement applies starting September 1, 2027, according to the tax administration’s practical guide.

The practical consequence is twofold. First, an increasing share of your incoming invoices will already arrive in structured form, with no AI reading required. Second, the Factur-X format, the French-German standard for hybrid invoices, embeds a readable PDF and XML data for automated processing in the same file, with several data profiles. When a PDF contains this XML, reading it directly is safer than having a model guess it. The pipeline below therefore applies this priority order: structured data if available, native PDF text next, vision as a last resort.

i
What this guide does not cover
Receiving invoices through an approved platform, e-reporting, and issuance requirements depend on your choice of platform and accounting software provider. This pipeline covers invoices that still arrive as simple PDFs or scans: foreign suppliers, small service providers, receipts, and digitized paper documents.

#Choosing the model: how much memory vision requires

For reading invoices, two families in the Ollama library are suitable. Qwen 3.5 with 9 billion parameters weighs 6.6 GB, accepts text and images, and advertises a 256,000-token window; its 27B version weighs 17 GB. Gemma 4, with its e4b variant, occupies between 6.6 and 9.5 GB according to the Ollama specifications and advertises 128,000 tokens of context, with text and image support. A 12 GB card is therefore enough for the 9B or e4b variant, with room for page images.

Which tool for which type of PDF
File typeToolWhy
Factur-X PDF or embedded XMLDirect XML parsingExact data, no risk of misreading
Native PDF with selectable textText extraction, then text LLMFast, no image processing
Scanned PDF or clean photoVision model (Qwen 3.5 9B, Gemma 4)Reads the layout and tables
Poor-quality scanTesseract OCR, followed by human proofreadingVision gets distracted by noise; a human will make the final call

Site memory reference: a 9-billion-parameter model in Q4 weighs about 5 to 6 GB, plus the context cache and the page images. Multipage invoice pages cost more than others: a batch of ten pages in one request can exceed Ollama's default context window, which remains well below the 256,000 tokens advertised by the model until you increase it.

#Sort PDFs before reading: native, scanned, structured

Detecting the file type avoids unnecessarily sending an image to a vision model. With the PyMuPDF library, a PDF whose extracted text is more than a few hundred characters is native; a PDF with almost no text is a scan. Factur-X files contain an XML attachment that can be listed before any other processing.

Route based on the PDF type
import fitz  # PyMuPDF

def type_pdf(chemin):
    doc = fitz.open(chemin)
    pieces = doc.embfile_names()  # pièces jointes intégrées
    if any(n.lower().endswith('.xml') for n in pieces):
        return 'structure'
    texte = ''.join(page.get_text() for page in doc)
    return 'natif' if len(texte.strip()) > 300 else 'scan'

For poor-quality scans, Tesseract remains a free, offline fallback: you then need to install the French language file and scan at 300 dots per inch. The Tesseract guide details the settings. Text produced by poor OCR must never pass through without review: the human proofreading queue takes over.

Install Tesseract with French
# macOS
brew install tesseract tesseract-lang

# Ubuntu / Debian
sudo apt install tesseract-ocr tesseract-ocr-fra

#Extract data with a vision model and a required schema

Ollama can constrain the model's response to a JSON schema: according to its documentation, you provide a schema in the format field, and it is also recommended to repeat it in the prompt to anchor the response. This is much more robust than asking for “JSON” in free-form text, because the keys and types are enforced. For images, the REST API expects base64-encoded images in the message's images field.

Structured invoice extraction (Ollama, vision)
import base64, json, requests
from io import BytesIO
from pdf2image import convert_from_path

SCHEMA = {
  'type': 'object',
  'properties': {
    'fournisseur': {'type': 'object', 'properties': {
      'nom': {'type': 'string'},
      'siret': {'type': ['string', 'null']},
      'tva_intra': {'type': ['string', 'null']}}},
    'facture': {'type': 'object', 'properties': {
      'numero': {'type': 'string'},
      'date': {'type': 'string'},
      'echeance': {'type': ['string', 'null']}}},
    'montants': {'type': 'object', 'properties': {
      'ht': {'type': 'number'}, 'tva': {'type': 'number'},
      'ttc': {'type': 'number'}, 'devise': {'type': 'string'}}},
    'lignes': {'type': 'array', 'items': {'type': 'object', 'properties': {
      'description': {'type': 'string'}, 'quantite': {'type': 'number'},
      'pu_ht': {'type': 'number'}, 'total_ht': {'type': 'number'}}}}
  },
  'required': ['fournisseur', 'facture', 'montants']
}

PROMPT = ("Tu extrais les données d'une facture française. "
  "Réponds uniquement par un JSON conforme à ce schéma : " + json.dumps(SCHEMA) +
  ". Si une donnée est absente ou illisible, mets null. Les montants sont des nombres (1234.56), "
  "les dates au format AAAA-MM-JJ. N'invente rien.")

def pages_b64(pdf, dpi=200):
    sortie = []
    for img in convert_from_path(pdf, dpi=dpi):
        buf = BytesIO(); img.save(buf, format='PNG')
        sortie.append(base64.b64encode(buf.getvalue()).decode())
    return sortie

def extraire(pdf):
    r = requests.post('http://localhost:11434/api/chat', json={
      'model': 'qwen3.5:9b', 'stream': False, 'format': SCHEMA,
      'messages': [{'role': 'user', 'content': PROMPT, 'images': pages_b64(pdf)}],
      'options': {'temperature': 0, 'num_ctx': 16384}})
    return json.loads(r.json()['message']['content'])

Three choices deserve an explanation. A temperature of zero makes the output reproducible. The num_ctx parameter increases the window because images from several pages quickly consume the default window of Ollama, which is reduced; the guide to the context window explains this mechanism in detail. Finally, the instruction “Don't invent anything” with a null value allowed reduces the most serious risk: a model filling an unreadable field with a plausible value. A strict schema forces the response's form, not the accuracy of its content.

#Extend the schema to your real-world cases

The basic schema covers most common vendor invoices. The extensions you need depend on your business. Add them one at a time and measure each one's effect on your sample: an overloaded schema degrades extraction of essential fields.

Down payment and balance
Add an acompte_paye field. Without it, the total including VAT does not match the amount remaining to be paid.
Command references
A bon_commande_ref field lets you match records to your orders.
Shipping costs and discounts
Dedicated line or port_ht field; otherwise the line-item total will not match the pre-tax total.
VAT breakdown
A list of rates, base amounts, and totals. Essential when the invoice combines multiple rates.
Accounting allocation
Do not ask the model to guess the account: produce a proposal, clearly marked as such, for your business rule or a human to confirm.

#Validate before importing: checks that catch errors

This is the step that determines the quality of the assembled result. Each check is deterministic and therefore more reliable than the model. The first is arithmetic: pre-tax total plus VAT must equal the tax-included total, with a tolerance of a few cents for rounding. The second concerns the line items: their sum must equal the pre-tax total. The third concerns dates: neither earlier than a reasonable limit nor in the future. The fourth concerns the SIRET.

SIRET validation deserves special attention. The number has 14 digits, and its last digit is a check digit calculated using the Luhn formula, according to Wikipedia. There is one exception: La Poste establishments, whose SIREN is 356000000, follow a different rule in which the sum of the 14 digits must be a multiple of 5. A naive Luhn check would therefore incorrectly reject La Poste invoices. The following code handles both cases.

Consistency checks for an extracted invoice
from datetime import date

def luhn_ok(s):
    total = 0
    for i, c in enumerate(reversed(s)):
        d = int(c)
        if i % 2 == 1:
            d *= 2
            if d > 9:
                d -= 9
        total += d
    return total % 10 == 0

def siret_valide(s):
    if not (s.isdigit() and len(s) == 14):
        return False
    if s.startswith('356000000'):  # La Poste : somme multiple de 5
        return sum(int(c) for c in s) % 5 == 0
    return luhn_ok(s)

def valider(f):
    erreurs = []
    m = f['montants']
    if abs(m['ht'] + m['tva'] - m['ttc']) > 0.02:
        erreurs.append('HT + TVA différent du TTC')
    lignes = f.get('lignes') or []
    if lignes and abs(sum(l['total_ht'] for l in lignes) - m['ht']) > 0.05:
        erreurs.append('somme des lignes différente du HT')
    siret = (f['fournisseur'].get('siret') or '').replace(' ', '')
    if siret and not siret_valide(siret):
        erreurs.append('SIRET invalide : ' + siret)
    try:
        d = date.fromisoformat(f['facture']['date'])
        if d.year < 2000 or d > date.today():
            erreurs.append('date suspecte : ' + str(d))
    except ValueError:
        erreurs.append('date illisible')
    return erreurs
!
Never import an inconsistency
An invoice whose total doesn’t match should go into the “to review” queue. An extraction that passes every check isn’t necessarily correct, but those that fail them definitely need review: human triage focuses there.

#Measure accuracy on your own invoices

No published accuracy figure can replace a measurement on your corpus, because a firm processing wholesale invoices does not have the same documents as an association. Build a sample of about fifty representative invoices, enter the essential fields by hand, then compare them field by field.

  1. 01
    Build the sample
    Use a varied set of real invoices: scans, native PDFs, multiple pages, and multiple vendors. Anonymize them if you share the results.
  2. 02
    Capture ground truth
    Manually note the fields to automate: number, date, pre-tax amount, VAT, total including tax, SIRET.
  3. 03
    Compare field by field
    Calculate accuracy per field, not overall: a model may read dates perfectly but get VAT numbers wrong.
  4. 04
    Set an automation threshold
    Decide which fields can be imported without review. Amounts, if they pass the arithmetic check, are good candidates; accounting classification, never.
  5. 05
    Replay on every change
    Change the model, resolution, or prompt: rerun the sample before moving to production.

#Batch-process a folder of invoices

Batch processing chains sorting, extraction, and validation, then files each invoice according to the result. Always keep two files side by side: the original PDF and the extracted JSON.

Batch processing with a review queue
from pathlib import Path
import json

def traiter(dossier_in, dossier_ok, dossier_revue):
    for pdf in Path(dossier_in).glob('*.pdf'):
        try:
            f = extraire(str(pdf))
            erreurs = valider(f)
        except Exception as e:
            f, erreurs = {}, ['échec extraction : ' + str(e)]
        dest = Path(dossier_ok if not erreurs else dossier_revue)
        (dest / (pdf.stem + '.json')).write_text(
            json.dumps({'donnees': f, 'erreurs': erreurs}, ensure_ascii=False, indent=2))
        pdf.rename(dest / pdf.name)
        print(('OK ' if not erreurs else 'A VERIFIER ') + pdf.name)

#Import data into accounting software

Each editor has its own import format, and its specifications evolve: start with your tool’s documentation, not a generic model. Accounting software generally offers a structured import par file or an application programming interface. The safest approach is to produce an import conforme file, load it into a test folder, then compare the entries it creates with those you would have entered manually.

Maintain the audit trail: the original PDF, the extracted JSON, the model version, and the processing date. In the event of an audit, you must be able to trace an accounting entry back to its supporting document. Invoices contain personal and business data: keeping them local avoids entrusting that data to a third party, but file storage is still governed by your retention and security policies.

FAQ
Can a local LLM read a scanned invoice?+
Yes, a vision model such as Qwen 3.5 9B or Gemma 4 reads the image of a page and extracts its fields. Reliability depends on scan quality and layout. You must therefore validate the result with arithmetic checks and send any invoice that fails to a human.
Do you need OCR in addition to the vision model?+
Not necessarily. A vision model reads the image directly, avoiding the information loss of a separate OCR step. Tesseract remains useful as a fallback for heavily degraded scans and for native PDFs where you can simply extract the text. Test both on your sample instead of assuming.
How do you guarantee valid JSON output?+
Ollama lets you pass a JSON schema in the API's format field, constraining the response to that structure. This guarantees the shape, not the correctness of the values. So always add consistency checks in code: subtotal before tax plus VAT, SIRET, dates, and line-item totals.
What does mandatory e-invoicing change for this type of tool?+
Since September 1, 2026, all businesses must be able to receive electronic invoices, and issuing them becomes mandatory for SMEs in September 2027. Structured invoices, such as Factur-X, can be read without AI. The pipeline is mainly for suppliers that still send simple PDFs or paper invoices.
What graphics card do you need to extract invoices?+
A 9-billion-parameter model in Q4 weighs about 6 GB, plus page images and context. A 12 GB card is suitable; an 8 GB card is enough for one-page invoices with a reduced context. Without a GPU, processing works but becomes slow: plan for overnight processing.
Can you trust the accounting classification proposed by the model?+
No, not without validation. The model can suggest an account based on the vendor label, but that suggestion must be confirmed by a business rule or the accountant. Misclassification is silent: it passes all arithmetic checks. Treat it as the one area that always requires human review.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.