Accounting: extraction of factures
To extract invoices locally, serve a vision model (Qwen 3.5 9B, 6.6 GB, or Gemma 4) with Ollama, enforce a JSON schema on its response, then validate in code: subtotal plus VAT equals total including VAT, SIRET, dates. Invoices that fail go for human review. Since September 2026, structured electronic invoices arrive without AI: this pipeline is for simple PDFs and scans.
Entering invoices by hand is slow and error-prone, but sending accounting documents to an online service raises a privacy issue. A local vision model reads the page, a schema enforces the output format, and deterministic checks filter out errors. This guide builds the complete pipeline, from sorting PDFs to import comptable, and explains what e-invoicing changes starting in September 2026.
#What we automate, and how reliably
Given a supplier invoice PDF, the goal is to obtain structured JSON (supplier, number, date, amounts excluding tax, VAT and including tax, line items) that can be batch-imported into accounting software. A vision model served by Ollama reads the page image directly, without going through a separate OCR step. The rule that makes the setup reliable can be stated in one sentence: the model reads, the code checks. No extraction is ever imported without passing arithmetic and identifier checks, and anything that fails goes into a human review queue.
This guide focuses on a concrete case: a firm or small business receiving a few hundred invoices per month on a workstation or small local server. No accuracy figure is promised here: accuracy depends on your suppliers, scan quality, and the model. The measurement method appears below, using a sample of your own invoices.
#Electronic invoicing: what changes in 2026 and 2027
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Before building a reading pipeline, you need to look at what the reform does to incoming invoices. Since September 1, 2026, according to impots.gouv.fr, large companies and mid-sized companies must issue their invoices through an approved platform, and all companies must be able to receive electronic invoices. For SMEs, small businesses, and micro-enterprises, the issuance requirement applies starting September 1, 2027, according to the tax administration’s practical guide.
The practical consequence is twofold. First, an increasing share of your incoming invoices will already arrive in structured form, with no AI reading required. Second, the Factur-X format, the French-German standard for hybrid invoices, embeds a readable PDF and XML data for automated processing in the same file, with several data profiles. When a PDF contains this XML, reading it directly is safer than having a model guess it. The pipeline below therefore applies this priority order: structured data if available, native PDF text next, vision as a last resort.
#Choosing the model: how much memory vision requires
For reading invoices, two families in the Ollama library are suitable. Qwen 3.5 with 9 billion parameters weighs 6.6 GB, accepts text and images, and advertises a 256,000-token window; its 27B version weighs 17 GB. Gemma 4, with its e4b variant, occupies between 6.6 and 9.5 GB according to the Ollama specifications and advertises 128,000 tokens of context, with text and image support. A 12 GB card is therefore enough for the 9B or e4b variant, with room for page images.
| File type | Tool | Why |
|---|---|---|
| Factur-X PDF or embedded XML | Direct XML parsing | Exact data, no risk of misreading |
| Native PDF with selectable text | Text extraction, then text LLM | Fast, no image processing |
| Scanned PDF or clean photo | Vision model (Qwen 3.5 9B, Gemma 4) | Reads the layout and tables |
| Poor-quality scan | Tesseract OCR, followed by human proofreading | Vision gets distracted by noise; a human will make the final call |
Site memory reference: a 9-billion-parameter model in Q4 weighs about 5 to 6 GB, plus the context cache and the page images. Multipage invoice pages cost more than others: a batch of ten pages in one request can exceed Ollama's default context window, which remains well below the 256,000 tokens advertised by the model until you increase it.
#Sort PDFs before reading: native, scanned, structured
Detecting the file type avoids unnecessarily sending an image to a vision model. With the PyMuPDF library, a PDF whose extracted text is more than a few hundred characters is native; a PDF with almost no text is a scan. Factur-X files contain an XML attachment that can be listed before any other processing.
For poor-quality scans, Tesseract remains a free, offline fallback: you then need to install the French language file and scan at 300 dots per inch. The Tesseract guide details the settings. Text produced by poor OCR must never pass through without review: the human proofreading queue takes over.
#Extract data with a vision model and a required schema
Ollama can constrain the model's response to a JSON schema: according to its documentation, you provide a schema in the format field, and it is also recommended to repeat it in the prompt to anchor the response. This is much more robust than asking for “JSON” in free-form text, because the keys and types are enforced. For images, the REST API expects base64-encoded images in the message's images field.
Three choices deserve an explanation. A temperature of zero makes the output reproducible. The num_ctx parameter increases the window because images from several pages quickly consume the default window of Ollama, which is reduced; the guide to the context window explains this mechanism in detail. Finally, the instruction “Don't invent anything” with a null value allowed reduces the most serious risk: a model filling an unreadable field with a plausible value. A strict schema forces the response's form, not the accuracy of its content.
#Extend the schema to your real-world cases
The basic schema covers most common vendor invoices. The extensions you need depend on your business. Add them one at a time and measure each one's effect on your sample: an overloaded schema degrades extraction of essential fields.
- Down payment and balance
- Add an acompte_paye field. Without it, the total including VAT does not match the amount remaining to be paid.
- Command references
- A bon_commande_ref field lets you match records to your orders.
- Shipping costs and discounts
- Dedicated line or port_ht field; otherwise the line-item total will not match the pre-tax total.
- VAT breakdown
- A list of rates, base amounts, and totals. Essential when the invoice combines multiple rates.
- Accounting allocation
- Do not ask the model to guess the account: produce a proposal, clearly marked as such, for your business rule or a human to confirm.
#Validate before importing: checks that catch errors
This is the step that determines the quality of the assembled result. Each check is deterministic and therefore more reliable than the model. The first is arithmetic: pre-tax total plus VAT must equal the tax-included total, with a tolerance of a few cents for rounding. The second concerns the line items: their sum must equal the pre-tax total. The third concerns dates: neither earlier than a reasonable limit nor in the future. The fourth concerns the SIRET.
SIRET validation deserves special attention. The number has 14 digits, and its last digit is a check digit calculated using the Luhn formula, according to Wikipedia. There is one exception: La Poste establishments, whose SIREN is 356000000, follow a different rule in which the sum of the 14 digits must be a multiple of 5. A naive Luhn check would therefore incorrectly reject La Poste invoices. The following code handles both cases.
#Measure accuracy on your own invoices
No published accuracy figure can replace a measurement on your corpus, because a firm processing wholesale invoices does not have the same documents as an association. Build a sample of about fifty representative invoices, enter the essential fields by hand, then compare them field by field.
- 01Build the sampleUse a varied set of real invoices: scans, native PDFs, multiple pages, and multiple vendors. Anonymize them if you share the results.
- 02Capture ground truthManually note the fields to automate: number, date, pre-tax amount, VAT, total including tax, SIRET.
- 03Compare field by fieldCalculate accuracy per field, not overall: a model may read dates perfectly but get VAT numbers wrong.
- 04Set an automation thresholdDecide which fields can be imported without review. Amounts, if they pass the arithmetic check, are good candidates; accounting classification, never.
- 05Replay on every changeChange the model, resolution, or prompt: rerun the sample before moving to production.
#Batch-process a folder of invoices
Batch processing chains sorting, extraction, and validation, then files each invoice according to the result. Always keep two files side by side: the original PDF and the extracted JSON.
#Import data into accounting software
Each editor has its own import format, and its specifications evolve: start with your tool’s documentation, not a generic model. Accounting software generally offers a structured import par file or an application programming interface. The safest approach is to produce an import conforme file, load it into a test folder, then compare the entries it creates with those you would have entered manually.
Maintain the audit trail: the original PDF, the extracted JSON, the model version, and the processing date. In the event of an audit, you must be able to trace an accounting entry back to its supporting document. Invoices contain personal and business data: keeping them local avoids entrusting that data to a third party, but file storage is still governed by your retention and security policies.
- Tesseract OCR: read a scan locally
- PaddleOCR: OCR that understands the page
- Docling: convert PDFs for local AI
- Local multimodal LLM with Ollama
- Structured JSON outputs with Ollama
- Understanding the context window
- Source: impots.gouv.fr, electronic invoicing
- Source: practical guide to e-invoicing (DGFiP)
- Source: the Factur-X format (FNFE-MPE)
- Source: structured outputs Ollama
- Source: Qwen 3.5 in the Ollama library
Can a local LLM read a scanned invoice?+
Do you need OCR in addition to the vision model?+
How do you guarantee valid JSON output?+
What does mandatory e-invoicing change for this type of tool?+
What graphics card do you need to extract invoices?+
Can you trust the accounting classification proposed by the model?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.