Intermediate 11 minRH

HR: resume screening automated

Direct response

To screen resumes locally, extract the text with PyMuPDF, redact identities, have a model such as Qwen 3.5 9B extract the facts using a required JSON schema, calculate experience and score in code from a criteria grid, and keep a human in the loop for the decision. Assisted screening is not prohibited; purely automated decision-making is (GDPR, Art. 22), and Annex III of the AI Act classifies these tools as high risk.

Sorting 200 résumés to call ten people is repetitive work where assistance can genuinely save time, provided you don’t turn a language model into an opaque decision-maker. This guide builds a local pipeline whose every result is explained and verifiable, sets the record straight on what the law actually says, and offers tests to detect bias before it affects candidates.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What we’re building: a triage aid, not a decision-maker

A local resume-analysis pipeline reads a folder of PDFs and a job description, extracts verifiable facts for each candidate (positions, dates, cited skills, education), compares them against criteria written by the recruiter, and produces an explained ranking. The difference from a simple “AI score” comes down to two choices: calculations happen in code, not in the model, and each criterion is tied to a passage in the resume that the recruiter can review. The result is an ordered list of candidates to review, never a rejection list.

Local execution closes one door: sending candidate data to a third-party provider. It doesn't close others: the model may reproduce biases, misread a resume formatted in columns, or invent a skill. This guide therefore gives equal attention to the pipeline mechanics and the safeguards that make its use defensible.

The Local Agents Kit

Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

You may sometimes read that sending resumes to a hosted AI service has been prohibited in France since 2024. No law imposes any such general prohibition. The real issues are more nuanced, and they also apply locally. The first is Article 22 of the GDPR: data subjects have the right not to be subject to a decision based solely on automated processing, including profiling, that produces legal effects or similarly significantly affects them. A hiring rejection falls within this scope if no human actually makes the decision.

The second is the European AI Act. Its Annex III classifies systems intended for the recruitment or selection of people as high-risk systems, including systems for analyzing and filtering applications and evaluating candidates. The timeline has shifted: according to the Bird & Bird law firm's analysis, the provisional agreement of May 7, 2026, on the so-called Digital Omnibus package postpones the Annex III high-risk obligations to December 2, 2027. Check the Official Journal of the European Union for the final adopted text before setting your compliance timeline: this guide is not legal advice.

Requirements to integrate into the pipeline
TopicWhat you needPractical translation
Automated decision-making (GDPR, Art. 22)A human actually decidesThe output is a list to review, with no automatic refusal
High risk (AI Act, Annex III)Risk management, human oversight, documentation, traceabilityInput and output logs, review procedure
Candidate informationInform them about the processingMention in the job offer and hiring policy
MinimizationProcess only the data relevant to the jobRedacting contact details and personal information before the model
Non-discriminationNo prohibited criteriaJob description reviewed, regular bias testing
!
Two caveats
This table summarizes principles, not exhaustive requirements: retention periods, the legal basis, and the impact assessment are the responsibility of your data protection officer. And running locally does not exempt you from compliance: processing job applications remains personal data processing, regardless of where it runs.

#The stack: PDF reader, local model, required format

A PDF reader such as PyMuPDF extracts the text. Ollama serves the model: Qwen 3.5 with 9 billion parameters weighs 6.6 GB, supports a 256,000-token context, and fits on an 8 GB card; Mistral Small 24B (14 GB) requires more memory. Ollama also lets you constrain the response to a JSON schema: according to its documentation, you pass the schema in the format field. This is more reliable than simple JSON mode because the keys and types are enforced.

A résumé is a short document: two pages amount to a few thousand tokens. So the problem is not the context window, unless you also attach cover letters. The difficulty lies in the layout, which is handled in the next step.

#Extract the text and mask what is unnecessary

Many resumes use two columns, with a sidebar for skills and languages. Reading a PDF then depends on the order of the blocks in the file. PyMuPDF offers a sorting option that orders text by vertical and then horizontal coordinates; on a two-column resume, it can mix lines from both columns. The default behavior, which follows the file order, often reads two-column resumes better. Test both options on your most complex resumes and keep the one that produces coherent text.

Masking unnecessary data at the workstation is a matter of minimization: email address, phone number, mailing address, links to profiles, date of birth. Regular expressions catch regular formats; they do not find a first or last name in the middle of a text. A reliable approach is to process the CV’s first line, which almost always contains the person’s identity, and replace the civil-status details with a case identifier.

Extraction and redaction (PyMuPDF)
import fitz  # PyMuPDF
import re

MOTIFS = [
    (r'[\w.+-]+@[\w-]+\.[\w.-]+', '[EMAIL]'),
    (r'(?:\+33|0033|0)[\s.-]?[1-9](?:[\s.-]?\d{2}){4}', '[TEL]'),
    (r'https?://\S+|www\.\S+', '[LIEN]'),
    (r'\b\d{1,2}[/.]\d{1,2}[/.](?:19|20)\d{2}\b', '[DATE]'),
]

def lire_cv(chemin):
    doc = fitz.open(chemin)
    texte = '\n'.join(page.get_text() for page in doc)  # essayez aussi sort=True
    if len(texte.strip()) < 200:
        raise ValueError('CV sans texte : scan ou image, passer par un OCR')
    lignes = texte.split('\n')
    lignes[0] = '[IDENTITE]'  # la première ligne porte presque toujours le nom
    texte = '\n'.join(lignes)
    for motif, remplacement in MOTIFS:
        texte = re.sub(motif, remplacement, texte)
    return texte
i
Don't confuse masking with anonymization
A redacted resume still contains identity clues: institutions, organizations, years of study, neighborhoods. Under the GDPR, it remains personal data. Redaction reduces exposure and the risk of name-related bias; it does not make processing anonymous.

#Extract facts, calculate in code

The model must extract what is written, not calculate. An instruction such as “annees_experience = somme des expériences” produces incorrect totals: models are bad at adding durations that overlap or are expressed in months and years. So we ask the model for the positions with their start and end dates, and calculate the total in Python by merging overlapping periods.

Structured extraction and experience calculation
import json, requests
from datetime import date

SCHEMA = {'type': 'object', 'properties': {
  'competences': {'type': 'array', 'items': {'type': 'string'}},
  'langues': {'type': 'array', 'items': {'type': 'object', 'properties': {
      'nom': {'type': 'string'}, 'niveau': {'type': ['string', 'null']}}}},
  'formation_plus_haute': {'type': ['string', 'null']},
  'postes': {'type': 'array', 'items': {'type': 'object', 'properties': {
      'titre': {'type': 'string'}, 'entreprise': {'type': ['string', 'null']},
      'debut': {'type': ['string', 'null']}, 'fin': {'type': ['string', 'null']}}}}
}, 'required': ['competences', 'postes']}

PROMPT = ('Extrais du CV les informations demandées. Réponds par un JSON conforme à ce schéma : '
  + json.dumps(SCHEMA) + '. Dates au format AAAA-MM ; fin = null si le poste est en cours. '
  'Si une information est absente, mets null ou une liste vide. N\'invente rien.\n\nCV :\n')

def extraire(texte, modele='qwen3.5:9b'):
    r = requests.post('http://localhost:11434/api/chat', json={
        'model': modele, 'stream': False, 'format': SCHEMA,
        'messages': [{'role': 'user', 'content': PROMPT + texte}],
        'options': {'temperature': 0, 'num_ctx': 8192}})
    return json.loads(r.json()['message']['content'])

def mois(s):
    a, m = s.split('-')
    return int(a) * 12 + int(m)

def annees_experience(postes):
    aujourdhui = date.today(); fin_defaut = aujourdhui.year * 12 + aujourdhui.month
    plages = []
    for p in postes:
        try:
            d = mois(p['debut']); f = mois(p['fin']) if p.get('fin') else fin_defaut
        except (ValueError, KeyError, AttributeError, TypeError):
            continue
        if f >= d:
            plages.append((d, f))
    total, courant = 0, None
    for d, f in sorted(plages):
        if courant is None:
            courant = [d, f]
        elif d <= courant[1]:
            courant[1] = max(courant[1], f)
        else:
            total += courant[1] - courant[0]; courant = [d, f]
    if courant:
        total += courant[1] - courant[0]
    return round(total / 12, 1)

#Compare against the criteria: a rubric instead of an invented score

Asking the model for “a score between 0 and 100” produces a number that looks precise but isn't: two runs can diverge, and no one knows what it measures. A rubric is more robust. The recruiter defines the criteria once and for all, distinguishing must-haves from nice-to-haves. For each criterion, the model returns only three possible statuses: met, not met, or not documented, quoting the resume sentence that justifies the result. The score is then calculated in code, using a rule you can explain to a candidate.

Criterion-by-criterion evaluation
CRITERES = [
  {'id': 'sql', 'libelle': 'Pratique de SQL en contexte professionnel', 'obligatoire': True},
  {'id': 'anglais', 'libelle': 'Anglais professionnel', 'obligatoire': False},
  {'id': 'encadrement', 'libelle': 'Encadrement d\'une équipe', 'obligatoire': False},
]
SCHEMA_EVAL = {'type': 'array', 'items': {'type': 'object', 'properties': {
    'critere': {'type': 'string'},
    'statut': {'type': 'string', 'enum': ['satisfait', 'non_satisfait', 'non_documente']},
    'preuve': {'type': ['string', 'null']}},
  'required': ['critere', 'statut', 'preuve']}}

def evaluer(texte_cv, modele='qwen3.5:9b'):
    consigne = ('Pour chaque critère, indique si le CV le satisfait, ne le satisfait pas, ou ne permet pas de savoir. '
      'Pour satisfait, recopie mot pour mot la phrase du CV. N\'utilise que le CV. Critères : '
      + json.dumps(CRITERES, ensure_ascii=False))
    r = requests.post('http://localhost:11434/api/chat', json={
        'model': modele, 'stream': False, 'format': SCHEMA_EVAL,
        'messages': [{'role': 'user', 'content': consigne + '\n\nCV :\n' + texte_cv}],
        'options': {'temperature': 0, 'num_ctx': 8192}})
    return json.loads(r.json()['message']['content'])

def score(evaluation, texte_cv):
    par_id = {e['critere']: e for e in evaluation}
    a_examiner, points = [], 0
    for c in CRITERES:
        e = par_id.get(c['id'], {'statut': 'non_documente', 'preuve': None})
        ok = e['statut'] == 'satisfait' and e['preuve'] and ' '.join(e['preuve'].split()) in ' '.join(texte_cv.split())
        if ok:
            points += 2 if c['obligatoire'] else 1
        elif c['obligatoire']:
            a_examiner.append(c['id'])
    return points, a_examiner

This approach has three strengths. The evidence is checked programmatically: a sentence the model claims to quote that doesn’t appear in the résumé doesn’t count. The ranking is reproducible because the score is a formula. And a required criterion that isn’t found doesn’t eliminate the candidate: it’s flagged to the recruiter, who reads the résumé, because the information may simply be phrased differently.

#Generate the list and keep a record

The final list sorts candidates by score, but also displays at the top those whose mandatory criterion is “needs review.” Keep a dated record for each resume: identifier, model version, criteria used, raw outputs. This record helps answer a candidate who asks how they were evaluated and demonstrate genuine human oversight, required for a high-risk recruitment system.

List and audit log
from pathlib import Path
from datetime import datetime

def traiter(dossier, sortie='audit.jsonl'):
    lignes = []
    for chemin in sorted(Path(dossier).glob('*.pdf')):
        try:
            texte = lire_cv(str(chemin))
            faits = extraire(texte)
            evaluation = evaluer(texte)
            points, a_examiner = score(evaluation, texte)
        except Exception as e:
            lignes.append({'fichier': chemin.name, 'erreur': str(e)}); continue
        lignes.append({'fichier': chemin.name, 'points': points, 'a_examiner': a_examiner,
                       'experience_ans': annees_experience(faits['postes']),
                       'evaluation': evaluation, 'date': datetime.now().isoformat(),
                       'modele': 'qwen3.5:9b'})
    with open(sortie, 'a', encoding='utf-8') as f:
        for l in lignes:
            f.write(json.dumps(l, ensure_ascii=False) + '\n')
    return sorted((l for l in lignes if 'points' in l), key=lambda l: (-l['points'], l['fichier']))
!
No automatic refusals
The script rejects no one and sends no messages. Candidates outside the top of the list remain available for review: the recruiter must be able to reread them, at least by sampling. That's what distinguishes decision support from an automated decision.

#Measure biases instead of assuming they are absent

Language-model bias in résumés is documented. A University of Washington study presented in October 2024 varied first names associated with white and Black people, men and women, across more than 550 real résumés: the tested models favored names associated with white people in 85% of cases and those associated with women in only 11%, and never preferred a name associated with a Black man over one associated with a white man. These models were ranking systems, not ones you would run; the lesson is that bias can exist without anyone seeing it.

Two measurements are essential. The first is a permutation test: take a few dozen résumés, replace the first name with names associated with other origins or the opposite gender, and compare the results. Without masking, the difference should be zero; if it isn’t, the pipeline is not usable. The second is outcome tracking: if you can reconstruct categories after sorting, compare the preselection rates. An unexplained difference should be addressed by revising the job description and scoring rubric.

First-name permutation test
def ecart_permutation(texte_cv, prenom, autres_prenoms):
    base = score(evaluer(texte_cv), texte_cv)[0]
    ecarts = []
    for p in autres_prenoms:
        variante = texte_cv.replace(prenom, p)
        ecarts.append(score(evaluer(variante), variante)[0] - base)
    return ecarts  # doit rester à zéro si l'identité n'influence pas le résultat

#Verify that the tool saves time

Compare the tool's ranking with that of a recruiter for about thirty resumes from a position that has already been filled. Count how many candidates selected by the recruiter appear near the top of the list, and which good candidates the tool placed too low. If too many good profiles are rejected, correct the rubric before expanding its use. A poorly worded criterion is the most common cause.

#What fails in practice

Creative resumes and scans
Graphical or scanned resumes produce poor text. The script rejects a file without text and reports it instead of evaluating it blindly.
Implicit skills
A candidate may be proficient with a tool without naming it. The undocumented status exists for that reason: it triggers human review, not rejection.
Languages and foreign-language workflows
Quality declines with multilingual resumes or degree equivalencies. Test them explicitly.
Scope drift
Vague or outdated criteria (“dynamic junior”) introduce age bias. Review them before injecting them.
Overconfidence
A ranking reassures people and encourages them not to reread. Measure the share of resumes actually reread by the recruiter.
FAQ
Can you screen resumes with AI in France?+
Yes, under certain conditions. The GDPR prohibits basing a decision solely on automated processing, and the European AI regulation classifies these tools as high-risk systems. You must inform candidates, minimize data, keep a human who genuinely makes the decision, and log the processing. No law inherently prohibits AI-assisted screening.
Is it illegal to send résumés to ChatGPT?+
There is no general prohibition, but this is a transfer of personal data to a service provider, requiring a legal basis, notification of candidates, a data-processing agreement, and often an impact assessment. Local processing avoids this transfer without eliminating the other requirements: inform people, minimize data, and keep a human involved in the decision.
Which local model should you use to analyze resumes?+
A 9-billion-parameter model such as Qwen 3.5 9B (6.6 GB, 256,000 tokens advertised) is enough to extract facts from a résumé on an 8 GB card. Mistral Small 24B (14 GB) requires more memory. Compare them on your own résumés, using the rubric and permutation test in this guide.
Why not ask the model for a score out of 100?+
Because this number has no definition: it varies from one run to another and cannot be explained to a candidate. A criteria grid, where the model simply says satisfied, not satisfied, or not documented, with cited evidence, makes it possible to calculate a reproducible score and explain it.
How can you check that a model is not discriminatory?+
Test by permutation: replace the first name on a few dozen résumés with first names from different genders and backgrounds, then compare the results, which should remain identical. Next, mask the identity before the model, then track preselection rates by category if you can reconstruct them. An unexplained gap requires you to review the setup.
How much time does resume screening save?+
It depends on your volumes and cannot be predicted: measure it. For about thirty resumes from a position that has already been filled, compare the tool’s ranking with the recruiter’s, and time the review of the outputs. If reviewing the explanations takes almost as long as reading the resumes, the gain is small.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.