HR: resume screening automated
To screen resumes locally, extract the text with PyMuPDF, redact identities, have a model such as Qwen 3.5 9B extract the facts using a required JSON schema, calculate experience and score in code from a criteria grid, and keep a human in the loop for the decision. Assisted screening is not prohibited; purely automated decision-making is (GDPR, Art. 22), and Annex III of the AI Act classifies these tools as high risk.
Sorting 200 résumés to call ten people is repetitive work where assistance can genuinely save time, provided you don’t turn a language model into an opaque decision-maker. This guide builds a local pipeline whose every result is explained and verifiable, sets the record straight on what the law actually says, and offers tests to detect bias before it affects candidates.
#What we’re building: a triage aid, not a decision-maker
A local resume-analysis pipeline reads a folder of PDFs and a job description, extracts verifiable facts for each candidate (positions, dates, cited skills, education), compares them against criteria written by the recruiter, and produces an explained ranking. The difference from a simple “AI score” comes down to two choices: calculations happen in code, not in the model, and each criterion is tied to a passage in the resume that the recruiter can review. The result is an ordered list of candidates to review, never a rejection list.
Local execution closes one door: sending candidate data to a third-party provider. It doesn't close others: the model may reproduce biases, misread a resume formatted in columns, or invent a skill. This guide therefore gives equal attention to the pipeline mechanics and the safeguards that make its use defensible.
#The legal framework: what is established and what is often exaggerated
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
- Lifetime online access
- PDF + files
- Lifetime updates
You may sometimes read that sending resumes to a hosted AI service has been prohibited in France since 2024. No law imposes any such general prohibition. The real issues are more nuanced, and they also apply locally. The first is Article 22 of the GDPR: data subjects have the right not to be subject to a decision based solely on automated processing, including profiling, that produces legal effects or similarly significantly affects them. A hiring rejection falls within this scope if no human actually makes the decision.
The second is the European AI Act. Its Annex III classifies systems intended for the recruitment or selection of people as high-risk systems, including systems for analyzing and filtering applications and evaluating candidates. The timeline has shifted: according to the Bird & Bird law firm's analysis, the provisional agreement of May 7, 2026, on the so-called Digital Omnibus package postpones the Annex III high-risk obligations to December 2, 2027. Check the Official Journal of the European Union for the final adopted text before setting your compliance timeline: this guide is not legal advice.
| Topic | What you need | Practical translation |
|---|---|---|
| Automated decision-making (GDPR, Art. 22) | A human actually decides | The output is a list to review, with no automatic refusal |
| High risk (AI Act, Annex III) | Risk management, human oversight, documentation, traceability | Input and output logs, review procedure |
| Candidate information | Inform them about the processing | Mention in the job offer and hiring policy |
| Minimization | Process only the data relevant to the job | Redacting contact details and personal information before the model |
| Non-discrimination | No prohibited criteria | Job description reviewed, regular bias testing |
#The stack: PDF reader, local model, required format
A PDF reader such as PyMuPDF extracts the text. Ollama serves the model: Qwen 3.5 with 9 billion parameters weighs 6.6 GB, supports a 256,000-token context, and fits on an 8 GB card; Mistral Small 24B (14 GB) requires more memory. Ollama also lets you constrain the response to a JSON schema: according to its documentation, you pass the schema in the format field. This is more reliable than simple JSON mode because the keys and types are enforced.
A résumé is a short document: two pages amount to a few thousand tokens. So the problem is not the context window, unless you also attach cover letters. The difficulty lies in the layout, which is handled in the next step.
#Extract the text and mask what is unnecessary
Many resumes use two columns, with a sidebar for skills and languages. Reading a PDF then depends on the order of the blocks in the file. PyMuPDF offers a sorting option that orders text by vertical and then horizontal coordinates; on a two-column resume, it can mix lines from both columns. The default behavior, which follows the file order, often reads two-column resumes better. Test both options on your most complex resumes and keep the one that produces coherent text.
Masking unnecessary data at the workstation is a matter of minimization: email address, phone number, mailing address, links to profiles, date of birth. Regular expressions catch regular formats; they do not find a first or last name in the middle of a text. A reliable approach is to process the CV’s first line, which almost always contains the person’s identity, and replace the civil-status details with a case identifier.
#Extract facts, calculate in code
The model must extract what is written, not calculate. An instruction such as “annees_experience = somme des expériences” produces incorrect totals: models are bad at adding durations that overlap or are expressed in months and years. So we ask the model for the positions with their start and end dates, and calculate the total in Python by merging overlapping periods.
#Compare against the criteria: a rubric instead of an invented score
Asking the model for “a score between 0 and 100” produces a number that looks precise but isn't: two runs can diverge, and no one knows what it measures. A rubric is more robust. The recruiter defines the criteria once and for all, distinguishing must-haves from nice-to-haves. For each criterion, the model returns only three possible statuses: met, not met, or not documented, quoting the resume sentence that justifies the result. The score is then calculated in code, using a rule you can explain to a candidate.
This approach has three strengths. The evidence is checked programmatically: a sentence the model claims to quote that doesn’t appear in the résumé doesn’t count. The ranking is reproducible because the score is a formula. And a required criterion that isn’t found doesn’t eliminate the candidate: it’s flagged to the recruiter, who reads the résumé, because the information may simply be phrased differently.
#Generate the list and keep a record
The final list sorts candidates by score, but also displays at the top those whose mandatory criterion is “needs review.” Keep a dated record for each resume: identifier, model version, criteria used, raw outputs. This record helps answer a candidate who asks how they were evaluated and demonstrate genuine human oversight, required for a high-risk recruitment system.
#Measure biases instead of assuming they are absent
Language-model bias in résumés is documented. A University of Washington study presented in October 2024 varied first names associated with white and Black people, men and women, across more than 550 real résumés: the tested models favored names associated with white people in 85% of cases and those associated with women in only 11%, and never preferred a name associated with a Black man over one associated with a white man. These models were ranking systems, not ones you would run; the lesson is that bias can exist without anyone seeing it.
Two measurements are essential. The first is a permutation test: take a few dozen résumés, replace the first name with names associated with other origins or the opposite gender, and compare the results. Without masking, the difference should be zero; if it isn’t, the pipeline is not usable. The second is outcome tracking: if you can reconstruct categories after sorting, compare the preselection rates. An unexplained difference should be addressed by revising the job description and scoring rubric.
#Verify that the tool saves time
Compare the tool's ranking with that of a recruiter for about thirty resumes from a position that has already been filled. Count how many candidates selected by the recruiter appear near the top of the list, and which good candidates the tool placed too low. If too many good profiles are rejected, correct the rubric before expanding its use. A poorly worded criterion is the most common cause.
#What fails in practice
- Creative resumes and scans
- Graphical or scanned resumes produce poor text. The script rejects a file without text and reports it instead of evaluating it blindly.
- Implicit skills
- A candidate may be proficient with a tool without naming it. The undocumented status exists for that reason: it triggers human review, not rejection.
- Languages and foreign-language workflows
- Quality declines with multilingual resumes or degree equivalencies. Test them explicitly.
- Scope drift
- Vague or outdated criteria (“dynamic junior”) introduce age bias. Review them before injecting them.
- Overconfidence
- A ranking reassures people and encourages them not to reread. Measure the share of resumes actually reread by the recruiter.
- Local AI in the enterprise: GDPR and sovereignty
- The AI Act and open-weight models
- Privacy checklist
- Structured JSON outputs with Ollama
- Limit hallucinations in a local LLM
- Invoice extraction: the same control pattern
- Source: Annex III of the AI Act
- Source: Article 22 of the GDPR
- Source: Bird & Bird on the Digital Omnibus agreement
- Source: University of Washington research on resume-screening bias
- Source: structured outputs from Ollama
Can you screen resumes with AI in France?+
Is it illegal to send résumés to ChatGPT?+
Which local model should you use to analyze resumes?+
Why not ask the model for a score out of 100?+
How can you check that a model is not discriminatory?+
How much time does resume screening save?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.