How to Automate Resume Screening with AI in 6 Steps
Who this is forHR staff whose AI resume scores disagree with human reviewers, and anyone who wants to tune AI grading criteria with past screening results.
TL;DR: Automating resume screening has two stages: eligibility checks (quantitative) and cover letter scoring (qualitative). Quantitative checks can be handled with rules, but AI scores for qualitative review diverge from human scores and cluster especially at the middle grades. Using past screening documents and scores as questions and answers, I organized a six-step process that fixes the criteria with a mock exam (practice set) and measures once with the CSAT (South Korea’s national college entrance exam) as the test set. The goal is not for AI to make every cut, but to separate only the extremes and hand the middle to humans.
Contents
- Prepare the questions and answers
- Split the practice exams from the CSAT
- Set and measure quantitative criteria
- Set up and measure qualitative evaluation
- Find the gap between AI and human scores
- Fix the criteria and repeat steps 3 to 5
- AI doesn’t make the final cut
- Checking whether you got this far
I consulted with an HR manager who had added AI scoring to the screening stage for technical positions. The eligibility check matched the results of the previous screening exactly. The problem was the cover letter. Scores of 1 and 5 matched human reviewers, but most of the rest clustered at 3. To the manager, there were visible differences among the 3s, but the AI could not tell them apart.
Below are the six steps we worked out in that consultation. Steps 1 and 2 are done once; steps 3 through 6 are repeated.
I am publishing an example repository that follows the steps in this article exactly. The example data is entirely synthetic; no real applications were used.
1. Prepare the Questions and Answers
To refine AI scoring, you first need questions to grade and answers to check against.
| What | In resume screening |
|---|---|
| Question | Documents from applicants in the past screening (résumés, cover letters) |
| Answer | The scores reviewers gave those documents and the pass/fail decision |
| Volume | 100~200 items, with a similar count in each score band |
You don’t need to create new answers. The scores reviewers gave in the past screening are the answer key. If several reviewers scored the same applicant, use their average.
Do not use the whole set from the start. Pick samples so that each score band has a similar count, from 1 to 5. If you pick at random, most samples land in the middle, leaving too few examples to learn the boundary between 2 and 3. Sort by the score column and take the same number from each band.
In the example repository’s data, each person’s résumé and cover letter are linked under the same ID.
| ID | Major | Certification | Past decision | Past cover letter grade |
|---|---|---|---|---|
| A01 | Mechanical engineering | General Mechanical Engineer (Korean national license) | Pass | 5 |
| B04 | Industrial engineering | Hazardous Materials Industrial Engineer | Human review | 4 |
| C01 | Chemical engineering | Industrial Safety Engineer | Pass | 3 |
| D01 | Equipment engineering | Construction Machinery Facilities Engineer | Pass | 2 |
2. Split the Practice Exams from the CSAT
Having answers doesn’t mean you should use all of them to fix the criteria. Set some aside and look at them only once, at the end.
If you study from the CSAT answer key, your CSAT score goes up. But your ability hasn’t improved. The same goes for grading criteria. If you read the wrong answers in the test set and revise the criteria, the test set score improves, but there is no guarantee the same results will appear for applicants arriving for the first time this year.
- Set aside 20~30% as the test set.
- Split within each score band. If only 5s pile up on one side, that score tells you nothing.
- Do not open test set results while you revise the criteria. In the example repository, the score report also hides the test set until you add
--final.
The example repository has only 30 items, so I split them in half. With real data, set the test set at 20~30%.
In the consultation, we pulled some past screening scores and built the grading criteria to match them. To know whether the criteria actually improved, you need a test set that was not used to build them.
3. Set and Measure Quantitative Criteria
Years of experience, major, certifications, and training hours are lookups, not judgments. Filter them with rules instead of AI. If you use AI, the same applicant can get different results each run, and then you cannot explain the reason for rejection.
Three things need to be decided when setting the criteria:
- Calculate experience against a single reference date. If you use today’s date, the same applicant’s experience changes every day.
- A major not on the list goes to human review, not rejection. Department names differ from school to school.
- Falling slightly short of a threshold goes to human review. If the requirement is 3 years and the applicant has 2.7, a person weighs the circumstances.
Once the criteria are set, first measure how well they match the past screening decisions. Mismatches mean either the rules are wrong or a person considered circumstances outside the rules. A person decides which.
In the consultation too, the quantitative step matched the past results exactly. There is almost nothing to adjust in quantitative scoring here. The hard part is the next step.
4. Set Up and Measure Qualitative Evaluation
Cover letters are scored from 1 to 5 for each item. The example repository uses five items: motivation for applying, strengths and weaknesses, experience overcoming conflict, experience with challenges, and writing.
The grading criteria are a single document, not code. Swapping this document for your own company’s scoring sheet is how you adapt this method to your own hiring. For each grade, write example sentences that do and do not qualify.
### 5 points: Has verifiable results
- Qualifies: "Recurrence of leaks dropped from 4 a year to 1"
- Does not qualify: "Was in charge of maintenance for 3 years" (the number only points to tenure, not a result)
### 4 points: Own actions are specified down to the object
- Qualifies: "Standardized the maintenance history form so the team used the same format"
- Does not qualify: "Actively improved things" (no specifics on what was improved)
Grade with one session per cover letter. This was the first thing recommended in the consultation. If you put several letters in one conversation, earlier letters can influence the scores of later ones.
5. Find the Gap Between AI and Human Scores
When grading is done, put the AI scores and past scores side by side. For each cover letter, average the five item scores, round to a grade, and compare it with the reviewer’s grade. Don’t look at the match count first.
| What to check | What it tells you |
|---|---|
| Direction of errors | If all errors are on the low side, the criteria are stingy |
| Bias by grade | If only certain grades are wrong, the boundary sentences for that grade are too loose |
| List of mismatched letters | Material for the next fix. Look only at the practice set |
6. Fix the Criteria and Repeat Steps 3 to 5
Read the mismatched letters, fix only one thing at a time in the criteria, then run again from step 3. If you fix two things at once, you can’t tell which one made the difference.
The place that takes the most work is the boundary between 2, 3, and 4 points. The AI cannot decide on its own why this letter is a 2 and that one is a 3. The criteria have to spell it out in writing, and a person has to find the sentences. Group the letters that reviewers gave 2 points with those given 3 points, ask the question below, and have a person read the answer and write it into the criteria.
Below are N cover letters that reviewers gave [2 points] and N cover letters they gave [3 points].
Find the differences that separate the two groups. Write up to 5 sentences in the form "If ~, 3 points; if it only reaches ~, 2 points" describing features that repeatedly appear in only one group, and attach two letter IDs as evidence for each sentence. Discard features that appear in only one or two letters.
The example repository’s grading criteria were revised twice, producing three versions.
| Criteria | What changed | Practice set matches (15 items) | Test set matches (15 items) |
|---|---|---|---|
| v1 | Item descriptions only, no grade boundaries | 4 | 9 |
| v2 | Added grade boundaries, instruction “if conditions are met, choose the higher grade” | 11 | 10 |
| v3 | Removed the global instruction, added examples for each grade | 14 | 10 (rerun: 12) |
The example changed two things each in v2 and v3, so you cannot tell which change made the difference in v3.
Now open the test set once. Practice-set matches rose by 10, from 4 to 14, but test-set matches rose by only 1, from 9 to 10. When I reran the same v3, the test set came to 12, so a difference of 1 to 3 items in the test set falls within run-to-run variation.
Improvement on the practice set does not carry over fully to documents the AI has never seen. This is why we set aside the test set in step 2. What clearly changed on the test set is the direction of errors: cases scored stricter than humans fell from 6 in v1 to 1 in v3.
For automating the repetition, you can look at the approach in autoresearch released by Andrej Karpathy. The evaluation code is left unchanged; the AI edits just one file, runs it, and keeps the change only when the result improves.
7. AI Doesn’t Make the Final Cut
If you try to cut sharply between 2, 3, and 4 points, the responsibility becomes far too large for the time invested. After the consultation, I passed this on to the manager like this:
Let me stress again: the task and the responsibility are large compared with the effort, so treat the AI workflow as support, and don’t try to be responsible for everything. If you get to a quantitative filter plus qualitative advice, consider that a success. Start small and expand from there. Especially with hiring decisions, human responsibility matters a great deal.
(Message sent to the manager after the consultation, by the author)
So the example repository gives decisions in only three categories.
| Decision | Criterion (average of 5 items) | Who decides |
|---|---|---|
| Pass | 4.0 or higher | AI decision, confirmed by a person |
| Human review | 2.0 or higher, below 4.0 | Reviewer reads it directly |
| Reject | Below 2.0 | AI decision, confirmed by a person |
In the 30 synthetic items, about 40% went to human review. This is the number you need when planning how many reviewers to staff. Narrowing the automatic decision bands (raising the pass threshold and lowering the reject threshold) increases the number of items people review and reduces the AI’s misjudgments.
Decisions are much more stable than grades. When I ran the same criteria twice, 6 rounded grades changed, but only 1 pass, review, or reject decision changed. That is why results are delivered as decisions rather than scores.
The delivery format agreed on in the consultation is a single Excel file. Open the CSV that the example repository produces in Excel, put the quantitative and qualitative grading criteria on the first sheet, and place each applicant’s scores and decisions on the next sheet. Reviewers read the first sheet and look directly at the applicants in the human-review rows.
8. Checking Whether You Got This Far
The figures in this article come from 30 synthetic cover letters. The grades were set first and the letters were written to match them, so the differences between grades may be sharper than in real applications. The same person set the answer grades and revised the grading criteria, and the grade example sentences quoted in step 4 come from practice-set cover letters. Do not read the match rates here as real hiring performance. The way to use this article is to measure again from step 1 with your own company’s past screening data.
Frequently asked questions
- What split should practice and test sets use for AI resume screening?
- Use 70~80% of past documents as the practice set and 20~30% as the test set. Sample 100~200 items evenly across score bands, and grade the test set only once.
- How is bias measured when comparing AI and human resume scores?
- Bias is the average of AI score minus human score. A negative value means the AI graded stricter, and a positive value means it graded more generously.
Want the full system? The Claude Code & Codex Skills guidebook collects the skills and subagents behind this blog, from $19.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›