Guides

How to Automate Resume Screening with AI in 6 Steps

13 min read#ai-resume-screening#hr-automation#ai-automation#guide

Who this is forHR staff whose AI resume scores disagree with human reviewers, and anyone who wants to tune AI grading criteria with past screening results.

TL;DR: Automating resume screening has two stages: eligibility checks (quantitative) and cover letter scoring (qualitative). Quantitative checks can be handled with rules, but AI scores for qualitative review diverge from human scores and cluster especially at the middle grades. Using past screening documents and scores as questions and answers, I organized a six-step process that fixes the criteria with a mock exam (practice set) and measures once with the CSAT (South Korea’s national college entrance exam) as the test set. The goal is not for AI to make every cut, but to separate only the extremes and hand the middle to humans.

Contents

  1. Prepare the questions and answers
  2. Split the practice exams from the CSAT
  3. Set and measure quantitative criteria
  4. Set up and measure qualitative evaluation
  5. Find the gap between AI and human scores
  6. Fix the criteria and repeat steps 3 to 5
  7. AI doesn’t make the final cut
  8. Checking whether you got this far

I consulted with an HR manager who had added AI scoring to the screening stage for technical positions. The eligibility check matched the results of the previous screening exactly. The problem was the cover letter. Scores of 1 and 5 matched human reviewers, but most of the rest clustered at 3. To the manager, there were visible differences among the 3s, but the AI could not tell them apart.

Below are the six steps we worked out in that consultation. Steps 1 and 2 are done once; steps 3 through 6 are repeated.

1. Prepare questions and answers Documents + past review scores 2. Split practice exam and CSAT Practice 70~80% / Test 20~30% Repeat zone: practice set only 3. Set and measure quantitative criteria Career and major by rules 4. Set up and measure qualitative evaluation One session per cover letter 5. Find gap between AI and human scores Direction of errors, bias by grade 6. Fix the scoring criteria One change at a time Back to step 3 If the criteria stop changing Grade once with the test set The CSAT is taken only once
Overall six-step flow. Steps 3 to 6 inside the dashed box are repeated; the white box is passed only once, at the end.

I am publishing an example repository that follows the steps in this article exactly. The example data is entirely synthetic; no real applications were used.

1. Prepare the Questions and Answers

To refine AI scoring, you first need questions to grade and answers to check against.

What In resume screening
Question Documents from applicants in the past screening (résumés, cover letters)
Answer The scores reviewers gave those documents and the pass/fail decision
Volume 100~200 items, with a similar count in each score band

You don’t need to create new answers. The scores reviewers gave in the past screening are the answer key. If several reviewers scored the same applicant, use their average.

Do not use the whole set from the start. Pick samples so that each score band has a similar count, from 1 to 5. If you pick at random, most samples land in the middle, leaving too few examples to learn the boundary between 2 and 3. Sort by the score column and take the same number from each band.

In the example repository’s data, each person’s résumé and cover letter are linked under the same ID.

ID Major Certification Past decision Past cover letter grade
A01 Mechanical engineering General Mechanical Engineer (Korean national license) Pass 5
B04 Industrial engineering Hazardous Materials Industrial Engineer Human review 4
C01 Chemical engineering Industrial Safety Engineer Pass 3
D01 Equipment engineering Construction Machinery Facilities Engineer Pass 2

2. Split the Practice Exams from the CSAT

Having answers doesn’t mean you should use all of them to fix the criteria. Set some aside and look at them only once, at the end.

Past documents + reviewer scores (100~200 items evenly sampled across score bands) Practice set = mock exam 70~80% Test set = CSAT 20~30%, locked Solve repeatedly Grade and compare scores Read the wrong answers yourself Change one grading criterion You may look at the answers and revise. This is where skill is built. Once, at the end Check only the score Don't read the essays If you look and fix it's not skill. Scores earned by studying the answer key won't show up in front of applicants you've never seen
Practice sets are solved and revised repeatedly. The test set is locked away, and only its score is checked at the end.

If you study from the CSAT answer key, your CSAT score goes up. But your ability hasn’t improved. The same goes for grading criteria. If you read the wrong answers in the test set and revise the criteria, the test set score improves, but there is no guarantee the same results will appear for applicants arriving for the first time this year.

  • Set aside 20~30% as the test set.
  • Split within each score band. If only 5s pile up on one side, that score tells you nothing.
  • Do not open test set results while you revise the criteria. In the example repository, the score report also hides the test set until you add --final.

The example repository has only 30 items, so I split them in half. With real data, set the test set at 20~30%.

In the consultation, we pulled some past screening scores and built the grading criteria to match them. To know whether the criteria actually improved, you need a test set that was not used to build them.

3. Set and Measure Quantitative Criteria

Years of experience, major, certifications, and training hours are lookups, not judgments. Filter them with rules instead of AI. If you use AI, the same applicant can get different results each run, and then you cannot explain the reason for rejection.

Three things need to be decided when setting the criteria:

  • Calculate experience against a single reference date. If you use today’s date, the same applicant’s experience changes every day.
  • A major not on the list goes to human review, not rejection. Department names differ from school to school.
  • Falling slightly short of a threshold goes to human review. If the requirement is 3 years and the applicant has 2.7, a person weighs the circumstances.

Once the criteria are set, first measure how well they match the past screening decisions. Mismatches mean either the rules are wrong or a person considered circumstances outside the rules. A person decides which.

Terminal output from running the quantitative filter. Each applicant gets a line with a pass, human review, or reject decision and the reason, and the last line shows that all 30 of 30 matched the past decisions.
Output from running the quantitative filter in the example repository. I trimmed 18 lines from the middle.

In the consultation too, the quantitative step matched the past results exactly. There is almost nothing to adjust in quantitative scoring here. The hard part is the next step.

4. Set Up and Measure Qualitative Evaluation

Cover letters are scored from 1 to 5 for each item. The example repository uses five items: motivation for applying, strengths and weaknesses, experience overcoming conflict, experience with challenges, and writing.

The grading criteria are a single document, not code. Swapping this document for your own company’s scoring sheet is how you adapt this method to your own hiring. For each grade, write example sentences that do and do not qualify.

### 5 points: Has verifiable results
- Qualifies: "Recurrence of leaks dropped from 4 a year to 1"
- Does not qualify: "Was in charge of maintenance for 3 years" (the number only points to tenure, not a result)

### 4 points: Own actions are specified down to the object
- Qualifies: "Standardized the maintenance history form so the team used the same format"
- Does not qualify: "Actively improved things" (no specifics on what was improved)

Grade with one session per cover letter. This was the first thing recommended in the consultation. If you put several letters in one conversation, earlier letters can influence the scores of later ones.

Cover letter 1One applicant
→
New sessionNo memory of earlier chat
→
Scores per item5 items, 1–5 points
→
Session endsNext applicant starts fresh

5. Find the Gap Between AI and Human Scores

When grading is done, put the AI scores and past scores side by side. For each cover letter, average the five item scores, round to a grade, and compare it with the reviewer’s grade. Don’t look at the match count first.

What to check What it tells you
Direction of errors If all errors are on the low side, the criteria are stingy
Bias by grade If only certain grades are wrong, the boundary sentences for that grade are too loose
List of mismatched letters Material for the next fix. Look only at the practice set
Score report for 15 practice-set items graded with the first criteria version. 4 matched, 11 were scored lower than humans, 0 higher, and bias is minus 0.73. All 11 mismatches were scored lower than humans, and the test set results are hidden.
Score report for the first grading criteria, which did not write out grade boundaries. Only the practice set is shown; the test set is hidden. This output was generated from saved grading results, and the earlier grading progress lines are trimmed.

6. Fix the Criteria and Repeat Steps 3 to 5

Read the mismatched letters, fix only one thing at a time in the criteria, then run again from step 3. If you fix two things at once, you can’t tell which one made the difference.

The place that takes the most work is the boundary between 2, 3, and 4 points. The AI cannot decide on its own why this letter is a 2 and that one is a 3. The criteria have to spell it out in writing, and a person has to find the sentences. Group the letters that reviewers gave 2 points with those given 3 points, ask the question below, and have a person read the answer and write it into the criteria.

Below are N cover letters that reviewers gave [2 points] and N cover letters they gave [3 points].
Find the differences that separate the two groups. Write up to 5 sentences in the form "If ~, 3 points; if it only reaches ~, 2 points" describing features that repeatedly appear in only one group, and attach two letter IDs as evidence for each sentence. Discard features that appear in only one or two letters.

The example repository’s grading criteria were revised twice, producing three versions.

Criteria What changed Practice set matches (15 items) Test set matches (15 items)
v1 Item descriptions only, no grade boundaries 4 9
v2 Added grade boundaries, instruction “if conditions are met, choose the higher grade” 11 10
v3 Removed the global instruction, added examples for each grade 14 10 (rerun: 12)

The example changed two things each in v2 and v3, so you cannot tell which change made the difference in v3.

Each time the grading criteria changed, the direction of errors shifted Practice set, 15 items; bar length is proportional to count Scored lower than humans Match Scored higher than humans Bias v1 No grade boundaries 11 4 −0.73 v2 Added 'higher grade' instruction 11 4 +0.27 v3 Examples for each grade 14 1 +0.07 Source: synthetic cover letters from the example repository, grading results results/v1~v3. © 2026 BuildnWrite
As the criteria changed, the direction of errors in the practice set shifted each time.

Now open the test set once. Practice-set matches rose by 10, from 4 to 14, but test-set matches rose by only 1, from 9 to 10. When I reran the same v3, the test set came to 12, so a difference of 1 to 3 items in the test set falls within run-to-run variation.

Improvement on the practice set does not carry over fully to documents the AI has never seen. This is why we set aside the test set in step 2. What clearly changed on the test set is the direction of errors: cases scored stricter than humans fell from 6 in v1 to 1 in v3.

For automating the repetition, you can look at the approach in autoresearch released by Andrej Karpathy. The evaluation code is left unchanged; the AI edits just one file, runs it, and keeps the change only when the result improves.

7. AI Doesn’t Make the Final Cut

If you try to cut sharply between 2, 3, and 4 points, the responsibility becomes far too large for the time invested. After the consultation, I passed this on to the manager like this:

Let me stress again: the task and the responsibility are large compared with the effort, so treat the AI workflow as support, and don’t try to be responsible for everything. If you get to a quantitative filter plus qualitative advice, consider that a success. Start small and expand from there. Especially with hiring decisions, human responsibility matters a great deal.

(Message sent to the manager after the consultation, by the author)

So the example repository gives decisions in only three categories.

Decision Criterion (average of 5 items) Who decides
Pass 4.0 or higher AI decision, confirmed by a person
Human review 2.0 or higher, below 4.0 Reviewer reads it directly
Reject Below 2.0 AI decision, confirmed by a person

In the 30 synthetic items, about 40% went to human review. This is the number you need when planning how many reviewers to staff. Narrowing the automatic decision bands (raising the pass threshold and lowering the reject threshold) increases the number of items people review and reduces the AI’s misjudgments.

Decisions are much more stable than grades. When I ran the same criteria twice, 6 rounded grades changed, but only 1 pass, review, or reject decision changed. That is why results are delivered as decisions rather than scores.

The delivery format agreed on in the consultation is a single Excel file. Open the CSV that the example repository produces in Excel, put the quantitative and qualitative grading criteria on the first sheet, and place each applicant’s scores and decisions on the next sheet. Reviewers read the first sheet and look directly at the applicants in the human-review rows.

8. Checking Whether You Got This Far

The figures in this article come from 30 synthetic cover letters. The grades were set first and the letters were written to match them, so the differences between grades may be sharper than in real applications. The same person set the answer grades and revised the grading criteria, and the grade example sentences quoted in step 4 come from practice-set cover letters. Do not read the match rates here as real hiring performance. The way to use this article is to measure again from step 1 with your own company’s past screening data.

Frequently asked questions

What split should practice and test sets use for AI resume screening?
Use 70~80% of past documents as the practice set and 20~30% as the test set. Sample 100~200 items evenly across score bands, and grade the test set only once.
How is bias measured when comparing AI and human resume scores?
Bias is the average of AI score minus human score. A negative value means the AI graded stricter, and a positive value means it graded more generously.

Want the full system? The Claude Code & Codex Skills guidebook collects the skills and subagents behind this blog, from $19.