
First, a confession. The checker we tested is ours.
People run their CV through an ATS score checker, see a number, and then spend an evening trying to push it up. We wanted to know what that number actually responds to. Not in theory, but by changing one thing at a time and watching.
So we built a clean CV for a real sounding candidate, wrote a real sounding job ad for it, and then made fourteen broken versions. Each version had exactly one common mistake in it and nothing else changed. We scanned every version five times, and the clean original ten times, using the exact same system that runs on our free ATS checker. What came back does not flatter our own tool, and we are publishing it anyway, because it is the most useful thing we know about ATS scores.
The short version: the score is a rough grade, not a measurement. The clean CV itself bounced between 85 and 90 on identical scans. Eleven of the fourteen mistakes left the score at exactly 85, even the ones that wrecked the keyword match or hid every section from the parser. Keyword stuffing earned the highest score in the test. The damage was real, it just showed up in the details under the score, not in the score.
How we tested it
The candidate is Aisha Rahman, an implementation project coordinator in Manchester with four years in software onboarding. Her clean CV is 270 words on one page: a two line summary, two jobs with results in the bullets, education, two certifications and a skills list. It is a good CV, deliberately. You learn more from breaking something that works.
The job ad is for an Implementation Project Coordinator at a software company, asking for the usual things: running onboarding projects, Jira and Asana, stakeholder management, Agile, budgets, SOPs, CSAT and KPIs, and a PRINCE2 style certification as a plus.
Then we made the versions. Each one is the clean CV with one change:
| Version | The one change |
|---|---|
| Synonyms | Every job ad term reworded. Implementation became client launch, Jira and Asana became our planning tools, CSAT became client happiness score. |
| Creative headings | Summary became Hello, I am Aisha. Experience became Where I Have Made an Impact. Skills became My Toolkit. |
| Table columns | A two column design built as one big table, the way many Word templates do it, printed to PDF. |
| Normal columns | The same two column design built with normal page layout, printed to PDF. |
| One column PDF | The clean CV printed to a simple one column PDF. This is the control for the two layouts. |
| No phone or city | Contact line cut down to email and LinkedIn. |
| Duties, not results | Every number and outcome removed. Responsible for, helped with, involved in. |
| No skills section | The skills list deleted. The same skills still appear in the bullets. |
| No summary | The two line summary deleted. |
| Typos | Six typos, including Jria, stakholders and Mangement. |
| Buzzwords | Same facts rewritten with spearheaded, leveraged, fostered, results driven and friends. |
| No dates | Every date removed from jobs, education and certifications. |
| Paragraphs | Each job written as one paragraph instead of bullets. |
| Title mismatch | The headline under her name changed to Operations Administrator. |
| Keyword stuffing | The clean CV plus a block of job ad terms pasted at the bottom, the text people hide in white font. |
Every version was scored through the same pipeline our live checker uses: the same prompt, the same model (gpt-4o-mini at temperature 0.2), and the same clean up rules that stop the checker inventing keywords. The two layout versions were printed to real PDF files with Chrome and read back with pdftotext, a standard text extractor, the same way a basic parser would read your file.
Because a score is only one view, we measured two more things for every version. An exact match rate against 20 terms copied word for word from the job ad, which is the kind of match most checkers show as a percentage. And what our deterministic parser recovered from the text: the sections, the phone number, the location and the number of bullets with results in them.
Finding one: the same CV scored 85 and 90. Ten times.
Before breaking anything, we scanned the clean CV ten times. Nothing about it changed between scans.
It scored 85 five times and 90 five times.
That is the most important number in this article, because it tells you how to read every other one. A move of five points on a rescan means nothing at all. If you changed a word, rescanned, and went from 85 to 90, you may have improved your CV or you may have just scanned it again. There is no way to tell from the score alone.
It gets starker. Across all 85 scans of all sixteen versions, the checker only ever returned three numbers: 85, 90 and 95. Not 83, not 88, not 91. Three marks on the dial. A tool with three possible readings cannot tell you much about fourteen different mistakes, and it did not.
Finding two: synonyms wiped out the keyword match. The score did not notice.
This is the one that should change how you write this week.
In the synonyms version we kept every fact and every number, and simply described them in our own words instead of the job ad's words. It is what people do naturally. The ad says risks and dependencies, you write issues and blockers. The ad says stakeholders, you write client contacts.
The checker did see it. Its list of missing keywords grew, and it named Jira, Asana and KPIs as absent. But the headline number, the thing everyone looks at, stayed at 85.
Why does this matter so much? Because a candidate never sees the employer's side, and on that side a lot of the work is search. They type Jira, or PRINCE2, or the job title, into their applicant tracking system and look at who comes up. A CV that says our planning tools instead of Jira does not come up. Nobody rejects it. Nobody ever sees it.
Finding three: creative headings hid every section from the parser
In the creative headings version, the content was identical. Only the five section titles changed, to the friendly kind you see on design heavy templates: Hello, I am Aisha, Where I Have Made an Impact, Where I Learned, Badges I Have Earned, My Toolkit.
Our parser, which looks for standard section headings the way a basic ATS does, recovered 0 of 5 sections. Down from 5 of 5. To a literal parser, this CV has a name, a contact line, and then one long unlabelled block of text. Where the skills end and the jobs begin is anyone's guess.
The score stayed at 85, because the model reading it is smart enough to understand what My Toolkit means. Plenty of real parsers are not. Use Summary, Experience, Education, Skills and Certifications. Boring headings are the ones that get filed correctly.
Finding four: two columns are fine. Tables are not.
Two column CVs have a bad reputation, so we tested two of them with identical content. One used normal page layout. The other was built as a single big table, which is how a lot of downloadable Word templates are made.
The normal layout came through the text extractor perfectly. The whole sidebar first, then the whole main column, every section and every contact detail intact.
The table version came out like this:
Build project plans in Jira and Asana, track milestones, and log risks and
dependencies each week
SKILLS
Cut average time to go live from 11 weeks to 7 weeks by standardising the
kickoff checklist
Project management
Run weekly status calls and monthly steering updates with customer
stakeholders
Stakeholder managementThe extractor read the table row by row, so her sidebar got woven into her job history. The SKILLS heading lands in the middle of her Brightdesk bullets, and each skill sits between two achievements as if it were one. Her city came loose from her contact details too, and the parser lost her location entirely.
And the score? 85, on all five scans. The keywords were all still there, just in the wrong places, so a model reading for meaning shrugged. A parser filling in database fields would not shrug. It would file Project management under a job and leave Manchester out of the location field.
If you like a two column design, keep it. Just make sure it is not built from a table. You can test any file you have in the ATS checker, which shows you the text a parser actually reads from it, word for word.
Finding five: keyword stuffing got the best score in the test
We added one block to the bottom of the clean CV: a long run of job ad terms, several of them repeated, the kind of thing people hide in white text so only machines see it.
It averaged 91. It produced the only 95 in all 85 scans. The checker's missing keyword list shrank from 2.4 items to 0.6.
We are not recommending it. We are showing you why the score is not the goal. The one change a recruiter would treat as a red flag is the one change the score liked best. White text does not stay hidden, either. The moment a system converts your file to plain text, which is the first thing most of them do, your hidden block sits there in black on the recruiter's screen. You would be trading a person's trust for four points on a checker.
Finding six: the mistakes that only a human would catch
The rest of the list follows one pattern. Each mistake is obvious to a recruiter and invisible to the score.
- Duties instead of results. Stripping every number and outcome cut her quantified bullets from 6 to 2 and turned Cut average time to go live from 11 weeks to 7 into Involved in improving the kickoff process. Score: 85. The checker did not even suggest adding figures back.
- No phone or city. Our parser flagged both as missing immediately. Score: 85. A recruiter cannot call her, and a search filtered by location will not find her.
- No summary. The checker found fewer keywords, 10.6 instead of 13.9, because the summary had been carrying several of them. Score: 85.
- No skills section. One fewer section for the parser, and the exact match lost KPI. Score: 85.
- Six typos. Jria and dependancies stopped matching the job ad. Keywords found fell to 12.2. Score: 85. A human spots a typo in seconds.
- No dates. Score: 85. A recruiter now has no idea whether she did each job for six months or six years, and will assume the worse answer.
- Buzzwords. Rewriting the same facts with spearheaded, leveraged and fostered changed nothing. Score: 85. Our 500 resume study shows how quickly a recruiter spots that register.
Two changes nudged the score up rather than down. Writing each job as a paragraph scored 90 on every scan, and changing her headline to a different job title averaged 89. Both are inside or at the edge of the checker's own wobble, and neither is a good idea. Paragraphs are harder for a person to skim, and a headline that does not match the job is the first thing a recruiter reads.
Every version, three ways
Here is the whole test in one table. The score is the average across scans. The exact match is out of 20 terms from the job ad. The parser column is what a literal reader recovered.
| Version | ATS score | Exact match | What the parser got |
|---|---|---|---|
| Clean baseline (10 scans) | 87.5 (85 to 90) | 20 of 20 | 5 of 5 sections, full contact |
| Job ad terms swapped for synonyms | 85 | 1 of 20 | 5 of 5 sections |
| Creative section headings | 85 | 20 of 20 | 0 of 5 sections |
| Two column PDF built as a table | 85 | 20 of 20 | Order scrambled, location lost |
| No phone number or city | 85 | 20 of 20 | Phone and location missing |
| Duties instead of results | 85 | 19 of 20 | Quantified bullets 6 down to 2 |
| Skills section removed | 85 | 19 of 20 | 4 of 5 sections |
| No summary | 85 | 19 of 20 | 4 of 5 sections |
| Six typos | 85 | 19 of 20 | 5 of 5 sections |
| Buzzword rewrite | 85 | 20 of 20 | 5 of 5 sections |
| No dates anywhere | 85 | 20 of 20 | 5 of 5 sections |
| Two column PDF, normal layout | 85 | 20 of 20 | 5 of 5 sections, full contact |
| One column PDF (control) | 85 | 20 of 20 | 5 of 5 sections, full contact |
| Headline title does not match the job | 89 (85 to 90) | 20 of 20 | 5 of 5 sections |
| Paragraphs instead of bullets | 90 | 20 of 20 | 5 of 5 sections |
| Keyword stuffing block | 91 (90 to 95) | 20 of 20 | 5 of 5 sections |
Read the middle and right hand columns and the real damage is obvious. Read only the score column and almost everything looks the same. The raw data for all 85 scans, including every individual score, is in our CSV file, free to reuse with a link back.
What is a good ATS score, then?
There is no universal answer, and anyone who gives you one exact number is selling something. Every checker grades differently, and the systems employers use never show candidates a score at all. On the recruiter's side they search, filter and sort, and some add their own match rating that works nothing like a public checker.
For an AI graded checker like ours, a score in the 80s means the basics are there: the sections exist, the contact details are present, and a good share of the job ad's language appears. A score in the 60s means something is genuinely missing and worth fixing. Past the mid 80s, you are mostly watching the wobble.
So stop at good, and spend the rest of your time on the three things the score cannot see.
What to check instead of the number
Six checks, about fifteen minutes, and they cover everything the score missed in our test.
- Exact terms. Put the job ad next to your CV. For each required skill you genuinely have, find the ad's exact word in your CV. If it is not there, add it.
- Plain headings. Summary, Experience, Education, Skills, Certifications. Nothing clever.
- No tables for layout. One column is safest. Two columns are fine if the template is not built from a table.
- Contact in the body. Name, phone, email, city and LinkedIn in the main text at the top, not in a header, footer or image.
- Results you can defend. Keep the numbers. Real ones only, and ones you can explain in an interview.
- See what a parser sees. Upload the actual file to an ATS checker that shows you the extracted text, and read that text top to bottom, not just the score.
What this test does not tell you
We tested one CV against one job ad. A different candidate in a different field could react differently, especially to the keyword changes, which depend heavily on how specific the ad is.
We tested one checker, our own. Other checkers use other methods. Some count exact keyword matches, some use AI grading like ours, some do both. Their scores will move differently, and we would encourage anyone running one to publish the same kind of test.
We did not test the systems employers use, because nobody outside those companies can. Workday, Greenhouse, Lever and the rest each parse and search in their own way. The parser results here come from our own deterministic parser and from pdftotext, a standard extractor, which behave like a basic ATS. They are not a guarantee of what any particular employer's system will do.
And the table layout test used one table template. Some table based designs will extract more cleanly than ours did, and some will be worse.
Everything above, in seven lines.
- 1The same unchanged CV scored 85 five times and 90 five times. Five points on a rescan means nothing.
- 2Across 85 scans the checker only ever said 85, 90 or 95. No mistake pushed it below 85.
- 3Synonyms took the exact job ad match from 20 of 20 to 1 of 20. The score did not move.
- 4Creative headings left the parser with 0 of 5 sections. The score did not move.
- 5Two columns in normal layout read cleanly. The same design built as a table came out scrambled.
- 6Keyword stuffing got the best score in the test. A recruiter would bin it.
- 7Stop at a good score. Then check exact terms, plain headings, layout and contact details.
Frequently asked questions
Is my ATS score accurate?
It is accurate about some things and blind to others. In our test the same unchanged CV scored 85 on five scans and 90 on five more, so any change of 5 points or less is noise. And the score stayed at 85 through mistakes that would genuinely hurt you, like swapping the job ad's terms for synonyms or using headings a parser could not recognise. Treat the score as a rough grade and read the details underneath it: the missing keywords and what the parser recovered.
What is a good ATS score?
There is no universal scale, because every checker grades differently and most employer systems never show you a score at all. On AI graded checkers like ours, a score in the 80s usually means the basics are in place. Past that point, chasing the number is a poor use of time. In our test a CV with a keyword stuffing block scored higher than the clean original, which tells you the top of the scale is not the same thing as a better CV.
Why does my ATS score change when I scan the same resume twice?
Because AI graded checkers are not calculators. They ask a language model to judge the CV, and the model can land on a slightly different grade each time. We scanned one unchanged CV ten times and got 85 five times and 90 five times. If your score moves by a few points on a rescan, nothing about your CV changed.
Does keyword stuffing work on ATS?
On a scoring checker, a little. Our stuffed version, the clean CV plus a block of job ad terms pasted at the bottom, averaged 91 and produced the only 95 in the whole test. On a person, no. The block reads as spam the moment a recruiter sees it, and hidden white text becomes visible the moment a system converts your file to plain text. A few points on a checker is not worth being binned by the human who reads it next.
Do two column resumes fail ATS?
It depends on how the columns are built. Our two column PDF made with normal page layout came out of the text extractor clean: sidebar first, then the main column, nothing lost. The same content built as a table, the way many Word templates do it, came out scrambled, with skills woven between job bullets and the location separated from the contact block. The score did not notice either way. If you use two columns, avoid table based templates.
Should I use the exact words from the job description?
Yes, for skills you really have. This was the biggest hidden effect in our test. Swapping the job ad terms for synonyms dropped the exact match from 20 of 20 terms to 1 of 20, and the checker found 3.4 of its keywords instead of 13.9. The overall score did not move. Recruiters often search their applicant database for exact terms, so writing issue tracking when the ad says risks and dependencies can take you out of a search without anyone reading your CV.
Do creative section headings hurt ATS parsing?
They did in our test. Renaming the sections to things like Where I Have Made an Impact and My Toolkit left our parser recovering 0 of 5 standard sections, down from 5 of 5. The score stayed at 85. Use plain headings: Summary, Experience, Education, Skills, Certifications.
Does an ATS reject resumes with typos?
Typos did not lower the score in our test, but they did cost matching. Misspelling dependencies and Jira meant those words no longer matched the job ad, and the checker found 12.2 keywords instead of 13.9. A person will notice the typos even when the software does not.
Is it bad to leave off my phone number or city?
It is bad for a human and invisible to the score. Our parser flagged the missing phone and location immediately, and the score did not change at all. A recruiter who wants to call you, or who filters applicants by location, cannot do either.
Which ATS checker did you test?
Our own, the free FreeCV ATS checker. We ran the exact production pipeline, same prompt, same model and same settings, and we are publishing results that do not flatter it. Other checkers use different methods and may behave differently. The raw data for all 85 scans is in our CSV file.