After I published the piece tracing the 6-second resume rule back to its source, a reader sent me a short, slightly panicked email. She had used ChatGPT to tighten the bullet points on her resume, kept every fact her own, and then read a LinkedIn post claiming applicant tracking systems now flag AI-written applications automatically. Her question was simple: can they tell? I spent a week trying to answer it properly, which meant reading the actual detection literature and looking at what hiring software vendors have actually shipped rather than what career blogs assume they shipped. The answer turned out to be more interesting than yes or no.
First, the thing almost every article gets wrong: your ATS is not running an AI detector
I went looking for a mainstream applicant tracking system that scores resume prose for AI authorship, and I could not find one. What the big vendors actually announced in 2025 was something adjacent but categorically different. Greenhouse launched Real Talent, later extended through a partnership with CLEAR that lets candidates verify themselves by matching a government ID to a selfie. Read the feature list closely and it is identity verification, bot and spam filtering, and fraud pattern detection. Those features answer the question "is there a real, singular human behind this application," which matters enormously in a market flooded with automated submissions. They do not answer "did a language model help write this sentence," and the vendors are not claiming they do.
That distinction gets flattened constantly in job-search content, usually into some version of "ATS now detects AI resumes." It is worth holding the two apart, because the fraud problem is real and growing while the prose-detection problem is, as far as the evidence goes, not solvable at resume length.
What happens when you point a real AI detector at real human writing
The best study on this remains a 2023 paper published in Patterns, a Cell Press journal, by Weixin Liang and colleagues at Stanford. They ran seven widely used commercial GPT detectors against two sets of unambiguously human writing: essays by US eighth-graders and TOEFL essays written by non-native English speakers under supervised exam conditions. The detectors handled the eighth-graders with near-perfect accuracy. On the TOEFL essays they misclassified more than 61% as AI-generated00130-7), and one detector flagged close to 98% of them.
The mechanism is the part worth understanding, because it explains everything that follows. These tools largely work on perplexity, a measure of how statistically surprising the next word in a sentence is. Text that uses common words in common arrangements scores as low-perplexity, and low-perplexity is what these tools treat as the signature of a machine. Writers with a smaller working vocabulary in English produce lower-perplexity text. The researchers confirmed this directly: when they enriched the vocabulary in the non-native samples, misclassification dropped, and when they simplified the native-speaker samples, misclassification rose. The detector is not measuring who wrote the text. It is measuring how ordinary the word choices are.
OpenAI graded its own homework and failed
In January 2023 OpenAI shipped an AI Text Classifier and published its own numbers: it correctly identified 26% of AI-written text while incorrectly flagging 9% of human-written text as machine-generated. It also stated the tool was unreliable on anything under 1,000 characters. On July 20, 2023 the company withdrew it, citing the low accuracy rate. The organization with the most complete access to how these models generate text could not build a reliable detector for them.
A resume is precisely the kind of document these tools decline to score
Here is the part I have not seen anyone work through, and it comes straight from the detector vendors' own documentation. Turnitin, the most widely deployed AI writing detector in the world, raised its minimum submission length from 150 to 300 words specifically to reduce false positives, and it only analyzes text written in standard grammatical sentences. Bullet points, lists, and other non-sentence structures are excluded from analysis entirely.
Now hold a resume up against that rule. A typical one-page resume runs 500 to 700 words, and the overwhelming majority of those words sit inside bullet fragments that start with a verb and never form a complete sentence. The only continuous prose on the page is usually the professional summary, which most style guides (including ours) cap at three or four sentences, roughly 60 to 80 words and often under 500 characters. Under Turnitin's own stated rules, nearly the entire document is unscoreable and the one scoreable section falls to a fifth of the minimum length. Under OpenAI's stated rules, the summary sat below the character floor even before the tool was withdrawn.
| Tool | Stated requirement | What a resume gives it | Result |
|---|---|---|---|
| Turnitin AI writing detection | 300+ words, complete grammatical sentences only; bullets and lists excluded | A 60 to 80 word summary; everything else is verb-first fragments | Most of the document is never analyzed |
| OpenAI AI Text Classifier (withdrawn 2023) | 1,000+ characters, and still only 26% accurate on AI text | A summary section often under 500 characters | Below the floor, and the tool no longer exists |
| The seven commercial detectors tested in Patterns (2023) | No published length floor | Short, standardized, deliberately low-perplexity text | Flagged 61%+ of human non-native essays; a resume is more formulaic still |
And there is a deeper irony in that last row. Every piece of advice that makes a resume machine-readable pushes it toward exactly the profile these detectors read as artificial: standardized section headings, past-tense action verbs, terminology mirrored from the job description, no stylistic flourishes. A resume that scores well against an ATS-friendly format checklist is, almost by construction, low-perplexity text. We spent fifteen years telling candidates to write like a machine parses, and now we are surprised when it reads like one.
So what are the 49% of hiring managers actually noticing?
The statistic driving most of the anxiety comes from a January 2025 Resume.io survey of 3,000 hiring managers, which reported that 49% would automatically dismiss a resume they believed was AI-generated. A separate Resume Now report put the figure at 62% for AI-generated resumes submitted without personalization. Both numbers are worth taking seriously as a read on attitudes. Both also come from companies that sell resume services, the same structural conflict of interest behind the 75% ATS-rejection claim and the 6-second scan rule, so treat them as sentiment data rather than measurement.
The more useful part of the Resume.io data is not the headline percentage, it is the description of how respondents say they spot it: buzzword overload, vague claims with no numbers behind them, a voice that shifts between sections, and phrasing copied wholesale out of the job posting. Roughly a third said they could tell within about twenty seconds. Read that list again and notice what it describes. Not one of those signals is evidence of a language model. Every one of them is evidence of writing with nothing specific in it.
A 51-point spread that gives the game away
In the same survey, the share of hiring managers who said they would reject an AI-written resume outright ranged from 71% in Iowa down to about 20% in New Hampshire. If this were a detection capability, geography would be irrelevant. A 51-point gap between two states is a map of cultural attitudes toward AI, not a map of anybody's ability to identify machine-generated prose.
This gives us two testable predictions, and in my experience reading resumes both hold. A human who writes a bland, buzzword-stuffed resume with no numbers in it gets suspected of using AI. And an AI-assisted resume packed with specific systems, figures, timeframes, and named outcomes does not, because there is nothing generic left to trip the pattern-match.
The part that actually bothers me: the cost lands on the people it helps most
The strongest evidence about AI writing help in hiring is not a survey at all. Emma van Inwegen, Zanele Munyikwa, and John Horton ran a randomized field experiment on an online labor market with roughly half a million jobseekers, later published in Management Science and available as NBER working paper 30886. One group got algorithmic writing assistance on their resume profiles; the control group did not. The treated group was hired about 8% more often, earned 8.4% higher wages ($18.62 per hour versus $17.17), and, contrary to the worry that polished writing destroys a useful signal, the researchers found no evidence that employers ended up less satisfied with who they hired.
The detail that stops me every time I reread it: non-native English writers made up more than 80% of that sample, and the effect was largest for them. So set the two studies side by side. The population that gains the most from AI writing assistance is the same population that commercial detectors falsely accuse at rates above 61%. A hiring norm built on the folk belief that AI writing is detectable does not distribute its costs evenly. It concentrates them on people whose English is a second language and whose resumes were already being read with a heavier thumb on the scale, which is the same disparate-impact pattern I found when looking at age signals in resume screening.
What to do with your resume, given all of this
The practical advice does not change much whether you use AI or not, which is itself the tell that authorship was never the real variable. My colleague Marcus has written the recruiter's view on whether you should use AI to write your resume; here is the version that falls out of the detection data specifically.
- 1Run the specificity test on every bullet. Each one should contain at least one fact a language model could not have produced without your data: a number, a system or tool name, a customer segment, a timeframe, a named outcome. A model can polish that sentence, but it cannot invent it, and its presence is what makes the writing read as yours.
- 2Write the summary yourself, then let AI edit it. That is the only continuous prose on the page, the one section a human reads word for word, and the only part a detector would even attempt to score. Keep your cadence in it.
- 3Delete sentences you lifted from the job posting. Mirroring keywords is how tailoring is supposed to work. Mirroring whole sentences is the single most-cited tell in the survey data, and it reads as lazy whether a machine or a person did it.
- 4Never let a model generate a number. The real risk is not being detected, it is walking into an interview unable to defend a metric you did not produce. Quantify from your own records and nowhere else.
- 5Do a rhythm check before you send. Three-item lists in every bullet, identical sentence lengths throughout, and a pile-up of "spearheaded," "leveraged," and "utilized" are the surface patterns humans register as artificial. Vary the shape and the impression goes away.
This is the problem Resume Leap is built to solve from the other end: instead of generating impressive-sounding prose, it pulls the specific, quantified results already in your history and surfaces the ones a given job description is actually asking about.
Key takeaway
Nobody has a working AI detector pointed at your resume, and the published evidence says one is not currently buildable at resume length: OpenAI withdrew its own after a 26% catch rate, Turnitin will not score bullet points or anything under 300 words, and the detectors that do run flag more than 61% of human writing by non-native English speakers. What hiring managers detect is genericness, which is a solvable problem. Put a fact in every bullet that only you could know, and the question of who typed it stops mattering.