In 2023, seventy-two applied linguists, people who review journal submissions for a living, were shown four abstracts and asked which had been written by ChatGPT. They got it right 38.9% of the time. Not one of them got all four. Thirteen percent got all four wrong. A coin would have beaten them.
That number deserves a second look, because worse than chance is informative. Random guessing gives you fifty percent. Scoring below it means the reviewers were using a theory, and the theory was pointing the wrong way.
That is the first half of the story. The second half matters more and gets less attention: whether a reader believes a text came from a machine changes how they judge it, even when the text is word for word the same. Once you hold both halves, most of the advice circulating about “how to sound human” turns out to be advice for the wrong problem.
Detection
How bad is it, really
Bad, and getting worse as models improve, which is not the direction anyone hoped for.
The largest study is Jakesch, Hancock and Naaman’s in PNAS (2023): 4,600 people judging 7,600 short self-descriptions, the kind you write for a dating site, a hosting profile, or a job platform. Accuracy came out between 50 and 52 percent. Paying people for correct answers did nothing. Telling them immediately whether they were right did nothing. No age group, education level, or degree of tech familiarity did better than the rest.
Earlier work on GPT-3 (Clark et al., 2021) found the same coin-flip baseline, and found that training evaluators with guidelines and worked examples moved them to about 55 percent. Five points for a training session.
Then the models got better. In Jones and Bergen’s three-party Turing test (FAccT 2025), an interrogator talks to a human and a model at the same time and has to say which is which. GPT-4 was picked as the human 54% of the time; the actual humans, 67%. A year later, GPT-4.5 with a persona prompt was picked as the human 73% of the time, more often than the person sitting in the other chat window. ELIZA, the 1960s pattern-matcher, scored 23%, so the interrogators weren’t guessing blindly. They had a model of what a human sounds like. GPT-4.5 fit it better than humans do.
Poetry is the case that should bother anyone who writes. Porter and Machery (2024) showed readers poems by well-known poets alongside AI imitations. Readers identified the AI poems 46.6% of the time, below chance again, and they preferred them: higher on rhythm, higher on beauty. The explanation in the paper is simple and a little grim. The AI poems are easier. Real poetry is dense and strange, and readers took the strangeness for machine incoherence.
Other media follow the same curve. A meta-analysis of 56 studies and 86,155 participants (Diel et al., 2024) puts human deepfake detection at 55.5% overall: audio 62%, video 57%, images 53%, text 52%. Audio is the one place people keep some edge, and newer voice clones are eroding it. Machines now beat humans on faked images (up to 97%); humans still beat machines on faked video, roughly two-thirds of the time, because we notice when motion is slightly wrong. That won’t last either.

Folk theory
The checklist in your head is wrong
If accuracy is at chance you might expect the judgments to be scattered at random. They aren’t. In the Jakesch data, raters agreed with each other about which texts felt synthetic. The agreement was slight (a Fleiss’ kappa of 0.07) but with that many judgments it was nowhere near random. One of the sources I drew on calls this “high consensus,” which oversells a kappa of 0.07; the honest description is faint but real. Either way the implication holds. People share a theory of what machine writing looks like. The theory is just wrong.
Here is what the theory said in 2022, and what the data said back:
| What raters used | What they took it to mean | What was actually true |
|---|---|---|
| Typos and grammar errors | Machine glitch. Raters were about 5% more likely to call flawed text AI. | AI text had fewer errors than human text (roughly 15% vs 20% of samples). An error was evidence of a person. |
| “I”, “me”, “my” | A self. | Models use them at the same rate humans do. |
| Stories about family | A life. | Produced on request; matched or exceeded human rates. |
| Contractions, casual phrasing | Spontaneity. | Models learned from conversational text. They contract. |
| Rare or long words | A database showing off. | Humans used rare word pairs more often. Standard decoding trims the improbable tail, which is exactly where unusual words live. |
| Repetition | A loop. | Correct. AI text really was more repetitive, and raters who leaned on this did better. |
| Sentences that don’t quite mean anything | Lost context. | Correct. Models produced more of these. |
Two of seven cues were real. The other five were useless or pointed backward, and the backward ones were the ones people trusted most.
The researchers then ran what I think is the single most important experiment in this area. They took the folk theory and handed it to the model as instructions: be casual, use “I”, mention your family, contract your verbs. The resulting profiles were judged human 65.7% of the time. Profiles written by actual people about their actual lives: 51.7%. On professional profiles the optimized machine text hit 71%. If you know what people think a human sounds like, you can out-human them.
The flip side: when raters were restricted to the two cues that work (repetition, nonsense), accuracy rose to 58.8%. Not good. Better than the intuition it replaced.
The checklist moved
I had the grammar row backwards in an earlier draft of this piece. I’d written that people read typos as human. They do now. They didn’t in 2022, when the mental model of AI was still “computers make mistakes.” The correction matters beyond the row itself, because it shows the folk theory isn’t fixed. It tracks the last machine you were annoyed by.
In 2022: errors mean machine. Wrong; machines made fewer.
In 2026: smoothness means machine. Em dashes, balanced clauses, tidy structure, an Oxford comma, the word “delve.” Pew’s Data Labs (Bestvater et al., August 2026) measured the shift on the open web: em dashes about twice as frequent since ChatGPT, Oxford commas up 63%, “delve,” “interplay” and “testament” more than doubled, something like 35% of new pages showing signs of machine involvement. So the new theory is not crazy. At the level of the whole internet, those features really did get more common.
It’s still useless for judging one page. The em dash predates the transistor. Plenty of people write balanced sentences on purpose; it’s called editing. What the 2026 checklist does in practice is punish anyone who writes carefully. NPR has already covered the writers organizing to keep their dashes, and the same self-censorship has been reported for semicolons and for “whilst.” A theory that started as “look for the mistakes” has become “look for the absence of mistakes,” and both versions fail for the same reason: they describe an average, then get applied to an individual.
Mechanism
What you’re actually reacting to
Ask people why a passage felt synthetic and they’ll talk about content. Watch what they react to and it’s mostly rhythm and manners.
Rhythm first. A language model picks each next word from a probability distribution, and the standard way of doing that (nucleus sampling, top-p) deliberately cuts off the improbable tail. The upside is fewer howlers. The downside is text with no surprise in it, sentence after sentence of the most likely thing. Combine that with similar sentence lengths and symmetrical clauses and you get prose that reads as metronomic even when the vocabulary is high. Human writing lurches. A clause piles up, then a fragment. Then a long reconsidering sentence with a parenthesis in the middle of it (like this one), because the writer changed their mind halfway through and didn’t go back.
Manners second, and I think this is the bigger one. Models trained with human feedback are rewarded for seeming helpful, so they over-fulfil. They explain the premise you already know, add a summary you didn’t ask for, close with a reassurance. Grice would say they violate the maxim of quantity; the reader’s translation is “this thing is performing helpfulness at me.” The same training produces a flat affability: no irritation, no irony, no “actually, I disagree.” Humans hedge, but they also commit, and they occasionally get bored of you. A text with no visible temperament reads as a customer-service script.
Then there’s the paradox that the too-perfect theory gets half right. Human writing carries the marks of having been made: a shift in tone where the writer’s mood changed, a sentence that strains because the idea was hard, an argument that doesn’t quite close. When a text has none of that, when every paragraph is the same size and every claim has a counter-claim and the conclusion is upbeat, readers don’t find fault with it. They find it suspiciously complete.
Underneath all of this sits an asymmetry that behavioural economists call algorithm aversion (Dietvorst, Simmons and Massey, 2015). A person’s mistake gets read as tiredness; a machine’s, as a design flaw. So a machine is punished for errors and a human is punished for having none. Those two rules together make a trap: the better you write, the more you look like the thing that can’t make mistakes.
The experts
Who can actually tell
Almost nobody, with one real exception.
Russell, Karpinska and Iyyer (ACL 2025) recruited five people who use ChatGPT heavily for their own writing and had them classify articles. Individually they caught about 93% of the machine text with a false-positive rate around 3 to 4%. As a majority vote they misclassified one article in three hundred. That beat nearly every commercial detector, and it held up when the AI text was paraphrased or run through “humanizer” tools: the humans still caught 100% of the humanized samples, where two well-known detectors caught 7% and 23%.
Nobody trained these people. They had read enough machine prose to know it, the way you know a colleague’s emails without the signature. What separates them from novices is instructive: novices flag any fancy word; experts flag the specific words machines over-use. It’s the difference between “this sounds smart” and “this sounds like it.”
Everyone else is at chance. Age, gender, education: small effects or none, across student, teacher and general-population samples. Younger people are more comfortable with machine text and women are more skeptical of it, and neither group is better at spotting it. The one cross-cultural detection study I know of predicted an East-West gap and didn’t find one.
There’s a group worth flagging even though the data are thin. Non-native English writers, and writers whose natural register is precise and unpadded (one of my sources names autistic writers specifically), produce text that sits exactly where the 2026 checklist points: controlled vocabulary, regular syntax, no social filler. For non-native writers the harm is documented, below. For neurodivergent writers the mechanism has been argued convincingly and the numbers not at all. I’d be surprised if they weren’t being flagged. Someone should measure it.
The label does more than the text
Everything above is about whether readers can tell. Here is the part that should reorganize how anyone thinks about this: it barely matters whether they can tell, because what they believe does most of the work.
Raj, Berg and Seamans ran sixteen preregistered experiments with 27,491 participants (Journal of Experimental Psychology: General). Same creative writing; one group told a person wrote it, the other told an AI was involved. The AI-labelled group rated it about 6% lower, and the drop ran through perceived authenticity. Then the authors tried to make it go away. Framing the AI as a tool. Framing it as a collaborator. Giving it a personality. Telling participants about machine emotional capacities. Nothing worked. The penalty lands where the mind-perception literature would predict: emotional, first-person writing takes the hit; technical description mostly doesn’t.
Schilke and Reimann (OBHDP, thirteen experiments) found the same shape in professional settings and supplied a mechanism: people who disclose AI use are trusted less because the use is seen as illegitimate, socially wrong for the role. The effects are not small. Professors known to grade with AI assistance, d around 0.8. Founders, 0.93. Disclosure hurt whether it was voluntary or required. The only thing that reliably shrank it was making AI use feel normal (d fell from 1.12 to 0.72 under that priming), and even then it stayed large.
Yin, Jia and Wakslak (PNAS 2024) is the study that’s hardest to dismiss. They had an AI write emotional-support messages and compared them with messages from ordinary, untrained people. Recipients felt more heard by the AI, mostly because it didn’t rush to give advice. Then they were told. The “less heard” penalty after the reveal roughly cancelled the gain. The message was the same. The feeling changed.
News runs the same way (Toff and Simon 2025; Altay and Gilardi 2024). AI-labelled stories are trusted less even when readers rate them just as accurate and just as fair, and the drop is largest among the people who trust news most. The Reuters Institute’s 2025 figures across six countries lay out the gradient: 12% comfortable with news made entirely by AI, 21% with a human somewhere in the loop, 43% with a human leading, 62% with no AI at all.
And there’s a finding from before any of this that explains the current mood better than anything since. In 2019 Jakesch and colleagues showed people host profiles and told some of them the profiles might be machine-written. Trust dropped for the profiles readers suspected, and it dropped specifically in the mixed condition, when nobody knew which was which. They called it the Replicant Effect. You don’t need a machine in the room to get it. You need the possibility of one. That’s the internet in 2026: every long, well-organised argument now has to clear a suspicion before it can be read, and a lot of readers are scanning for the excuse not to.
One interpretive note. Some researchers (Bellaiche and others) argue this isn’t animus toward machines so much as a premium on humans: people value the effort, the risk, the fact that someone spent an afternoon on this. A small study (N=70, so hold it loosely) found 73% of participants would pay more for human-made work, and that showing the process (a time-lapse, a stack of drafts, a note about what was hard) raised perceived authenticity, while visible imperfections on their own did little. I find that convincing, and it’s the hinge for the advice below.
The people getting hurt
The detector industry has produced a specific harm, and it lands on the least powerful writers.
Liang and colleagues (Patterns, 2023) ran seven commercial detectors on essays by American eighth-graders and on TOEFL essays by non-native speakers. The eighth-graders sailed through. The TOEFL essays were flagged as machine-written 61.3% of the time on average; 97.8% were flagged by at least one tool. The reason is the same low-perplexity, regular-syntax profile the tools were built to find. When the authors had a language model enrich the vocabulary of the same essays, the false-positive rate fell to 11.6%. Read that again. The fix for being wrongly accused of using AI was to use AI.
A 2026 paper argues, from first principles, that this can’t be engineered away. Any text-only detector powerful enough to catch models will misfire wherever human style overlaps model style, and human style overlaps everywhere, because the models learned it from us. The trade press meanwhile reports a “polish paradox”: the most disciplined writers are the most likely to be flagged. Even the five expert readers above, the best detectors on record, wrongly accused a human 3 to 4% of the time. Scale that across a semester of submissions.
The institutional response has been to reverse the burden of proof. Students keep version histories and screen recordings. Scholars are asked to explain why their transitions are smooth. Writers describe removing the em dash, then the semicolon, then the long sentence, until the prose is as plain as the accusation demands.
Signals
What actually reads as human
If competence is what machines have, humanness is whatever is left over, and the mind-perception literature (Gray, Gray and Wegner, 2007) says what that is. People grant machines agency: planning, reasoning, self-control. They withhold experience: sensation, vulnerability, wanting things, being somewhere in particular. Fluency signals agency. It buys nothing on the other axis. What does:
Being somewhere. A room. The specific street. Models average over corpora; they aren’t standing anywhere. A physical detail that could only be known by someone who was present is the hardest thing to fake and the easiest thing to recognise.
Saying something you could be wrong about. Models default to the middle of the road and a disclaimer. A committed claim, an unpopular one, one that costs the writer something if it’s wrong, reads as a person with a stake.
Leaving the seams in. Not typos. The trace of a mind changing: a counter-argument taken seriously, an admission of ambivalence, a paragraph that doesn’t tie off because the thought didn’t finish. The correction I made above about the grammar row is doing this work. The HCI studies on typo correction found the same thing at the smallest scale: visibly fixing a mistake read as more human than never making one. It’s the fixing that signals, not the flaw.
A temperament. Humor, irritation, a preference, an address to the reader as a particular person rather than an audience.
And process. Drafts and dates. A note on what was hard. That’s the one item here with direct experimental support, in the struggle-premium study, and it doubles as protection: version history is the only thing that has reliably gotten a falsely accused student off.
Notice what isn’t on the list. Fake typos. Faux-casual phrasing. Exactly the things the 2022 folk theory rewarded and the 2023 experiment exploited. Anyone can prompt for them.
Practice
What I’d actually do
For writers. Write about specific things in specific places. Take a position and pay for it. Cut the both-sidesing, the summary paragraph, the upbeat close; end on the claim. Keep your drafts. Don’t manufacture errors, and be careful about manufacturing rhythm: varying your sentence length is what thinking does on its own, but done as a technique it’s a surface feature, the exact one the humanizer tools inject and the detectors chase, and you’ll be in an arms race you can’t win. If you’re a non-native writer or a naturally precise one, and you’re anywhere that runs detectors, save everything.
For anyone evaluating. Stop using detector scores and gut feelings as evidence. Not “use them carefully.” Stop. The false-positive rates are high, they’re concentrated on identifiable groups, and there’s a decent argument they can’t be fixed. Assess the process instead: drafts, a conversation about the work, an oral defense. If you want to sharpen your own judgment, the one method with evidence is repeated practice with feedback on each guess (Milička et al., 2025, about 255 people), which improved accuracy some and calibration a lot; without feedback, people were most wrong when they were most sure. Jakesch found feedback didn’t help, and I don’t think the two studies are fully reconciled. My guess is that feedback fixes overconfidence rather than perception, which would make it worth doing anyway. Be most suspicious of yourself when you feel certain.
On disclosure. Saying that AI was used costs trust, reliably, in every study above. Saying how the work was made and checked (which sources, what a human verified) recovers some of it. If you write policy, don’t mandate the first without the second, and expect the penalty to shrink slowly, if at all, as the practice becomes ordinary.
For institutions:
| Where | What’s happening now | What would be better |
|---|---|---|
| Classrooms | Assignments run through detectors; students presumed guilty | Version history, scaffolded drafts, oral defense; audit the detector on known-human work before trusting it |
| Peer review | Rejection on “reads like AI” | Rubrics limited to method, data and argument |
| Workplaces | Bans, or quiet shame | A stated policy on what assistance is fine, with the human still accountable for every claim |
| Public literacy | Tell-tale-sign checklists | Teaching that the checklists don’t work and the accusations do damage |
| Platforms | Bare “made with AI” labels | Labels that say what was done and who checked it |
Where I’m still not sure
Whether the disclosure penalty erodes as machine writing becomes ordinary. The collective-validity result says a little. Schilke and Reimann found experienced AI users penalised just as hard, while Yin’s data suggest people with warmer attitudes soften. I lean toward “it will shrink for functional writing and persist for anything that claims to be felt,” which is roughly what mind perception predicts, but that’s a prediction, not a finding.
Whether any surface signal survives. Every feature that reassures a reader can be prompted for. The experts in the Russell study weren’t fooled by humanization, which suggests they’re reading something deeper than features, and I can’t say what.
Whether the right question is even “is it AI.” For essays, for poems, for a message from someone you love, provenance matters and I don’t think that can be argued away. For a product manual, a meeting summary, a page of documentation, it’s hard to see why it should, and the energy spent policing it would be better spent asking whether the thing is true and useful. That’s an easier position to hold about a manual than about a eulogy.
Sources
What this leans on
- Jakesch, Hancock & Naaman, PNAS 2023
- Jakesch, French, Ma, Hancock & Naaman, CHI 2019
- Clark et al., ACL 2021
- Jones & Bergen, FAccT 2025 and PNAS 2025
- Casal & Kessler, Research Methods in Applied Linguistics 2023
- Porter & Machery, Scientific Reports 2024
- Diel et al., Computers in Human Behavior Reports 2024
- Russell, Karpinska & Iyyer, ACL 2025
- Liang et al., Patterns 2023
- Raj, Berg & Seamans, J. Experimental Psychology: General
- Schilke & Reimann, OBHDP
- Yin, Jia & Wakslak, PNAS 2024
- Toff & Simon, Int’l J. Press/Politics 2025
- Altay & Gilardi, PNAS Nexus 2024
- Reuters Institute, Digital News Report 2025
- Pew Research Center Data Labs (Bestvater et al.), 2026
- Milička et al., 2025
- Dietvorst, Simmons & Massey, JEP: General 2015
- Gray, Gray & Wegner, Science 2007
- Stein & Ohler, Cognition 2017
- Bellaiche et al.; “Struggle premium,” arXiv 2026 (N=70)
- Frank et al., cross-country detection study
- NPR, November 2025, on the em dash.
Disclosure
In case you were wondering
Whether an agent drafted, reviewed or edited this: of course one did. The more useful question is what it did. It did not write the argument or pick the sources. It found papers I then read, checked numbers against the originals, and argued with me about the grammar row, which I had backwards until it pushed. Every figure here I verified myself. The sentence you are reading now is mine, and so is the one that was wrong.



