Prose Coach · Blog

By Prose Coach · July 26, 2026

Undetectable AI Writing: What Holds Up and What Doesn't

Undetectable is a claim about the future. It says a piece of text will keep clearing a scanner that its vendor retrains on a schedule you don't control.

Nobody can sell you that. What can be sold, and what's worth buying, is prose that doesn't read as machine-made to a person. That property is stable, because human readers don't ship model updates.

TL;DR: Detectors measure how predictable your word choices are and how much that predictability varies. Paraphrasers change the words and leave the structure, which is why they expire. Changing sentence structure, paragraph shape, and specificity holds up, because it removes the signal instead of masking it.

Illustration of a repeating wave flattening across a page with one orange break in the pattern, representing statistical regularity in machine-written text

The two things detectors measure

Perplexity is how surprising each word is given the words before it. Language models pick high-probability continuations by design, so their output scores low. A person writing about their own subject reaches for the precise term, the odd aside, the word that only makes sense if you know the context. That raises the number.

Burstiness is how much that predictability moves across a document. Human writing surges and settles. A long qualifying sentence, then four words. A dense paragraph, then a thin one. Models hold a steady level, paragraph after paragraph, and the steadiness is the signal.

Side-by-side explanation of perplexity and burstiness, the two statistical properties AI detectors measure.

Neither measures authorship. They measure statistical regularity, which correlates with machine generation and also with careful formal prose, non-native English phrasing, and any writer working from a strong template. That's the whole mechanism behind detectors flagging writing people actually wrote.

Why paraphrasers expire

The standard humanizer replaces words and reorders clauses. It leaves the shape alone.

Graphic showing that paraphrasing tools replace vocabulary while leaving the paragraph structure detectors measure intact.

Shape is the durable fingerprint. Paragraphs that all run four sentences. Openings that all state a claim. Closings that all restate the opening. Transitions announcing turns the sentences already make. Swap every noun for a rarer synonym and every one of those properties survives intact, which is why a text can clear one scanner and fail the next.

There's a second cost that gets less attention. Substitution damages the prose. Rare synonyms in place of ordinary words read as strained to a human editor, so you trade a scanner problem for a reader problem. Vocabulary variety on its own doesn't beat detection, and it makes the copy worse on the way to not working.

What actually changes the measurement

Four changes move both metrics, and they're the same changes that make writing better to read.

Restructure clauses, not vocabulary. Move the subordinate clause to the front. Split a compound sentence into two uneven ones. Let one sentence run long because the thought is genuinely long, then stop the next one early.

Vary paragraph length on purpose. A one-line paragraph is a legitimate move. So is a nine-line one. Uniform blocks are the most visible tell on the page and the easiest to fix.

Replace the general claim with the specific one. "Improves engagement" is predictable in every context. "Cut the reply time on support tickets from two days to four hours" is not, because the numbers came from somewhere the model couldn't guess.

Delete the connective scaffolding. Words like "Additionally" and "Furthermore" exist in the draft because the model was signposting, not because the argument needed a bridge. Transition filler is the cheapest thing on this list to remove.

Run those and the score moves for a reason that survives a retrain: the underlying regularity is gone, not disguised.

The detectors disagree with each other

Detector Primary audience Main signal Practical note
GPTZero Education Perplexity and burstiness Flags formal prose more often than most
Turnitin Academic institutions Linguistic pattern matching Delivered inside an LMS review process
Originality.ai SEO and content marketing Multi-model probability Updates its model more often than free tools
Copyleaks Academic and enterprise AI signals plus plagiarism Two separate scans sold together
ZeroGPT General, free tier Statistical analysis Least consistent of the group

The same paragraph can score very differently across these, which is the strongest evidence that a single number isn't a verdict. If you're checking your own work, run two tools that use different approaches and treat wide disagreement as information about the tools. A fuller breakdown lives in the detector accuracy comparison.

Where volume changes the math

Everything above is manual work, and manual work has a ceiling.

At blog length you can read the draft aloud and fix the rhythm by ear. At book length you can't. Writers producing serialized fiction or long manuscripts find that the revision pass costs more time than the drafting pass saved, which removes the reason for using a model in the first place. That constraint is what AI novel writing platforms are designed around, and it's the same reason the useful intervention at volume happens during generation rather than after it. A rule the model follows while drafting scales. A checklist you apply by hand does not.

What Prose Coach does about it

Prose Coach installs inside ChatGPT, Claude, or Cursor as a writing skill. It works on structure: it scans a draft for the patterns above, points at the specific sentence producing each one, and applies revision rules during generation so fewer of them appear in the first place.

It's deterministic. There's no model guessing whether your text is machine-written, because that isn't the job. The job is naming the flat rhythm, the repeated paragraph shape, and the vocabulary reaching for the safest option, so you can decide what to do about them.

Free tier at $0, PRO as a one-time $39 purchase. What it won't do is guarantee a detector result, and any product that offers you one is describing a promise it has no way to keep.

Can AI writing be made permanently undetectable?

No. Detectors retrain, so any claim of permanence is a prediction about someone else's release schedule. Structural revision holds up better than word substitution because it removes the signal rather than masking it.

What is the difference between perplexity and burstiness?

Perplexity measures how predictable individual word choices are. Burstiness measures how much that predictability varies across a document. Machine-written text usually scores low on both.

Do humanizer tools work?

They move scores in the short term by substituting vocabulary and reordering clauses. They leave paragraph shape and rhythm untouched, so results decay as detectors update, and the substitutions often read as strained to a human editor.

Should I trust a single AI detection score?

No. The same passage frequently scores differently across tools, and passages under roughly 150 words produce unreliable results everywhere. Treat a score as one data point.

Is using an AI humanizer legitimate?

It depends entirely on the disclosure rules that apply to you. Refining your own work with AI assistance is different from presenting fully generated text as your own where that's prohibited, and no tool resolves that question for you.