How to Spot AI Written Content — And Why Detector Tools Keep Getting It Wrong

How to Spot AI Written Content

Why generated prose has a recognisable rhythm, what perplexity and burstiness actually measure, why detectors fail, and what search guidance really says.

Spotting_AI_written_content

Last Updated: August 2026

🔴 Why Generated Prose Has a Shape

A language model does one thing repeatedly: given everything written so far, it produces a probability for every word that could come next, then picks one. The properties people notice all follow from that loop.

The safe word wins

At each step there is usually one obvious continuation and a long tail of unlikely ones. A model set to produce coherent, agreeable text takes the obvious one most of the time. That is what makes the output readable, and it is also what makes it flat.

Human writers make locally strange choices constantly. We reach for the odd word because it sounds better, break a sentence in half because we ran out of breath, use a fragment for emphasis. Like that. Those choices are exactly the low-probability continuations the model is least likely to pick.

Alignment training flattens it further

Models are tuned on human ratings of what makes a good response. Raters reward answers that are balanced, hedged, structured and inoffensive. Repeated across millions of comparisons, that produces a house style: the qualified opening, the tidy three-item list, the summary paragraph restating what was just said.

This is why generated text from different vendors reads similarly despite different training data. They were shaped by similar preferences about what a helpful answer looks like.

The result is measurable rhythm

Take any page and list its sentence lengths. Human writing scatters — a nine word sentence next to a thirty-four word one next to four words. Generated text clusters, typically between fifteen and twenty-five words, sentence after sentence.

You can put a number on that. Divide the standard deviation of sentence length by the mean and you get a coefficient of variation. Above roughly 0.5 reads naturally. Below 0.35 the prose feels mechanical even to a reader who could not explain why. This is the single most visible difference, and unlike a detector score it is something you can check and fix yourself.

🟡 What Detectors Actually Measure

Scatter plot showing human and generated text overlapping on a perplexity scale

Almost every detector rests on two statistics, and understanding them explains the failures.

Perplexity

Perplexity asks how surprised a reference model is by each word. If the text keeps choosing exactly what the model would have chosen, perplexity is low. Human writing tends to be higher because we make unusual choices.

The flaw is immediate. Low perplexity does not mean machine-written; it means predictable. A legal disclaimer, a recipe, a product specification, an instruction manual — all predictable by nature, all written by people, all scoring as generated.

Burstiness

Burstiness is the variation described above, applied across a document. Human writing is bursty: dense paragraphs beside short ones, complex sentences beside blunt ones. Generated text is smoother.

The same flaw applies. Technical documentation is deliberately uniform because uniformity aids comprehension. Consistency is a virtue there, and the detector reads that virtue as evidence of a machine.

Who this hurts

Studies of detector behaviour have found a consistent bias against writing by non-native English speakers. A writer with a smaller working vocabulary and simpler sentence construction produces exactly the low-perplexity, low-burstiness signature the detector was built to catch.

In an academic setting that becomes an accusation. In an editorial setting it means someone rewrites perfectly good work to appease a number. Neither outcome is acceptable, and both follow from treating a probabilistic signal as a verdict.

Watermarking, and why it is not the answer yet

A more promising approach embeds a statistical signature during generation — biasing word choice imperceptibly so a detector holding the key can recognise it. It works well in controlled tests and survives light editing.

Two problems keep it from mattering in practice. It requires the generating model to cooperate, which open-weight models running on someone’s own hardware never will. And heavier paraphrasing removes it. Watermarking may become useful for verifying that something is from a particular source; it will not tell you that arbitrary text is not machine-written.

🟢 What Search Guidance Actually Says

This is where the common assumption is simply wrong.

What actually gets penalised

🔵 Pages that answer nothing. Five hundred words that restate the question and never resolve it.

🟠 Pages with no first-hand knowledge. A synthesis of other pages, adding nothing the sources did not already contain.

🟣 Pages written for a keyword. Where the topic was chosen by a search volume figure rather than because anyone had something to say.

🔵 Pages at scale with no editorial pass. Volume without a human deciding whether each one was worth publishing.

A machine can produce all four quickly, which is why they correlate with generated content. But a person can produce all four too, and plenty have. The failure is thin content, not the tool that made it.

The practical implication

If you are checking your own pages, “does this read as machine-written” is the wrong question. The right one is “does this page contain anything a reader could not get from the first three results”. A real number, a limitation you found the hard way, a case where the obvious approach failed. That is what no model can supply, because it did not do the work.

🔴 A Checklist That Works

Rather than a score, six things worth reading for:

🔵 Is there a specific number anywhere? Not “significantly faster” — nine seconds instead of forty.

🟠 Does it admit a limitation? Generated marketing copy rarely says what a thing cannot do. Real experience always knows.

🟣 Is there a first-hand detail? A moment that could only come from having done it, and could not be inferred from a specification.

🔵 Do the sentences vary? Read three paragraphs aloud. If the rhythm never changes, that is the tell.

🟠 Does it hedge everything? “May potentially help in some cases” is a sentence that commits to nothing.

🟣 Would deleting a paragraph lose anything? If not, it was padding whoever wrote it.

None of these gives a percentage, and that is the point. They give you something to change. The Content Forensics Studio automates the countable ones and names every instance, so a page can be fixed rather than merely scored. For the wider argument about why these checks run in the browser instead of on a server, there is a longer piece on secure offline web development utilities, and the general background is covered in Wikipedia’s article on large language models.

❓ Frequently Asked Questions

Are AI content detectors accurate?

Not reliably. OpenAI withdrew its own after measuring roughly a quarter accuracy. Independent testing finds high false-positive rates, especially on non-native English writing.

What is perplexity?

A measure of how predictable text is to a reference model. Low perplexity means predictable, which is not the same as machine-written — recipes and legal text score low too.

What is burstiness?

How much sentence length and complexity vary across a document. Human writing scatters; generated text clusters. Technical documentation is deliberately uniform and gets misread.

Does Google penalise AI-written content?

Not for being AI-written. Automation aimed at manipulating rankings is against the guidelines; automation used to produce genuinely helpful pages is not, in itself, a violation.

Why does generated text all sound similar?

Models pick high-probability words, and alignment training rewards balanced, hedged, tidily structured answers. Both push different systems toward the same house style.

What is the most visible tell?

Sentence rhythm. Human writing varies length constantly; generated prose settles near one comfortable length. Reading three paragraphs aloud usually reveals it.

Can watermarking solve this?

Partly. It needs the generating model to cooperate, which open-weight models will not, and paraphrasing removes it. Useful for proving origin, not for proving absence.

Should I run my own pages through a detector?

Not for a percentage. Check instead for specifics, admitted limitations and varied rhythm — things you can act on rather than a number you cannot verify.

What actually makes a page thin?

Containing nothing a reader could not get from the first three results. No number, no first-hand detail, no honest limitation. A person can write that just as easily as a machine.

Choose a language

Top Tools Ranking

Network Total Views
14,348
Tracking Since
Jul 9, 2026

Click any tool to open in a new window