Estimated reading time at 200 wpm: 9 minutes
Over the past three months, a series of custom Python scripts were built to manage AI-generated text. One attempted to detect slop phrases. Another tried to enforce cleaner output. A third monitored register creep across longer documents. None of them performed well enough to be reliable. Pangram, a commercial AI detection service, continued to flag the output as AI-generated or AI-assisted regardless of the filtering and corrections applied.
Whether or not you agree our Fat Disclaimer applies
This prompted a different approach. The original AI-assisted material was set aside and used only as a structural guide. In effect this was a test. The same content was then written entirely by hand with no intention to avoid AI Slop patterns. Much of the underlying information remained identical. There was negligible difference in word choices and sentence rhythms, and none that the human ‘eye’ could detect.
The results were stark. Both Pangram and Originality.ai scored the rewritten text at 99 to 100 per cent human-generated. These were the same detection services had flagged the earlier drafts at over 60 per cent AI probability. The content had barely changed. Instead it was realised that the statistical fingerprint had changed entirely. But that remained imperceptible to human reading.
That amazing find led to a deeper investigation into how these services actually work. Both Pangram and Originality.ai are themselves driven by AI models trained to recognise the patterns left behind by other AI models (AI Content Detection in 2026: Trends to Watch, Wellows, July 2025). The remainder of this article examines the specific mechanisms they use, drawing on published research and technical documentation.
1. The statistical baseline of machine text
Perplexity and predictability
AI detection services do not read for meaning. They analyse text as a sequence of mathematical probabilities. The most important of these is perplexity, which measures how predictable a given sequence of words is to a language model (How Do AI Detectors Work? Techniques, Limitations & More, GPTZero, October 2024).
Large language models generate text by selecting the most statistically probable next token at each step. This produces writing that flows smoothly and reads coherently, yet it also produces writing that is highly predictable. Every word choice sits close to the top of the probability distribution. The resulting text has low perplexity because very little about it surprises the model.
Human writing behaves differently. A person might reach for an unusual synonym, construct a clause in an unexpected order, or use a colloquialism that no language model would rank as the most likely continuation. These choices raise the perplexity score. The text becomes harder for a model to predict, and that unpredictability is one of the strongest signals a detector looks for (GPTZero and the challenges of AI detection in assessing, L Giray, 2026).
Burstiness and rhythmic variation
The second core metric is burstiness, which measures how much the perplexity varies across the length of a document. In practice, this translates to variation in sentence length, structure, and complexity (GPTZero Review 2026: Accuracy, False Positives and What Teachers Get Wrong, AI Brand Factory, May 2026).
AI-generated text tends to maintain a steady, uniform rhythm. Sentences cluster around a similar word count. Grammatical structures repeat in predictable cycles. The output reads as polished and consistent, which is precisely the problem. Real human writing is irregular. A long, winding sentence full of subordinate clauses might be followed by a short declarative statement. A paragraph might open with a fragment and close with a question. This unevenness is a natural product of human thought and it creates the high burstiness scores that detectors associate with genuine authorship.
The hand-rewritten text from the experiment described above likely succeeded because it carried this natural irregularity. The AI-assisted draft, despite being factually identical, had the even cadence and low burstiness that mark machine output. The detector was not evaluating the ideas. It was measuring the rhythm.
2. The mechanics of adversarial training
AI detection models continuously updated to recognise the outputs of text humanisers. This process is known as adversarial training. Developers feed their detection algorithms millions of examples of AI text that have been processed by popular rewriting tools. The model learns to identify the specific statistical fingerprints left behind by these humanisers.
Training on humanised datasets
Humaniser tools attempt to bypass detection by artificially inflating perplexity and burstiness. They replace common words with obscure synonyms or awkwardly invert sentence structures. Detectors are trained on these exact transformations. They learn to recognise when a common word has been systematically swapped for a statistically improbable alternative. The detector searches specifically for the mathematical disguise applied to that output (RADAR: Robust AI-Text Detection via Adversarial Learning, ResearchGate, June 2026).
Identifying algorithmic transformation patterns
Humanisation software relies on a limited set of transformation rules. It might consistently replace passive voice with active voice or insert specific transitional phrases to break up long sentences. These repetitive rules create their own detectable pattern. An empirical study of bypass tools shows that detectors analyse the text at the token level to find these residual algorithmic signatures (How to Bypass AI Detection in 2026: Empirical Study of 6, GitHub Community, September 2026). A human writer does not follow a rigid, repetitive formula for varying their sentence structure.
3. Contextual coherence and semantic analysis
Advanced detectors evaluate more than just the probability of individual words. They assess the semantic relationships between entire paragraphs. This requires a deeper level of natural language processing.
Evaluating semantic relationships
Human writing features a coherent progression of ideas. The text builds logically from one concept to the next. Detection frameworks now combine textual features with claim-level semantic analysis (Semantic Level Detection of AI Generated Peer-Reviews, ICML, July 2026). The model checks if the meaning of the text remains consistent and if the arguments flow naturally. AI-generated text, especially when passed through a humaniser, often loses this deep contextual coherence. The sentences might make sense individually but lack the intuitive semantic flow of a human author.
The uncanny valley of synthetic variance
A humaniser can create a sentence that is statistically unpredictable. It might use a rare word or an unusual grammatical structure. This successfully spikes the perplexity metric. However, it often degrades the overall readability and logical connection of the paragraph. Independent evaluations show that detectors are trained to spot this synthetic variance (Evaluating the accuracy and reliability of AI content detectors, M Hadra, 2026). The text has the statistical variance of a human but lacks the authentic semantic grounding. The detector recognises this as the textual equivalent of an uncanny valley.
4. Token-level analysis and syntactic fingerprints
Detectors look beneath the surface of the words to examine the structural skeleton of the text. Even when a humaniser changes the vocabulary, the underlying grammatical architecture often remains identical to the original machine output.
Grammatical skeleton preservation
Language models generate text using specific syntactic templates. These templates dictate how clauses are nested and how phrases are ordered. Research into syntactic templates reveals that AI-generated text exhibits a highly regular, repetitive grammatical depth (How Syntactic Templates Reveal Patterns in AI-Generated Text, Complex Discovery, November 2024). A rewriting tool might swap a noun for a synonym, but it rarely alters the fundamental depth of the syntactic tree. Detectors map these trees to find the rigid, machine-like consistency that human writers naturally avoid.
Repetitive transformation rules
Human writers vary their syntax intuitively. They might follow a complex, multi-clause sentence with a simple subject and verb. Humanisation algorithms, however, apply repetitive transformation rules to the text. They might systematically convert passive constructions to active ones or insert transitional adverbs at fixed intervals. Token-level analysis shows exactly which words and structures create these detectable AI patterns (How AI Text Detectors Actually Work: Token-Level, Greg Bessoni, 2026). The detector identifies the mechanical repetition of these structural shifts, flagging the text as algorithmically processed rather than organically written.
5. Ensemble modelling and the limits of obfuscation
Modern detection platforms do not rely on a single algorithm to evaluate text. They deploy multiple models simultaneously to capture the full spectrum of machine-generated characteristics.
Ensemble detection architecture
An ensemble approach combines the outputs of several distinct detection models. One model might evaluate raw probability, while another checks for paraphrasing patterns, and a third analyses readability metrics. Originality.ai is widely noted for using this ensemble detection to combine multiple models into a single, comprehensive assessment (Originality.ai vs GPTZero vs Turnitin 2026, Free Academic Tools, May 2026). For a text to bypass the system, it must successfully evade every model in the ensemble at the same time. This multi-layered architecture makes simple obfuscation techniques highly ineffective.
The failure of mathematical disguise
The fundamental limitation of AI humanisers is that they attempt to solve a linguistic problem with a mathematical formula. They calculate the necessary adjustments to bypass a specific metric and apply those adjustments uniformly. Current detection technology has evolved to recognise this uniform application of mathematical disguise (AI Detection in 2026: What’s Changed & What’s Coming, UndetectedGPT, January 2026). Genuine human writing contains natural imperfections, intuitive leaps, and organic variations that no algorithmic formula can perfectly replicate. The detector ultimately succeeds because it measures the absence of mathematical calculation in the text itself.
Conclusion
AI text detection has evolved far beyond simple word frequency and pattern matching. Modern platforms evaluate semantic coherence and syntactic depth alongside token-level probabilities. They do not merely scan for key phrases. They assess the structural integrity and contextual flow of the entire document.
Custom Python scripts built to filter slop or prevent register creep apply rigid rules to a fluid linguistic challenge. These tools operate on fixed parameters and predictable transformations. Detection models are explicitly trained to recognise and penalise those specific algorithmic adjustments.
Attempting to outsmart these systems with secondary scripts is exactly what the detectors are ready for. The detectors measure the natural irregularity of human thought, a quality that cannot be engineered through code. The evidence suggests that producing text which passes modern detection requires the organic, unassisted process of actual human writing.
The evidence is that AI detectors are rapidly advancing in knowledge and methods being used. The message from all this is simple: Stop griping. Stop trying to defend yourself if your writing fails AI-detection tests. Knuckle down and write by true human effort.











