Coming soon! The Kael'Nyrin Scrolls: The Atlas Edict

Diagram explaining four levels of register creep in text

Captain Walker

Register Creep: Beyond Slop to the Four Levels of Text

AI models, creep, register, sentences, slop, training, words

Estimated reading time at 200 wpm: 36 minutes

I had been working on slop. In previous work I found eleven patterns. Words, lists and stock phrases were caught by scripts. But something kept getting through. Text with no slop words still read like something delivered from a lectern, or as a report when none of that was required.

Whether or not you agree our Fat Disclaimer applies

This post will be of importance for those who are training on how to assess AI responses. Flowery language and sentences structures have a way of ‘convincing’ people about the value of outputs. Underneath all that the construction may be pretty bare.

AI assistance will be of importance for the future and so too will the work of avoiding or cleaning up various patterns of text. Those patterns are almost hard-coded into AI models because of the data they were trained on i.e. sloppy human writing.

An exploration of the phenomena with one AI let to the concept of register creep. A follow up discussion with Claude sorted it into something I could use. This dissertation records that. It is for my purposes.

In tight summary: Slop is a fault in vocabulary, and patterns in use of words. Register creep is a fault in construction and in voice. A de-slop script cannot pick up register creep because it is not designed to do that.

What follows are my discoveries, how register creep works, and strategies for prevention. As a favourable side-effect slop words were reduced but still required the de-slop skill to be enforced in OpenWebUI. Users of standard web-based interfaces will not be able to address these issues without extreme effort.

1. The four levels of text

Text can go wrong at four levels. Each one sits above the last. Each one is harder to detect than the last.

1.1 Level 1: Lexical, which words

The choice of individual words and set phrases. “Delve”, “tapestry”, “testament to”, “it is important to note”. This is where most slop discussion lives. It is countable. You can point at the word. You can swap it out and the sentence is repaired.

1.2 Level 2: Syntactic, the shape of the sentence

How the sentence is built. Where the main clause sits. Whether the point is delivered at once or held back. Whether there is a contrast frame, a cleft, a runway before the content. Every word can be ordinary and the sentence can still perform. The shape does the work. Swapping words does nothing. The sentence has to be rewritten.

1.3 Level 3: Discourse, the shape of the paragraph

How sentences are arranged into a paragraph. Whether the paragraph builds to something. Whether it ends on a moral. Whether events are ordered for effect rather than for the record. Each sentence can be clean and the paragraph can still have the shape of an essay or a speech. The fix is cutting or reordering, not rewording.

1.4 Level 4: Stance, who is speaking and what they claim to know

The position the writer takes. A reporter states what was found. A narrator knows what people felt. A speaker tells the audience what to think. When a report drifts into narrator or speaker, the text starts asserting things the document cannot support. Beliefs nobody checked. Conclusions nobody drew. This is the most serious level. The fix is deletion. The claim should not be there.

2. An example of how it works, or rather not

Start with the sentence a report wants:

The statement was checked against the CCTV footage and did not match.

Level 1. Same shape, two words swapped for slop:

The statement was delved into against the CCTV footage and did not align.

A de-slop word list skill catches this. Take the words out and it is fine.

Level 2. No slop words. Sentence rebuilt:

Far from matching the CCTV footage, the statement, once checked, did not hold.

Every word is ordinary. The point is held back. There is a contrast frame setting up a small reveal. It is built to be spoken. A word list sees nothing.

Level 3. Two clean sentences and a coda:

The statement was checked against the CCTV footage. It did not match. And that, ultimately, told its own story.

The third sentence says nothing. It gives the paragraph a shape: setup, fact, meaning. A report does not need that shape.

Level 4. The writer becomes a narrator:

He was under no illusion that the statement would survive the CCTV footage. It did not.

The writer and reader are told what to believe and without substance. It may sound good and convincing but the sentences acquired a claim it cannot back.

Put together:

LevelFault inFix
1VocabularySubstitute
2Sentence constructionRewrite
3Paragraph constructionCut or reorder
4Writer’s positionDelete the claim

My scripts worked at level 1. What I was seeing was happening at 2, 3 and 4. That is why it did not look like slop but it was sloppy.

3. Why slop screens miss levels 2 to 4

A slop screen works by matching. It holds a list of words and phrases. It counts how often they appear. It reports density and clumping. This is the right tool for level 1. The fault is in the vocabulary and the vocabulary is on the surface.

Levels 2 to 4 are not on the surface in the same way.

A level 2 fault is a property of sentence architecture. The words are ordinary. What is wrong is their order and the frame they sit in. A matcher can catch some of the frames, because some frames have fixed openings. “Far from”, “not X but Y”, “it is against this backdrop”. But the frame can be built from any words at all. The list can never be complete.

A level 3 fault is a property of the paragraph, not of any sentence in it. Each sentence passes. The paragraph fails. A matcher reads one sentence at a time. It has no view of the whole.

A level 4 fault is a property of the relationship between the text and the world. “He was under no illusion” is a perfectly formed sentence. It is wrong because no evidence of his belief was cited and the document is a report. Knowing that requires knowing what kind of document this is, who “he” is, and what has been established. That is reading, not matching.

So the screens did their job. The job was smaller than the problem.

There is a second reason the fault escaped notice. Slop is ugly. Register creep is often fluent. The prose reads well. It has rhythm. It sounds confident. The reader is carried along. The fault hides behind competence.

4. Register creep: the umbrella term

Register creep is the drift of a text out of the register the document type demands and into an adjacent one. A report drifts towards the essay, the speech, the case study, the thriller. Nothing is added by way of content. The voice changes.

The drift has three recognisable forms. They often appear together. They are worth separating because they are detected differently and fixed differently.

4.1 Narrative gravity: people acquire interiority and stakes

The people in the document stop being subjects of a record and become characters. They notice, realise, understand, wrestle with. Their actions acquire motive. Their situations acquire tension. Events are given weight they did not have.

This is the form that does the most damage in a report. It converts a description of what happened into a story about what it was like. A story requires an inner life. The document supplies one. Nobody verified it.

4.2 Declamation: prose acquires cadence and an audience

The sentences begin to sound as if they are being delivered. Rhythm appears. Long sentence, short sentence. Three-part lists. Openings that build rather than state. Closings that land rather than stop.

This form is the least harmful in itself. The facts may be intact. But declamation carries a stance. Prose built to persuade tells the reader what to feel about the facts. A report should not do that.

4.3 Inflation: facts acquire gravitas

Small things are treated as large. A routine check becomes a critical juncture. A missing signature becomes a question that cannot be ignored. A procedural step becomes an insight.

This form is easy to miss because it is often true in a trivial sense. The check did matter. The signature was missing. What is wrong is the scale. The text has lost its sense of proportion.

4.4 How the three relate

Narrative gravity is a fault of stance. Declamation is a fault of syntax and discourse. Inflation is a fault of judgement about scale. One text can carry all three. A sentence like “He was under no illusion that this small discrepancy would prove decisive” has interiority (he was under no illusion), a contrast frame (small but decisive), and inflation (a discrepancy proves decisive). Three forms in fourteen words.

5. Why this is epistemic, not cosmetic

Register creep is usually treated as a style problem. The prose sounds wrong. It reads like a speech. It feels overblown. The instinct is to file it alongside slop: an aesthetic irritant, a sign of careless editing, a polish issue.

That understates it. Register creep does not just change how things sound. It changes what the text claims.

5.1 Unverifiable claims about mental states

When narrative gravity enters a report, mental states appear. “He was under no illusion.” “She was acutely aware.” “The team understood the gravity of the situation.” These are claims. They assert that a person held a particular belief, experienced a particular awareness, or reached a particular understanding. None of it is sourced. None of it is evidenced. It is presented as established fact.

In fiction this is normal. The narrator knows what characters think. In a report nobody has that access. The writer is attributing an inner life on the basis of nothing recorded. The document now contains assertions that cannot be checked.

5.2 Imported stance and implied conclusions

Declamation brings a voice. A voice brings a position. “It cannot be ignored that…” tells the reader this matters. “Far from being straightforward…” tells the reader this is complicated. “And that, in the end, is the point” tells the reader what to conclude.

None of these add a fact. All of them add a direction. They steer the reader towards a judgement the document has not argued for. In an essay or an opinion piece that is expected. The whole point is to take a position. In a report, a case summary, or a factual record, the reader is supposed to reach their own conclusions from the evidence presented. Imported stance takes that away.

Inflation does the same thing less visibly. When a routine step is framed as a critical juncture, the text is telling the reader that this moment mattered more than others. That is a judgement. It may or may not be right. Either way, it was slipped in, not argued.

5.3 Why reports suffer most

Not all document types are equally damaged by register creep. An essay can carry a voice. A speech is supposed to have cadence. A feature article can take a stance. Register creep in those contexts is harder to see because the creep is towards something the genre already allows.

Reports are different. A report exists to present findings so that someone else can act on them. Its authority depends on restraint. It says what was found, what was done, what was observed. It does not say what to feel about it or what it all means. It does not tell the reader who understood what.

When register creep enters a report, it does not just sound wrong. It weakens the document. A reader who spots the rhetoric starts to wonder what else has been shaped. A reader who does not spot it is being led without knowing it. In either case the document has failed at its job.

This is why the problem is epistemic. It is not about how the text reads. It is about what the text knows, claims, and implies without warrant.

6. Triggers

Register creep is not random. Certain features in the source material make it more likely. The model does not creep uniformly across all text. It creeps when something in the content pulls it towards a genre it handles more fluently.

6.1 Charged words

Certain words carry strong genre associations. “Witness”, “evidence”, “truth”, “justice”, “victim”, “accused”. These appear in reports. They also appear in courtroom dramas, legal thrillers, political speeches, and investigative journalism. When the model encounters a cluster of them, its output drifts towards the genre where those words appear most often with the most expressive prose around them.

This is the trigger identified earliest. It is real but it is not the only one, and probably not the strongest.

6.2 Named persons as subject

When a document is about a named individual, the risk of narrative gravity rises sharply. A name turns a process into a story. The model begins to treat the person as a character. Actions become motivated. Situations become dramatic. Decisions become weighted.

This is especially marked when the person is described doing something: making a statement, attending a meeting, raising a concern. The model reads these as plot events. The person is given an inner life to match.

Depersonalised text is far less affected. “The statement was taken” does not trigger the same drift as “Mr Ahmed gave his statement.” The subject matters.

6.3 Chronology read as plot

A sequence of events is the backbone of a report. It is also the backbone of a story. The model does not reliably distinguish between the two.

When a report lays out events in order, the model tends to shape them into a narrative arc. Early events become setup. Later events become complications or resolutions. The last event becomes the conclusion. Tension is added between steps. Consequences are foreshadowed.

This is the hardest trigger to manage, because chronology cannot be avoided. Events did happen in an order. The report must say so. The problem is not the sequence. The problem is the arc that gets draped over it.

6.4 The prompt as trigger

Section 6.1 to 6.3 cover triggers in the content. The prompt itself is also a trigger.

Certain instructions invite register creep. “Write a professional summary” nudges towards formality and performance. “Provide a thorough analysis” nudges towards depth signalling. “Draft a comprehensive report” nudges towards breadth signalling and inflation. The model reads the adjective and delivers the register it associates with that adjective, not the register the document type requires.

Neutral phrasing reduces the risk. “Summarise the findings” is safer than “craft a professional summary of the findings.” The verb “craft” alone may be enough to shift the output. The more the prompt sounds like a brief to a speechwriter, the more the output sounds like a speech.

This trigger is unlike the others. It is controllable. Content triggers are properties of the subject matter and cannot be removed. Prompt triggers are properties of the instruction and can be reworded. Whether rewording reliably prevents the drift is an empirical question that has not been systematically tested.

6.5 Positional drift

Register creep may not be uniform across a document. There is reason to suspect that models start flat and drift as output lengthens. Early paragraphs follow the instruction. Later paragraphs follow the momentum of the text already generated.

If this is true, the phenomenon has a positional dimension. A long report may pass inspection in its opening pages and fail in its closing ones. A detection system that samples only the first few paragraphs would miss the drift entirely.

This is testable. A paragraph-by-paragraph register-creep score across a long output would show whether the density of markers increases with position. That test has not yet been done. But the possibility should inform how documents are sampled for assessment. The end of a long output may be more revealing than the beginning.

7. Diagnostic signatures by level

Each form of register creep leaves traces. Some are surface patterns that a script can match. Others are structural and require wider context. They are listed here by what to look for, not by how to detect it. Detection is covered in section 8.

7.1 Grand negations and contrastive frames

Syntactical frames that create contrast for dramatic effect. “He had no illusion that…”, “Far from being a simple matter…”, “It cannot be ignored that…”, “Not X but Y”, “Less about A than B.”

These frames hold the main point back. They set up a gap between what might be expected and what is asserted. In a speech this builds tension. In a report it delays the content for no purpose.

Grand negations and contrastive frames are a single family. The negation is a special case of the contrast: the expected state is denied before the actual state is given.

7.2 Rhetorical runways and delayed main clauses

Sentence openings that build anticipation before delivering content. “It is against this backdrop that…”, “To fully grasp the nature of…”, “What becomes clear upon closer examination is…”, “In light of what had already transpired…”

The main clause sits at the end. Everything before it is preamble. In speech, the runway gives the audience time to prepare. In a report, the reader wants the point first. The runway is pure performance.

A related form is the cleft sentence used for emphasis: “What he understood was that…” instead of “He understood that…” or better, a plain statement of what was understood. The cleft puts a spotlight on the content. A report does not need a spotlight.

7.3 Interiority modifiers

Adverbs and verbs that attribute emotional depth or acute awareness to professional activities. “Keenly observed”, “acutely understood”, “deeply wrestled with”, “quietly recognised”, “came to appreciate.”

These turn actions into experiences. “He reviewed the notes” becomes “He carefully considered the notes.” The second version claims something about the quality of attention. That quality was not measured or recorded. It is decoration in the shape of a fact.

7.4 Abstract nouns as agents

Sentences in which documents, timelines, or evidence perform actions. “The evidence demanded a response.” “The timeline raised questions.” “The disclosure painted a picture.” “The findings spoke for themselves.”

In each case, an inanimate thing is given agency. This is a form of inflation. It makes the material sound significant. It also avoids stating who actually did the demanding, questioning, or interpreting. The real agent is hidden behind a metaphor.

7.5 Manufactured rhythm

Deliberate variation in sentence length for dramatic effect. Typically a long, complex sentence followed by a very short one.

“The report covered eleven separate incidents across three sites over a period of eighteen months, each one documented in detail by the safeguarding lead, cross-referenced against the original referrals, and reviewed by two independent consultants. None of it mattered.”

The short sentence lands like a punchline. It is built for impact. It belongs in journalism or fiction. In a report the rhythm is incidental, not designed.

This is detectable. Sentence-length variance within a paragraph can be measured. A paragraph with low variance throughout and one outlier sentence at the end has the signature.

7.6 Paragraph-final codas

A closing sentence that draws a moral, makes an observation, or wraps the paragraph in significance. “And that, ultimately, told its own story.” “It was a pattern that would prove difficult to ignore.” “The implications were hard to overstate.”

The coda adds no fact. It tells the reader how to feel about the facts just presented. It converts a report paragraph into an essay paragraph. It is identifiable by position (last sentence), by content (abstract, evaluative, no new information), and often by brevity.

7.7 Tricolon, anaphora, and historic present

Three rhetorical devices borrowed from oratory and narrative.

Tricolon. A three-part list with parallel structure. “It was thorough, it was documented, and it was ignored.” The rule of three is a persuasion technique. It belongs in speeches. In a report, a list has as many items as the facts require.

Anaphora. Repetition of a word or phrase at the start of consecutive sentences or clauses. “There was no follow-up. There was no escalation. There was no record.” The repetition builds force. That force is rhetorical, not evidential.

Historic present. A shift from past tense to present tense in a narrative passage. “He submitted the report on Monday. On Tuesday the panel convenes.” The tense shift creates immediacy. It pulls the reader into the moment. A report written in the past tense should stay there.

All three are high-confidence markers. They are rare in genuine report prose. Their presence is a strong signal that the text has drifted.

8. Detectability: what scripts can and cannot do

The diagnostic signatures vary widely in how tractable they are. Some can be caught by a pattern matcher. Some can be approximated by heuristics that will also flag clean text. Some require comprehension.

8.1 Scriptable

These have surface forms that a matcher can look for with reasonable precision.

Grand negations and contrastive frames. Many have fixed openings: “far from”, “it cannot be ignored”, “not X but Y.” A curated list of frame-starters will catch a useful proportion. It will miss novel frames but will have a low false-positive rate, because these phrases are uncommon in flat report prose.

Rhetorical runways. Similarly pattern-based. “It is against this backdrop”, “to fully grasp”, “what becomes clear” are rare in reports and common in register-crept text. A list works.

Tricolon. Detectable by syntactic parallelism across three consecutive clauses. Not trivial, but tractable with a dependency parser. Three clauses of similar length and similar structure in sequence is a flag.

Anaphora. Consecutive sentences starting with the same word or phrase. Simple string matching with a threshold of three or more repetitions.

Manufactured rhythm. Sentence-length variance within a paragraph, with special attention to a short outlier at the end. Measurable. Needs a threshold to avoid false positives on naturally varied prose.

Historic present. A tense shift within a past-tense passage. A POS tagger can identify present-tense verbs in a context where surrounding verbs are past tense. Some false positives from habitual present (“the policy requires”) but manageable.

8.2 Heuristic with false positives

These can be approximated but not reliably isolated.

Interiority modifiers. A list of candidate adverbs and verbs (“keenly”, “acutely”, “wrestled with”, “came to appreciate”) will catch some instances. But “carefully reviewed” may be legitimate. The word alone is not the fault. The fault is that the quality of attention is being claimed without evidence. A script can flag candidates. A human or model must judge.

Abstract nouns as agents. Partially detectable by looking for inanimate nouns in subject position with active verbs. “The evidence demanded” has an inanimate subject and an active verb. But “the report states” is standard usage. The line between acceptable and inflated is contextual.

Paragraph-final codas. A last sentence that is short, abstract, and contains no proper nouns, dates, or figures is a candidate. But short concluding sentences are not always codas. The heuristic will over-flag.

8.3 Judgement only

These cannot be scripted. They require understanding the document, its purpose, and what has been established.

Stance. Whether the text is telling the reader what to conclude can only be judged by someone who knows what the text is supposed to do. “It cannot be ignored” is stance in a report and normal practice in a policy brief.

Narrative gravity. Whether a mental state has been attributed without evidence requires knowing what evidence exists. “He understood the risk” might be evidenced by his own recorded words. It might not. The script cannot tell.

Genre baseline. Whether a feature is a fault depends on the document type. A tricolon in a speech is expected. In a report it is a marker. The script needs to know what kind of document it is reading, and that judgement is itself non-trivial.

This is picture. Scripts cover level 1 fully and parts of level 2. Heuristics extend the reach into the rest of level 2 and parts of level 3, at the cost of false positives. Level 4 belongs to human judgement or to a model asked to read as a human would.

9. The fix at each level: substitute, rewrite, cut, delete

Knowing what is wrong is not the same as knowing what to do about it. The fix depends on the level. Each level requires a different kind of intervention. Applying the wrong fix to the wrong level either fails or introduces new problems.

9.1 Level 1: Substitute

A slop word is removed and a plain word takes its place. “Delve into” becomes “examine.” “Align” becomes “match.” The sentence keeps its shape. Only the vocabulary changes.

This is the easiest fix. It can be automated with confidence. A substitution table and a script will handle it. The risk is low. A bad substitution sounds odd but does not change the meaning or the structure.

9.2 Level 2: Rewrite

A sentence built for performance must be rebuilt for function. “Far from matching the CCTV footage, the statement, once checked, did not hold” must become “The statement was checked against the CCTV footage and did not match.”

Substitution does nothing here. Every word in the original is acceptable. The fault is in their arrangement. The sentence must be taken apart and reassembled with the main clause first, the modifiers in service of clarity, and the contrast frame removed.

This is harder to automate. A model can do it. A script cannot, except for a narrow set of known frame types where the rewrite follows a predictable pattern. “Far from X, Y” can be mechanically inverted to “Y. It was not X.” But that mechanical inversion often sounds awkward. A good rewrite requires judgement about what the sentence was trying to say.

9.3 Level 3: Cut or reorder

A paragraph shaped for effect must be reshaped for the record. The coda must go. The moral must go. Events sequenced for suspense must be resequenced for chronology or logic.

Cutting is the simpler operation. If the last sentence adds no fact, remove it. If a sentence exists only to build tension before the next one, remove it. The test is: does this sentence carry information that the reader needs? If not, it is performing, and performance is not the job.

Reordering is harder. A paragraph that builds to a reveal has put the most important fact last. A report paragraph should usually put it first. But deciding what the most important fact is requires understanding the content, not just the structure.

There is a temptation to fix level 3 by fixing each sentence. That does not work. A paragraph of clean sentences can still have the shape of an essay. The fix is at the paragraph level: what is here, what order is it in, and does anything need to go.

9.4 Level 4: Delete

A claim about a mental state that is not evidenced must be removed. There is no softer option. It cannot be reworded into something acceptable, because the problem is not the wording. The problem is that the claim exists.

“He was under no illusion that the statement would survive scrutiny” cannot be fixed by changing “was under no illusion” to “believed.” The new version still claims knowledge of a mental state. The fix is deletion. If his belief matters, it must be evidenced: “In his written response, he stated that he did not expect the statement to withstand scrutiny.” If no such evidence exists, the belief is not part of the record.

The same applies to imported stance. “It cannot be ignored that…” cannot be fixed by changing “cannot be ignored” to “is notable.” Both versions tell the reader what to think. The fix is to remove the frame entirely and state the fact. The reader will decide whether to ignore it.

9.5 A note on over-correction

The fixes are not symmetrical in their risks. Substitution at level 1 rarely causes harm. Deletion at level 4 can cause harm if done carelessly. Removing a sentence that attributes a mental state may also remove a fact embedded in it. “He was under no illusion that three referrals had been made” contains a factual claim (three referrals) wrapped in an epistemic claim (his awareness). The deletion must preserve the fact and lose the framing. “Three referrals had been made” keeps what matters.

Over-correction at level 3 can flatten prose that was legitimately structured. Not every paragraph-final sentence is a coda. Not every chronological sequence has been arranged for suspense. A text stripped of all shape reads as a list of disconnected observations. The goal is a report voice, not no voice at all.

9.6 Prevention

Sections 9.1 to 9.5 deal with correction after the text exists. A prior question is whether register creep can be prevented before it appears.

Three levers are available. Prompting strategy: how the instruction is worded, what register is specified, whether examples of the desired output are provided. System instructions: standing directions to the model about tone, voice, and what to avoid. Model selection: different models may be more or less prone to the drift, depending on their training data and reinforcement history.

None of these has been systematically tested against the four-level framework. Anecdotal evidence suggests that explicit negative instructions (“do not use rhetorical questions, do not attribute mental states without evidence”) reduce level 2 and level 4 features. Whether that reduction holds across long outputs or under content-trigger pressure is unknown.

Prevention and detection are complementary. A prompt that reduces register creep makes the detection system’s job easier. A detection system that measures what gets through makes the prompt’s effectiveness measurable. Neither replaces the other.

Early first test

An early test of an inlet filter in OpenWebUI produced one observation worth recording. The constraint text was injected into the system prompt using a ‘function‘ before generation.

The Function

"""
title: Register Constraint
author: Captain Walker
version: 0.1.0
description: Inlet filter that injects register-discipline constraints into the system prompt to reduce register creep in model outputs.
"""

from pydantic import BaseModel, Field
from typing import Optional


class Filter:
  class Valves(BaseModel):
      enabled: bool = Field(
          default=True, description="Enable or disable the register constraint."
      )

  def __init__(self):
      self.valves = self.Valves()

  async def inlet(self, body: dict, __user__: Optional[dict] = None) -> dict:
      if not self.valves.enabled:
          return body

      constraint = """
## Register Discipline

Match the register to the task. Do not drift upwards in formality, drama, or rhetoric unless the task explicitly calls for it.

When writing reports, summaries, case notes, factual records, or any document whose purpose is to present findings:

Sentence construction:
- Put the main clause first. Do not delay it with preamble.
- Do not use contrastive frames for emphasis ("far from", "not X but Y", "it cannot be ignored that").
- Do not use cleft sentences for spotlight effect ("what became clear was...").
- Do not build sentences for spoken delivery. Build them for silent reading.

Paragraph construction:
- Do not end paragraphs with a moral, an aphorism, or an evaluative coda.
- Do not sequence events for suspense. Sequence them for the record.
- Do not manufacture rhythm through deliberate sentence-length contrast.
- Do not use tricolon or anaphora for rhetorical force.

Stance:
- Do not attribute mental states unless citing evidence ("he was under no illusion" is prohibited unless his own words are quoted).
- Do not give agency to abstractions ("the evidence demanded", "the timeline raised questions").
- Do not tell the reader what to conclude. Present the facts.
- Do not use evaluative intensifiers ("quietly devastating", "telling", "subtle but important").
- Do not add commentary, editorial observations, or analysis after the conclusion or decision. Once the findings and outcome have been stated, stop. If there is a learning or actions section, confine it to what was decided. Do not observe tensions, draw attention to contradictions, or note what an external reviewer might think. That is not the document's job.

When the task is creative writing, essays, speeches, or persuasive text, these constraints do not apply. Register creep is only a fault when the register does not match the task.
"""

      messages = body.get("messages", [])

      has_system = len(messages) > 0 and messages[0].get("role") == "system"

      if has_system:
          messages[0]["content"] = messages[0]["content"] + "\n\n" + constraint
      else:
          messages.insert(0, {"role": "system", "content": constraint})

      body["messages"] = messages
      return body

The output was flatter, more concise, and stopped when the facts were stated. That was expected. What was not expected was that the model’s chain-of-thought reasoning also changed. There was less performative deliberation. Less rhetorical planning. The model spent less effort deciding how to frame the material and more on what the material was.

This suggests that register constraints do not only suppress features in finished prose. They may steer the model away from the rhetorical planning that produces those features in the first place. The model does not build a runway because it never plans to land dramatically. The effect is upstream of the text.

If confirmed across more cases, this has a practical implication. Prevention may be cheaper than correction. A constrained generation that never produces register creep requires no editing. An unconstrained generation that does requires a second pass. The constraint adds a few hundred tokens to the system prompt. The second pass costs more.

Second test

A second observation from the same test. The constrained output also showed fewer level 1 faults. Slop words were reduced, though the constraint text said nothing about vocabulary.

This was not expected. The four levels are presented in this framework as independent. A register constraint targets levels 2 to 4. It should not touch level 1. But if slop words and register creep share a common upstream cause, the independence may hold as a description but not as a mechanism.

The candidate explanation is this. When the model plans rhetorically, it enters a mode that produces both dramatic structure and dramatic vocabulary. The constraint disrupts that mode at the planning stage. The structures and the words are symptoms of the same stance. Suppress the stance and both drop.

If that holds, it reframes the practical relationship between the Slop Assessor and the register constraint. A word list treats level 1 symptoms. A register constraint may treat the cause that produces them. The two tools are not redundant. But the constraint may do more of the level 1 work than a framework based on independent levels would predict.

This is a preliminary observation from a single test. It needs replication across models and task types. The independence of the four levels as descriptive categories is not in question. Their independence as generative mechanisms may be.

This is a single observation, not a finding. It needs replication across different models, different tasks, and different prompt types before it can bear weight.

10. Limitations and open questions

This framework is a working model. It is useful. It is not finished. Its known limitations are set out below. Several questions remain that only testing or further work can answer.

10.1 Where the drift originates

The framework assumes that register creep is a product of language model training. That is likely but not proven. The mechanism is unclear. It may come from pretraining on fiction, journalism, and speeches. It may come from post-training, where raters rewarded engaging prose. It may come from both, in different proportions for different forms. It may vary between models.

The framework does not need the causal story to work. The signatures are observable regardless of their origin. But understanding the cause would help predict which models are more susceptible and whether fine-tuning can reduce the drift.

10.2 Whether human writers do the same thing

Register creep is not unique to models. A human report writer who has read too many opinion columns may start writing like one. A clinician trained on case studies may write a factual record as if it were a teaching narrative. The four levels apply to human text as well.

This matters for detection. If a text shows level 2 and level 3 features, that may indicate AI generation. It may also indicate a human writer with a particular style. The features are evidence of register creep, not evidence of origin. The two questions must be kept separate.

10.3 The genre baseline problem

Section 8 noted that whether a feature is a fault depends on the document type. A tricolon in a speech is normal. In a report it is a marker. Any detection system needs to know, or be told, what it is reading.

How that genre identification works is unresolved. The user could declare it. The system could infer it from structural cues. A hybrid is possible. But without a genre baseline, every flag is conditional. The system can say “this feature is present” but not “this feature is a fault.” That distinction matters.

10.4 Domain-specific variation

The framework treats “report” as a single genre. It is not. A psychiatric report, a legal submission, a structural engineering report, and an audit summary each have different register norms. What counts as inflation in one may be standard practice in another.

A psychiatric report may legitimately discuss a subject’s apparent emotional state, because clinical observation of affect is part of the task. The same sentence in an engineering report would be absurd. A legal submission may use contrastive framing (“not X but Y”) as a standard argumentative structure. The same framing in a factual summary would be a register-creep marker.

This means that the diagnostic signatures from section 7 cannot be applied uniformly. Each one needs a domain modifier: how likely is this feature to appear legitimately in this type of document? Building those modifiers requires domain expertise. A general-purpose tool can flag. Only a domain-aware tool, or a domain-aware reader, can judge.

10.5 Cultural register norms

UK report prose and US report prose differ. UK professional writing tends towards understatement. US professional writing tends towards assertion. A sentence that reads as inflated in a UK context may read as normal in a US one. The reverse also holds: UK hedging (“it may be worth considering”) can read as evasive in a US context.

The framework as written reflects UK register expectations. That is its origin. Applying it to US-authored text without adjustment would produce false positives on features that are culturally normal rather than register-crept.

This is a calibration problem, not a structural one. The four levels hold across cultures. The thresholds differ. A tool intended for use beyond a single national context would need switchable baselines. How many baselines are needed, and how they differ in practice, is unexplored.

10.6 Interaction between levels

The four levels are presented as separate. In practice they interact. A level 2 frame (a rhetorical runway) may enable a level 4 fault (an unsupported stance). “It is against this backdrop that his frustration becomes understandable” has a runway at level 2 and an unverifiable mental state at level 4. Fixing the runway alone leaves the stance. Fixing the stance alone still leaves a sentence shaped for performance.

Whether to treat these as independent faults or as a single compound fault is undecided. For detection, counting them separately makes sense. For correction, they may need to be addressed together.

10.7 Multi-turn compounding

The framework assumes a single output. Much AI-assisted writing is iterative. A draft is generated, revised, regenerated, and revised again. Each pass takes the previous output as input.

If register creep is present in an early draft, the model reads it back as part of the baseline. The crept register becomes the norm for the next pass. Features that were marginal in the first output become established in the second. The drift compounds.

This has implications for both detection and correction. A document that has been through several revision passes may show a density of register-creep features that no single generation would produce. The features reinforce each other across turns. Correcting the final output may require more intervention than correcting the first draft would have.

It also has implications for prevention. If compounding is real, the most effective point of intervention is the first draft. Cleaning register creep early prevents it from seeding subsequent passes. Leaving it in place and correcting at the end is harder because the crept register has become structural.

Whether compounding occurs in practice, and how quickly the drift accelerates, has not been measured. It is a plausible mechanism. It has not been confirmed.

10.8 Thresholds and density

A single contrastive frame in a twenty-page report is not a problem. Twelve of them across three pages is. The framework needs thresholds, and those thresholds will differ by feature, by level, and by genre.

Setting those thresholds requires testing against real documents. The numbers cannot be derived from theory. They must be calibrated empirically, and that work has not yet been done.

10.9 Extensibility

The diagnostic signatures listed in section 7 are not a closed set. They were identified from observation of current model outputs. Models change. New signatures will emerge. The framework is designed to accommodate them. The four levels are structural categories. New signatures slot into whichever level they belong to. No redesign is needed.

The same applies to the levels themselves. If a fifth level is identified, it sits alongside the existing four. The architecture is a taxonomy, not a fixed list. That is deliberate.

11. Where this sits in the Slop Assessor

The Slop Assessor was built to catch slop. It works with patterns: word lists, phrase lists, density scoring, clumping. It flags text that is padded, repetitive, or filled with stock language. That is level 1 work. It does level 1 well.

Register creep is not slop. It is a different fault on a separate axis. Text can be slop-free and register-crept. Text can be full of slop and perfectly flat in register. The two problems are independent. But they belong in the same tool, because the user asking “is this text sound?” wants both answers, not one.

11.1 What can be added now

The level 2 signatures with fixed surface forms can be added to the existing pattern framework without changing the architecture. Grand negations, rhetorical runways, contrastive frames, anaphora, and historic present all have matchable surface strings or measurable structural features. They are new patterns in the existing system.

Manufactured rhythm (sentence-length variance with a short outlier at the end of a paragraph) can be added as a metric alongside the existing density and clumping measures. It requires paragraph segmentation rather than whole-document scoring, but the infrastructure for that is already present or close.

Tricolon detection requires a dependency parser. If spaCy is already in use, this is within reach. Three consecutive clauses of similar structure and length is a parseable pattern.

11.2 What needs a new layer

The heuristic signatures from section 8.2 do not fit cleanly into the existing pattern framework. Interiority modifiers, abstract nouns as agents, and paragraph-final codas can be flagged, but each flag needs a qualifying statement. “This may be an interiority modifier” is a different kind of output from “this is a slop word.” The assessor currently deals in findings. Heuristic flags are candidates, not findings.

This suggests a second tier of output: a register-creep panel, separate from the slop panel, that presents flags with lower confidence and explicit caveats. The user reads the slop panel for definite faults and the register panel for things worth checking. The two panels serve different purposes