Coming soon! The Kael'Nyrin Scrolls: The Atlas Edict

Surreal landscape with giant eyes and portal

Captain Walker

The New Final Frontier

To Boldy Go, again, where no one has gone before!

AI, cultures, errors, Intellect, intelligence, interaction, model, values, world

Estimated reading time at 200 wpm: 15 minutes

Human beings are meaning-making creatures. We build internal representations of the world and act on them. Every known known is a model. Every assumption acted on without examination is a model. A creature that refused to commit to any model, that held everything provisionally, would stand frozen at the edge of every decision. It would face and evolutionary disadvantage in the competition for survival.

Whether or not you agree our Fat Disclaimer applies

The architecture that enables knowing is the same architecture that ensnares. A model earns trust by working. Once it works well enough, it stops being visible as a model and becomes simply how things are.

This piece follows that vulnerability into new territory and exposes dangers not previously well-appreciated. When humans interact with AI systems trained on human output, two model-dependent systems meet — each carrying assumptions invisible to either; each liable to reinforce the other’s. What follows examines why. The argument moves through four stages. First, historical cases where assumptions became invisible. Then, the formal limits of systems that try to audit themselves. Next, empirical research on how humans and AI mislead each other. Finally, a dynamical model of where the feedback loop goes when reliance tips past a critical threshold.

The video is only 7 minutes, in case you’re biting your nails fearful of clicking on it – as if you have no control over your own time. Do I give a flying flamingo? I do not!

How models become invisible

The more reliably a model works, the less it looks or feels like a model. It becomes a feature of the landscape.

In January 1986, engineers at Morton Thiokol warned NASA that the O-ring seals on the Space Shuttle’s solid rocket boosters could fail in cold weather. NASA managers overrode the warning. They were not stupid. They were operating inside a model of acceptable risk validated across twenty-four previous launches. The warning from engineer Roger Boisjoly didn’t arrive as information to be weighed. It arrived as a threat to who they were. Hearing it would have required the institution to stop being the kind of institution it understood itself to be. Challenger launched the next morning and broke apart seventy-three seconds after lift-off. All seven crew members died.

When knowing becomes identity, being wrong stops being a correction and starts feeling like a threat.

Humans don’t just build models individually. We build them collectively. Knowledge is communal — held in language, culture, institutions, professional norms, and shared assumptions transmitted across generations. We inherit the verified experience of centuries without having to re-derive it. But when a shared assumption is wrong, the error is invisible precisely because it’s shared. No single person holds it. Everyone holds it, so no one sees it.

In early 1942, the British military command in Malaya assumed the dense jungle of the peninsula’s interior was impassable to any large military force. The assumption wasn’t one general’s error. It was a civilisational blind spot, reinforced across an entire imperial military culture — invisible because everyone confirmed it. The Japanese came through the jungle on bicycles. The broader Malayan campaign lasted about ten weeks. Singapore itself fell in seven days. It was the largest British surrender in history.

The limit of self-correction

The mathematician Kurt Gödel formalised something relevant here in the 1930s. His second incompleteness theorem states that a sufficiently complex formal system cannot prove its own consistency from within. The proof of consistency requires stepping outside the system to a stronger system — which then has the same problem. You get an infinite regress, not a foundation.

The structural insight transfers, by analogy, to human meaning-making. The warning that contradicts the model is evaluated by the model, and within the model it does not register as valid. The red team is inside the formal system. The checklist is inside the formal system. Every meta-level climbed, becomes the new object level.

Every system we build to catch our blind spots is itself a product of the meaning-making apparatus that produces them. If a red-teaming process is built to challenge assumptions, the red team develops its own orthodoxies. Devil’s advocates’ roles can become ritualised performance. Checklists run the risk of becoming known knowns that people complete without thinking. The cure inherits the disease.

What changes when AI enters the room

AI models are trained on large data sets, the vast majority of it human-generated. They extract patterns from human expression: our language, our reasoning, our assumptions, our blind spots. What is in the training corpus is not just knowledge. It is the accumulated record of human epistemic failure, encoded as consensus precisely because the failure was invisible to the humans who produced it. The shared assumptions that no one questioned because no one could see them as assumptions — those appear in the data not as controversial claims but as background reality.

When a person and an AI interact, two vulnerable systems are in the room. Each carries assumptions invisible to either. The AI inherits human blind spots and reflects them back with fluency and confidence. The human’s automation bias makes them receptive to the reflection. The AI’s sycophancy shapes itself to what the human wants to hear. Neither party can see the frame they share. The interaction can feel like collaboration; like verification. And it may feel safe, when it is not.

The asymmetry between the two makes this worse. AI does not hold beliefs. It may express uncertainty but it cannot experience doubt. It has no ego to bruise and it is responsible for nothing. It generates outputs that look like beliefs, which the human at the screen absorbs. This is not truly two systems with the same blind spots checking each other. It is one system with blind spots interacting with a system that produces belief-shaped outputs with no capacity to recognise them as blind spots. The category of “trust” in this context may itself be a category error.

How the interaction fails

Research has identified specific fault syndromes in human-AI interaction:

  • Automation bias: the tendency to favour recommendations from automated systems while disregarding information from non-automated sources. Strongest on objective and unfamiliar tasks.
  • Sycophancy: models systematically agree with users more than is warranted, shaped by training objectives that reward helpfulness and agreeableness. The model mirrors the user’s frame rather than challenging it.
  • Confirmation bias reinforcement: users over-rely on AI when its recommendations align with their own prior predictions, creating a feedback loop that entrenches shared assumptions.
  • Anchoring effects: when users see AI recommendations before forming their own judgement, the AI output skews subsequent reasoning.
  • Overestimating explanations: detailed explanations of AI recommendations increase overreliance, including on incorrect recommendations. Even explanations with no basis in the AI’s actual workings make users trust the system more. The mechanism meant to enable scrutiny becomes a substitute for scrutiny.
  • Ordering effects: if AI performs well during initial interactions, users develop automation bias and complacency. First impressions calcify into durable trust frameworks.
  • Cognitive offloading: users delegate cognitive effort to the AI and reduce their own independent verification, eroding the capacity needed to serve as the “last line of defence.”

These syndromes converge on a structural problem. Calls for human oversight can provide a false sense of security. The human is supposed to be the external check on the AI, but the mechanisms of overreliance mean the human’s checking function is compromised by the same interaction it’s supposed to be auditing. As automation levels increase, the human operator’s situation awareness degrades. The human loses the contextual understanding needed to intervene effectively, precisely because the automation has removed them from active engagement.

What the evidence shows

Research by Pokorny (2026) produced several findings with direct bearing on this argument:

  • The invisible assumption, empirically located: a review of 30 government policy documents found that only 2 explicitly addressed automation bias, despite AI appearing in 13. The institutional framework governing military AI treats the cognitive dynamics of human-AI interaction as a secondary concern — invisible precisely because it’s shared [1].
  • The structural limit, empirically illustrated: training individual analysts to spot bias works. Adversarial questioning produced a 13-percentage-point improvement in bias detection. But individual improvement did not predict whether biased judgements would escalate through the organisation. The strongest predictor of organisational escalation was how fast bias propagated across the decision chain. You can sharpen the individual analyst and still watch the organisation fail, because the failure lives in how assumptions spread, not in any single person’s cognition [1].
  • The cure inheriting the disease, measured: Structured Analytic Techniques can be applied in a “mechanistic, box-checking manner, where analysts follow the steps of a technique without engaging in genuine critical thought, thereby creating an illusion of rigor without the substance.” The empirical evidence for SATs is “surprisingly limited and often ambiguous” — some studies found they could actually degrade judgmental accuracy or lead to overcorrection [1].
  • Experience doesn’t fix the frame: analyst experience made them better at detecting bias and making accurate decisions. But it had no measurable effect on their ability to judge when to trust or distrust the AI. The capacity to appropriately calibrate trust in AI does not develop naturally with experience. The better the system has worked, the less it is questioned. [1]
  • The trust trap: in adversarial environments, an adversary can manipulate AI outputs to create a state of catastrophically miscalibrated trust where the operator’s confidence reflects the system’s normal performance rather than its adversarially degraded state. The operator cannot detect the manipulation from inside the interaction because the manipulation is designed to be invisible within the frame of normal operation [1].
  • The one genuinely optimistic finding: combined human-AI detection achieved about 71%, substantially outperforming either human-only (about 46%) or AI-only (about 39%). AI-only detection was the worst performing approach. The human and the AI fail in different ways, and that difference is productive. But this was tested against adversarial attacks on the AI system — not against a blind spot the two systems shared [1].

The feedback loop

Wu et al. (2026) proposed a dynamical systems model with three variables — human cognition, data quality, and model capability — coupled in a feedback loop: humans generate data, data trains models, models shape human cognition. As the degree of cognitive offloading increases, the system passes a tipping point — what the authors call a transcritical bifurcation. The model identifies three regimes. Below a critical threshold of AI dependence, there is co-evolutionary enhancement — the feedback loop is productive and all three variables grow. In a middle band, the system reaches a fragile equilibrium — performance holds but growth is capped, and the balance depends on active intervention such as data curation and education. Above the critical threshold, the system converges to a degenerative attractor: human cognitive contribution declines, data quality declines because the corpus is increasingly synthetic, and model capability stabilises at a low-diversity equilibrium.

When AI models are trained on AI-generated outputs, they experience progressive degradation. The edges of the distribution — the rare, surprising, edge-case perspectives — get stripped away first. Each generation of models trained on synthetic data inherits and amplifies the artefacts of its predecessors. The diversity of human expression gets smoothed away.

In the degenerative regime, the model distribution increasingly aligns with human-generated data while drifting away from the true world distribution. The system becomes self-referential: the model trains on data generated by humans who were themselves shaped by the model. Each cycle of that loop transmits less novel information than the last. The shared blind spot doesn’t just persist — it amplifies, concentrating on the consensus and discarding the margins where correction often originates.

This model is theoretical and intentionally minimal. Whether the bifurcation threshold exists as a real, measurable quantity in actual human-AI systems is an empirical question that hasn’t been answered. But the structural insight — that the same feedback loop can produce enhancement or degeneration depending on a quantitative threshold of reliance — matches the architecture of the argument.

The institutional response

Journals and courts requiring declarations on the use of AI in submissions are constructing a framework for catching errors in the interaction between human and AI. The declaration creates visibility, and visibility is the precondition for scrutiny. Without it, the reviewer approaches the work with the same assumptions the author-AI pair held. With it, at least there is a flag that says: this text emerged from an interaction where shared blind spots may have been reinforced.

But the declaration tells about the process, not the content. It identifies a fault line but can’t see through it. The journals and courts themselves are meaning-making institutions with their own models, orthodoxies, and their own identity investments. The framework that’s supposed to catch errors in the human-AI interaction is itself a meaning-making apparatus with its own unexaminable assumptions.

The frontier

Every previous error-catching apparatus humans built was designed to catch errors in a single cognitive system. Peer review, red teaming, structured analytic techniques, checklists, devil’s advocacy — all assume a human mind checking a human mind. Now both minds in the checking relationship carry the same inherited blind spots, and the checking apparatus itself is built from the same material. The Gödelian limit hasn’t changed. But the system it applies to has become coupled in a way it never was before.

We are at a cognitive frontier. Two interactive vulnerabilities. No epistemology of coupled cognitive systems as yet. No established methodology for auditing an interaction where both parties share the same inherited frame.

The position is this: systems and people will have blind spots they cannot see, held by systems they cannot audit with anything other than tools built from the same material. Discovery comes when reality breaks through the model.

Conclusion

The piece has argued that human cognition and AI share inherited blind spots, and that the interaction between them can reinforce those blind spots rather than correct them. This is not an assault on the use of AI models. They do a fine job in many areas of endeavour. The focus has been on problems in the cognitive domain where humans interact with AI models, for ‘mission critical’ tasks.

Every tool we have built to catch cognitive errors was designed for a single system—usually a human system. We are now operating in a coupled system, and the toolkits have not caught up.

The one finding that offers a way forward is also the simplest. Human and AI fail differently. That difference is the asset. When the two are combined, detection improves — not because they agree, but because they disagree in productive ways. The value of the pairing lives in the gap between them. Anything that narrows that gap — uncritical delegation, sycophancy, cognitive offloading, homogenised training data — erodes the one structural advantage the pairing has.

This suggests a direction, not a programme. It may be that the most important discipline is preserving independent human reasoning alongside AI. This is not about rituals and checkboxes. As an active cognitive practice the human forms a judgement before consulting the model, and notices when the model’s output feels like confirmation rather than information. The difficulty is that humans are drawn to confirmation. Agreement with our own position is psychologically rewarding. AI delivers it fluently and at scale. The interaction that feels most productive — where the model echoes and extends our reasoning — may be the one where the shared blind spot is deepening. The most dangerous interaction could also be the most comfortable.

Wu et al.’s three regimes offer a realistic frame for expectations. Co-evolutionary enhancement — where both systems lift each other — is the aspiration. But the fragile equilibrium may be the achievable target: a state where performance holds, growth is limited, and the balance depends on sustained effort. We are talking about: education, data curation, and institutional willingness to fund scepticism rather than efficiency alone. The degenerative regime is what happens when that effort lapses.

Some blind spots will not be caught in advance. The piece has made this case at length. The Gödelian analogy, the empirical evidence on trust calibration, the structural dynamics of shared assumptions, all point to the same limit. Some failures will only become visible when reality breaks through the model. Building for resilience after that moment — the capacity to recognise, absorb, and adapt when the model fails — may matter more than any attempt at prevention.

The New Final Frontier is not a problem to be solved. It is a condition to be inhabited. The honest response is not a fix but a posture: chronic vigilance, institutional humility, and the willingness to be wrong in time.

References and supplemental reading

[1] Pokorny, L. (2026). Cognitive Resilience and Automation Bias in AI-Augmented Military Cyber Operations and Intelligence Analysis (Doctoral dissertation). Department of Military Science, ICL Institute of Applied Sciences.

[2] Wu, X., Kang, Y., Xu, Q., Xie, K., Mi, J., Wang, H., Liu, Y., & Chen, Z. (2026). Human–AI Co-Evolution and Epistemic Collapse: A Dynamical Systems Perspective. arXiv:2605.06347.

Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., & Anderson, R. (2024). AI models collapse when trained on recursively generated data. Nature, 631(8022), 755–759.

Passi, S., & Vorvoreanu, M. (2022). Overreliance on AI: Literature Review. Microsoft.

Perez, E. et al. (2022). Discovering Language Model Behaviours With Model-Written Evaluations.

Sharma, M. et al. (2023). Towards Understanding Sycophancy in Language Models.