Sunday, July 12, 2026
spot_img

Why and How LLMs Fail: Context Degradation and the Evans Law Series

  1. This body of work began as an empirical investigation into repeated failure modes observed during sustained use of large language models. What first appeared to be isolated hallucinations or context-window problems became, through repeated testing across models, tasks, modalities, and interaction lengths, a broader pattern: LLM failures are not random surface defects but predictable consequences of the architecture itself.
  2. The central axiom developed from this work is Evans’ Law: the longer a model reasons, the more likely it is that it will produce an incorrect or degraded response, until the likelihood of an incorrect or degraded response becomes higher than the likelihood of a correct one. The early papers establish this threshold behaviour in long-context reasoning, showing that model performance does not simply improve with longer context windows or larger parameter counts. Instead, coherence degrades under sustained reasoning load, producing measurable collapse points, degraded reference stability, and eventual factual or semantic failure.
  3. The series then expands Evans’ Law from a threshold model into a broader reliability framework. Later work introduces the idea of a reliability surface, accounting not only for context length but also for task rigidity, input complexity, output complexity, and verification burden. This moves the theory beyond “how long can a model reason before it fails?” toward a more operational question: under what conditions does a model become unreliable, and how can that degradation be detected, measured, governed, or mitigated?
  4. A major branch of the research examines multimodal degradation. These papers argue that failures become more complex when models must process and reconcile text, image, document structure, layout, and contextual inference at the same time. The multimodal work introduces the cross-modal degradation tax: the added reliability burden created when a model must maintain coherence across different modes of representation. In this framework, multimodal failure is not just hallucination plus vision error; it is an additional class of degradation caused by fractured alignment between modes.
  5. The agentic AI work extends the same empirical concern into systems that claim to plan, act, remember, decide, and coordinate over time. These papers question whether “agentic AI” exists in the strong sense often implied by vendors, or whether current systems are better understood as probabilistic language models wrapped in orchestration layers. The research identifies architectural risks in agentic systems, including compounding error, unstable task interpretation, poor recovery from drift, brittle tool use, and the illusion of autonomy where no reliable internal agency exists.
  6. The hallucination papers form another core layer of the series. Rather than treating hallucination as a generic content problem, they examine it as a mechanistic failure of coherence, reference, significance, and repair. The work argues that source-grounding, retrieval, and longer context can reduce some classes of error but do not eliminate hallucination because they do not solve the deeper architectural problem: the model still lacks a stable native mechanism for determining what must remain true, what must remain attached to whom or what, and what must not be overwritten by statistically plausible continuation.
  7. From there, the research develops the concept of missing primitives in contemporary transformer systems. The papers on strict semantic dominance and revocable semantic dominance argue that LLMs lack native mechanisms for enforcing durable meaning relationships. This becomes especially visible in proper noun failures, identity drift, source confusion, and semantic authority failures, where models mishandle names, entities, relationships, and claims even when the relevant information is present. The proper noun work shows that names are not a minor edge case but a stress test for whether a system can preserve identity and reference under pressure.
  8. The S-vector papers propose a possible architectural and operational response. The S-vector, or significance vector, is introduced as a missing dimension in current language-model representations: a way to encode not merely what tokens mean, but how much they matter within a task, document, system, or institutional context. The research argues that current models are “flat” with respect to significance. They can represent association, probability, and semantic proximity, but they do not reliably distinguish between disposable details and high-consequence anchors such as names, legal terms, medical facts, financial figures, authorship, institutional authority, or safety-critical instructions. The S-vector work explores how significance weighting could improve long-context stability, retrieval, enterprise AI, agentic systems, and governance.
  9. Taken together, the technical papers form a unified empirical theory of LLM failure and evolution. They show that hallucination, long-context degradation, multimodal collapse, proper noun failure, agentic instability, and semantic drift are not separate problems. They are related expressions of a deeper architectural limitation: current models predict continuation without a sufficiently stable internal mechanism for preserving significance, authority, identity, and verification across time.
  10. The final stage of the series begins to connect these architectural findings to public policy, institutional procurement, and AI sovereignty. If LLMs fail in predictable ways, and if those failures are amplified in government, infrastructure, healthcare, justice, immigration, education, and public administration, then AI governance cannot be limited to privacy, procurement, ethics, or domestic model ownership alone. The sovereignty work asks who controls the models, infrastructure, standards, data flows, evaluation systems, and decision layers increasingly embedded in public life. It argues that AI sovereignty is not simply a matter of building national models; it is a question of whether governments can understand, audit, govern, and retain authority over the systems on which they are becoming dependent.
  11. Across the full series, the work moves from observation to axiom, from axiom to measurement, from measurement to architecture, and from architecture to governance. Its central claim is that LLM reliability cannot be solved by scale alone. The future of trustworthy AI will depend on whether systems can be designed, evaluated, and governed around the failure modes they actually exhibit: degradation over reasoning length, instability across modalities, weak semantic authority, poor significance weighting, fragile agency, and the unresolved problem of knowing what must remain true.

  12. Evans’ Law: A Predictive Threshold for Long-Context Accuracy Collapse in Large Language Models
  13. Evans’ Law v4.1 (Extended)
  14. Evans’ Law: Scaling, Coherence, and Governance Implications V4.0
  15. Evans’ Law 5.0: Long-Context Degradation in Multimodal Models and the Cross-Modal Degradation Tax
  16. AI’s Unmeasured Reality: How Users Are Left Behind
  17. AI’s Accountability Gap: A Policy Blueprint for Policymakers
  18. Does Agentic AI Exist? v6.0
  19. The When Where and How of LLM Failures, Measured
  20. Why Hallucinations Happen: Fracture and Repair in Transformer Systems v1
  21. The S-Vector: Topographic Attention and the Architecture of Intelligence
  22. The Mechanistics of Hallucination Version 3.0
  23. Why the S-Vector Matters: The Missing Dimension in Enterprise AI — and What Companies Should Try
  24. Research Summary: A Unified Theory of LLM Evolution
  25. Two Missing Primitives in Contemporary Language Models: Strict Semantic Dominance and Revocable Semantic Dominance
  26. The Missing Key to True LLM Intelligence 3.0: An Operational Roadmap for the S Vector
  27. Beyond Content: Proper Nouns and Semantic Governance Failures in LLMs
  28. Why Agentic AI Is Problematic: The Architectural Risks
  29. Source-Grounding Does Not Prevent Hallucinations: A Controlled Replication Study of Google NotebookLM
  30. Coordination, Significance and Manifold Efficiency: A Path to Transformative Intelligence
  31. Significance Weighting in Large Language Models: Cross-Architecture Behavioral Evidence
  32. Proper Noun Failure: An Empirical Update on Evans’ Law
  33. Nudgment Signal Discernment Framework
  34. Whose AI Runs the Government?

Featured

The Third Position: Collective Procurement and the NATO Maven System

Part of the Canadian AI sovereignty series My recent analysis...

Data Adjacency: How Canada Is Now Exposed to AI Systems It Never Procured

Part of the Canadian AI Sovereignty Series The implications of...

Models May Be Attempting Ambiguity Resolution, a Small Subset of Intelligence

On July 6, Anthropic published interpretability research identifying what...

Quick Take: Is a flock of AI unicorns a bubble?

As of June 2026, there are currently 1,778 startup...
Jennifer Evans
Jennifer Evanshttps://www.b2bnn.com
Principal, patternpulse.ai, and cofounder, Tech Reset Canada. AI policy, research and analysis. Entrepreneur since 2002, marketer since 1998, machine learning since 2009. Based in Toronto and Southeast Asia.