- This body of work began as an empirical investigation into repeated failure modes observed during sustained use of large language models. What first appeared to be isolated hallucinations or context-window problems became, through repeated testing across models, tasks, modalities, and interaction lengths, a broader pattern: LLM failures are not random surface defects but predictable consequences of the architecture itself.
- The central axiom developed from this work is Evans’ Law: the longer a model reasons, the more likely it is that it will produce an incorrect or degraded response, until the likelihood of an incorrect or degraded response becomes higher than the likelihood of a correct one. The early papers establish this threshold behaviour in long-context reasoning, showing that model performance does not simply improve with longer context windows or larger parameter counts. Instead, coherence degrades under sustained reasoning load, producing measurable collapse points, degraded reference stability, and eventual factual or semantic failure.
- The series then expands Evans’ Law from a threshold model into a broader reliability framework. Later work introduces the idea of a reliability surface, accounting not only for context length but also for task rigidity, input complexity, output complexity, and verification burden. This moves the theory beyond “how long can a model reason before it fails?” toward a more operational question: under what conditions does a model become unreliable, and how can that degradation be detected, measured, governed, or mitigated?
- A major branch of the research examines multimodal degradation. These papers argue that failures become more complex when models must process and reconcile text, image, document structure, layout, and contextual inference at the same time. The multimodal work introduces the cross-modal degradation tax: the added reliability burden created when a model must maintain coherence across different modes of representation. In this framework, multimodal failure is not just hallucination plus vision error; it is an additional class of degradation caused by fractured alignment between modes.
- The agentic AI work extends the same empirical concern into systems that claim to plan, act, remember, decide, and coordinate over time. These papers question whether “agentic AI” exists in the strong sense often implied by vendors, or whether current systems are better understood as probabilistic language models wrapped in orchestration layers. The research identifies architectural risks in agentic systems, including compounding error, unstable task interpretation, poor recovery from drift, brittle tool use, and the illusion of autonomy where no reliable internal agency exists.
- The hallucination papers form another core layer of the series. Rather than treating hallucination as a generic content problem, they examine it as a mechanistic failure of coherence, reference, significance, and repair. The work argues that source-grounding, retrieval, and longer context can reduce some classes of error but do not eliminate hallucination because they do not solve the deeper architectural problem: the model still lacks a stable native mechanism for determining what must remain true, what must remain attached to whom or what, and what must not be overwritten by statistically plausible continuation.
- From there, the research develops the concept of missing primitives in contemporary transformer systems. The papers on strict semantic dominance and revocable semantic dominance argue that LLMs lack native mechanisms for enforcing durable meaning relationships. This becomes especially visible in proper noun failures, identity drift, source confusion, and semantic authority failures, where models mishandle names, entities, relationships, and claims even when the relevant information is present. The proper noun work shows that names are not a minor edge case but a stress test for whether a system can preserve identity and reference under pressure.
- The S-vector papers propose a possible architectural and operational response. The S-vector, or significance vector, is introduced as a missing dimension in current language-model representations: a way to encode not merely what tokens mean, but how much they matter within a task, document, system, or institutional context. The research argues that current models are “flat” with respect to significance. They can represent association, probability, and semantic proximity, but they do not reliably distinguish between disposable details and high-consequence anchors such as names, legal terms, medical facts, financial figures, authorship, institutional authority, or safety-critical instructions. The S-vector work explores how significance weighting could improve long-context stability, retrieval, enterprise AI, agentic systems, and governance.
- Taken together, the technical papers form a unified empirical theory of LLM failure and evolution. They show that hallucination, long-context degradation, multimodal collapse, proper noun failure, agentic instability, and semantic drift are not separate problems. They are related expressions of a deeper architectural limitation: current models predict continuation without a sufficiently stable internal mechanism for preserving significance, authority, identity, and verification across time.
- The final stage of the series begins to connect these architectural findings to public policy, institutional procurement, and AI sovereignty. If LLMs fail in predictable ways, and if those failures are amplified in government, infrastructure, healthcare, justice, immigration, education, and public administration, then AI governance cannot be limited to privacy, procurement, ethics, or domestic model ownership alone. The sovereignty work asks who controls the models, infrastructure, standards, data flows, evaluation systems, and decision layers increasingly embedded in public life. It argues that AI sovereignty is not simply a matter of building national models; it is a question of whether governments can understand, audit, govern, and retain authority over the systems on which they are becoming dependent.
- Across the full series, the work moves from observation to axiom, from axiom to measurement, from measurement to architecture, and from architecture to governance. Its central claim is that LLM reliability cannot be solved by scale alone. The future of trustworthy AI will depend on whether systems can be designed, evaluated, and governed around the failure modes they actually exhibit: degradation over reasoning length, instability across modalities, weak semantic authority, poor significance weighting, fragile agency, and the unresolved problem of knowing what must remain true.
- Evans’ Law: A Predictive Threshold for Long-Context Accuracy Collapse in Large Language Models
- Evans’ Law v4.1 (Extended)
- Evans’ Law: Scaling, Coherence, and Governance Implications V4.0
- Evans’ Law 5.0: Long-Context Degradation in Multimodal Models and the Cross-Modal Degradation Tax
- AI’s Unmeasured Reality: How Users Are Left Behind
- AI’s Accountability Gap: A Policy Blueprint for Policymakers
- Does Agentic AI Exist? v6.0
- The When Where and How of LLM Failures, Measured
- Why Hallucinations Happen: Fracture and Repair in Transformer Systems v1
- The S-Vector: Topographic Attention and the Architecture of Intelligence
- The Mechanistics of Hallucination Version 3.0
- Why the S-Vector Matters: The Missing Dimension in Enterprise AI — and What Companies Should Try
- Research Summary: A Unified Theory of LLM Evolution
- Two Missing Primitives in Contemporary Language Models: Strict Semantic Dominance and Revocable Semantic Dominance
- The Missing Key to True LLM Intelligence 3.0: An Operational Roadmap for the S Vector
- Beyond Content: Proper Nouns and Semantic Governance Failures in LLMs
- Why Agentic AI Is Problematic: The Architectural Risks
- Source-Grounding Does Not Prevent Hallucinations: A Controlled Replication Study of Google NotebookLM
- Coordination, Significance and Manifold Efficiency: A Path to Transformative Intelligence
- Significance Weighting in Large Language Models: Cross-Architecture Behavioral Evidence
- Proper Noun Failure: An Empirical Update on Evans’ Law
- Nudgment Signal Discernment Framework
- Whose AI Runs the Government?

