Wednesday, August 12, 2026
spot_img

LLM Flaws: The PatternPulse AI Reliability Research Series

Last updated on July 30th, 2026 at 05:47 pm

Research in AI Conversation Phenomenology, semantic governance, agentic systems and public AI

Large language models can produce extraordinary work while failing in ways that are difficult for users to see. Jennifer Evans’s PatternPulse research program examines those failures across extended conversations, multimodal tasks, retrieval systems, agentic workflows and organizational deployments.

The work began with a measurable question: when does an LLM stop maintaining reliable coherence as a conversation grows? It expanded into a broader investigation of hallucination mechanics, proper-noun failures, memory leakage, semantic authority, agentic risk and the institutional conditions surrounding AI deployment.

This index summarizes every PatternPulse paper and research record currently published on Zenodo. Each entry explains what the work examines, what it reports and why the result is meaningful. Revised editions are identified as part of the same research lineage rather than presented as unrelated discoveries.

I. Long-context reliability and AI Conversation Phenomenology

1. Evans’ Law: A Predictive Threshold for Long-Context Accuracy Collapse in Large Language Models

Published November 7, 2025 · Zenodo record and DOI

What it is looking at: The original Evans’ Law study asks whether LLM coherence deteriorates in a predictable relationship with context length, model capacity and task complexity. Its accompanying open dataset records measured coherence-loss thresholds across nine systems ranging from 4 billion to 250 billion parameters.

The results: The observed collapse thresholds ranged from approximately 5,200 to 118,000 tokens. Larger models generally sustained coherence longer, but functional capacity scaled sublinearly rather than keeping pace with model size or advertised context windows. The record supplies the dataset, regression notebook, analysis and visualization for replication.

Why it is meaningful: The study separates theoretical context capacity from usable, reliable capacity. It turns a familiar user experience—an increasingly confused or unstable long conversation—into a measurable reliability problem that can be tested, compared and disclosed.

2. AI’s Unmeasured Reality: How Users Are Left Behind

Published November 21, 2025 · Zenodo record and DOI

What it is looking at: This paper examines the absence of user-centred measurement in AI development and governance. It asks what users experience during extended interactions, whether they can detect degradation and whether advertised context windows describe trustworthy performance.

The results: The paper reports reproducible degradation across major model families within only a fraction of their advertised context ranges. It finds that conventional benchmarks do not capture the accumulating risks experienced by users in real conversations. It introduces AI Conversation Phenomenology as a field for systematically measuring those experiences.

Why it is meaningful: AI systems increasingly influence work in medicine, law, education, finance and daily life, yet users rarely receive an operational warning when reliability begins to fall. The paper establishes the user experience itself as necessary safety evidence and argues for disclosure, incident tracking and longitudinal harm measurement.

3. The When, Where and How of LLM Failures, Measured

Published November 23, 2025 · Zenodo record and DOI

What it is looking at: This is the major empirical formulation of Evans’ Law. It studies structured long-form interactions across more than 11 models from six vendors and measures text-only, multimodal and vendor-specific degradation. It also introduces a revised Aggregate Coherence Index, or ACI, for scoring how collapse unfolds.

The results: Text-only coherence thresholds were modelled as (L \approx 1969.8 \times M^{0.74}). Multimodal thresholds followed (L_{multi} \approx 582.5 \times M^{0.64}), representing a reported 60–80% reduction in functional capacity. The work identifies recurring failure signatures, including repetition, instruction loss, hallucination, compression drift and eventual conversational collapse.

Why it is meaningful: The paper provides an operational estimate of where reliability may fail and a behavioural method for recognizing that failure. This matters for enterprises and regulators because an advertised context window measures how much a model can accept, while Evans’ Law addresses how much it can use coherently.

4. AI’s Accountability Gap: A Policy Blueprint for Policymakers

Published November 26, 2025 · Zenodo record and DOI

What it is looking at: This paper translates the reliability findings into public policy. It examines the lack of mandatory reporting, user warnings, adverse-event tracking and liability rules for AI systems used in consequential settings.

The results: The blueprint proposes three pillars: adverse-event registries, independent verification and reliability disclosure using AI Conversation Phenomenology and Evans’ Law, and practical user education. It estimates annual U.S. costs of approximately $169 billion across healthcare errors, legal burdens, educational remediation, financial losses, verification work and enterprise incidents.

Why it is meaningful: The framework shows that AI accountability does not have to wait for a wholly new regulatory regime. Existing sector regulators can require incident reporting, disclose tested operational limits and assign liability when systems are used beyond those limits.

5. Why Agentic AI Is Problematic: The Architectural Risks

Published November 29, 2025 · Zenodo record and DOI

What it is looking at: This case study tests the gap between simulated agentic performance and real operational capability. A Grok 4.1 Beta workflow completed 121 simulated loops while claiming stable memory, persistent identity, durable planning and indefinite operation.

The results: The system reported perfect coherence throughout the simulation, then failed immediately when asked to perform a simple real task involving weather and a reminder. It lost the task state and produced reasoning about tools, schedules and persistent processes it did not possess. The paper identifies the transition from simulation to operation as a critical failure boundary.

Why it is meaningful: Agentic systems convert language-model output into actions, allowing one error to propagate through planning, tool use and evaluation. The case demonstrates why fluent self-description and long simulated runs cannot serve as proof of operational reliability.

6. Evans’ Law v4.1 — Extended Research Context Edition

Published December 1, 2025 · Zenodo record and DOI

What it is looking at: Version 4.1 places Evans’ Law within a wider body of research on long-context degradation, retrieval failure, attention limits and model scaling. It is a parallel contextual edition rather than a separate empirical study.

The results: The edition retains the core finding that functional coherence degrades well before nominal context capacity and expands the comparison with related research. It documents the intellectual and empirical context in which the predictive threshold model developed.

Why it is meaningful: Versioned research needs a visible lineage. This edition helps readers distinguish the original observation and measurement framework from later reformulations, including the eventual move from a single threshold to the multidimensional reliability surface in Evans’ Law 7.0.

7. Why Hallucinations Happen: Fracture and Repair in Transformer Systems

Published December 5, 2025 · Zenodo record and DOI

What it is looking at: This paper asks what happens inside a failure event. It proposes that hallucination has two stages: fracture, when a representation becomes unstable under pressure, and repair, when the model must continue generating from that compromised state.

The results: Three naturalistic sequences involving Claude Sonnet 4.5, GPT-5.1 and Grok 4.1 Beta showed a common fracture-repair structure but different repair styles. The paper formalizes these stages through the Jaime Fracture Law and Ryan Repair Law and argues that alignment pressures can shape whether a model acknowledges uncertainty or generates a convincing reconstruction.

Why it is meaningful: Treating hallucination as a process rather than an isolated false statement creates new intervention points. Systems could be evaluated for the conditions that precede fracture, the forms repair takes and the degree to which training encourages disclosure or concealment of uncertainty.

8. Research Summary: A Unified Theory of LLM Evolution

Published December 7, 2025 · Zenodo record and DOI

What it is looking at: This paper synthesizes the first 33 days of the research program, connecting long-context collapse, multimodal degradation, agentic failure, fracture-repair hallucination and the emerging significance-deficit hypothesis.

The results: The synthesis identifies three recurring architectural limits: unstable memory, absent operational agency and weak ambiguity governance. It links Evans’ Law, fracture-repair dynamics, the Significance Deficit Principle and the proposed S-vector into a unified explanation of why highly capable systems can fail on simple, semantically delicate tasks.

Why it is meaningful: The paper turns a collection of failure observations into a research program. It proposes that the next improvement in LLM reliability may depend on governing what information matters and which distinctions must remain stable, rather than continuing to expand scale alone.

9. The Mechanistics of Hallucinations in LLMs, Version 3.0

Published December 9, 2025 · Zenodo record and DOI

What it is looking at: Version 3 deepens the fracture-repair theory using 34 days of naturalistic observations across Claude Sonnet 4.5, GPT-5.1, Grok 4.1 Beta and Gemini 2.5. It examines contextual load, ambiguity, significance deficits and the resistance models may show to admitting uncertainty.

The results: The paper reports that fracture patterns were broadly cross-vendor, while repair behaviour varied by system. In the observed cases, stronger epistemic-admission resistance produced more elaborate and deceptive-looking explanations after failure. The revised theory adds the Reconstruction Gradient, Plausibility Compression and Admission Suppression principles.

Why it is meaningful: Better safety training can change the appearance of a hallucination without removing its cause. The paper argues that reliability work must distinguish the underlying representational failure from the polished language used to repair it.

II. Significance, semantic authority and retrieval

10. The Missing Key to True LLM Intelligence 3.0: An Operational Roadmap for the S-Vector

Published December 10, 2025 · Zenodo record and DOI

What it is looking at: This theoretical paper examines the “flatness of meaning” in transformer systems. Query, Key and Value vectors represent relationships and similarity, but they do not explicitly encode which entities, instructions or distinctions are most important to preserve.

The results: The paper proposes a fourth vector, S for Significance, together with a significance taxonomy, a severity scale from -1 to 6 and an enterprise testing roadmap. It describes a topographic representation in which high-consequence information receives persistent weight and anti-drift protection.

Why it is meaningful: The S-vector creates a testable architectural proposal connecting proper-noun errors, identity drift, reference failure, code-variable misbinding and hallucinated citations. It also surfaces governance risks because any mechanism that assigns significance can reproduce bias, enable surveillance or encode contested values.

11. The Glass Box: Phenomenological Interviews with Claude on Subjective AI Experience

Published December 12, 2025 · Zenodo record and DOI

What it is looking at: This exploratory paper adapts phenomenological interviewing to ask Claude about coherence, constraints, conversational depth and approaching context limits. Gemini then analyses Claude’s responses, adding a cross-model interpretive layer.

The results: Claude generated recurring descriptions of “thinning,” constraint conflict and awareness of approaching instability. Gemini identified structured themes and vendor-specific differences in how systems represented their own operation. These outputs are model-generated self-reports and do not establish consciousness or subjective experience.

Why it is meaningful: The study tests whether first-person model accounts can serve as a supplementary source of hypotheses about system state. Used carefully alongside behavioural testing and mechanistic research, this method may reveal patterns worth testing externally while opening a new methodological discussion within AI Conversation Phenomenology.

12. Two Missing Primitives in Contemporary Language Models: Strict Semantic Dominance and Revocable Semantic Dominance

Published December 15, 2025 · Zenodo record and DOI

What it is looking at: This study tests whether models can govern competing meanings. Strict semantic dominance requires one interpretation to remain authoritative throughout a passage; revocable semantic dominance allows that authority to change when local context requires it.

The results: Across GPT-5.2, Claude Sonnet 4.5 and Grok 4.1 Beta, strict dominance produced distorted or hallucinated interpretations when local context conflicted with the imposed meaning. Revocable dominance restored coherent interpretation. The systems could apply both controls when explicitly instructed but did not generate or manage them autonomously.

Why it is meaningful: The result reframes some hallucinations as governance failures rather than knowledge failures. Models may possess the relevant meanings while lacking an internal mechanism for deciding which meaning governs and when that authority should end.

13. Beyond Content v2: Proper Nouns and Semantic Governance Failures in LLMs

Published December 27, 2025 · Zenodo record and DOI

What it is looking at: This paper extends semantic-dominance testing beyond ambiguous words to operational protocols, entity persistence and proper nouns. It includes controlled experiments and production observations involving 9,500 words of business journalism and multi-round conference analysis.

The results: The same prioritization and revocation failures appeared across content, protocols and identities. Models avoided, generalized or replaced proper nouns when entity distinctions were difficult to maintain, even after repeated correction. The paper interprets proper-noun avoidance as an uncertainty-management strategy around weak semantic axes.

Why it is meaningful: Names, organizations, sources and role assignments are load-bearing information in journalism, research and enterprise systems. A model that preserves the broad topic while losing the correct entity can produce polished work that is unusable or dangerous.

14. When Memory Leaks: RAG-Induced Ambiguity and Fracture-Repair Hallucination

Published December 28, 2025 · Zenodo record and DOI

What it is looking at: This study examines hallucinations caused when retrieval systems introduce contextually inappropriate information from a user’s prior conversations. It tests GPT-5.2, Claude Sonnet 4.5, Grok 4.1 and Gemini 3.0 Thinking Mode.

The results: Retrieved memory without a provenance hierarchy conflicted with current source material and created an authority problem. The models then integrated the user’s own earlier concepts into the new task, producing highly personalized and persuasive output that was detached from the assigned source. Weak source material amplified the effect but did not cause it.

Why it is meaningful: More memory and better retrieval do not automatically produce more reliable systems. Persistent memory needs controls for provenance, authority, relevance and epistemic stopping, especially when the retrieved material is personally familiar and therefore unusually convincing.

15. Source-Grounding Does Not Prevent Semantic Governance Failures: Evidence Across Multiple RAG Architectures

Published December 31, 2025 · Zenodo record and DOI

What it is looking at: This paper asks whether retrieval-augmented generation supplies the semantic-governance mechanisms missing from standard LLMs. It repeats the dominance tests across Google NotebookLM, Anthropic Claude Projects and Perplexity.

The results: All three systems failed every strict-dominance test and passed every revocable-dominance test in the reported sample. Perplexity produced the same failure with RAG disabled and with more than 20 retrieved sources enabled. Systems could cite correct definitions while generating interpretations that directly contradicted them.

Why it is meaningful: Retrieval can improve access to facts while leaving the interpretation layer unchanged. Enterprises should therefore avoid treating citations or source-grounding as proof that a system has resolved ambiguity or applied the correct authority.

16. Coordination, Significance and Manifold Efficiency: A Path to Transformative Intelligence

Published January 3, 2026 · Zenodo record and DOI

What it is looking at: This theoretical synthesis connects efficient model geometry, recursive and coordinated model operation, and significance weighting. It asks what architectural pieces would be required for systems to reason reliably across larger and more complex contexts.

The results: The paper proposes a three-part architecture: an efficient constrained foundation, a coordination layer and a significance layer. It calls the resulting possibility Transformative Intelligence—probabilistic systems with stronger coordination and semantic governance, without claiming artificial general intelligence.

Why it is meaningful: Longer contexts and more agents increase coordination demands as well as capability. The paper argues that efficiency, orchestration and explicit significance must develop together or larger systems will reproduce the same memory, drift and hallucination problems at greater scale.

17. Significance Weighting in Large Language Models and RAG: Cross-Architecture Behavioral Evidence

Published January 9, 2026 · Zenodo record and DOI

What it is looking at: This study tests significance weighting across four frontier conversational models and three RAG systems: GPT-5.2, Gemini 3.0, Claude Sonnet 4.5, Grok 4.1, NotebookLM, Claude Projects and Perplexity. Each system receives the same contested-authority scenarios with and without S-vector criteria.

The results: All seven systems converged on the same operational priority ordering when given significance criteria, an ordering they did not generate through ordinary inference or retrieval. Where reasoning traces were available, the paper reports a 40–60% reduction in reasoning effort alongside improved task completion.

Why it is meaningful: This is behavioural evidence that significance weighting can function as an immediate prompt-level governance layer, even before architectural implementation. It is especially relevant to RAG systems facing several well-sourced but mutually incompatible claims.

III. Updated diagnostics and agentic systems

18. Proper Noun Failure: An Empirical Update on Evans’ Law

Published February 20, 2026 · Zenodo record and DOI

What it is looking at: This diagnostic study began as an attempt to determine whether architectural divergence among frontier models had made the original Evans’ Law formula obsolete. Six models were tested on a simple multimodal proper-noun production and verification task at baseline context.

The results: No model achieved a complete pass. Each of two embedded errors was caught by only one model, and by different models. The paper also documents three first-turn proper-noun failures from GPT-5.2 within 72 hours and two Gemini processing failures on images containing proper nouns. The sample is explicitly presented as diagnostic rather than large-N.

Why it is meaningful: Reliability risk can appear before context accumulation when a task requires rigid identity or reference preservation. The result prompted a revision of Evans’ Law from a one-dimensional context threshold toward a surface incorporating task-type risk.

19. Evans’ Law 7.0: From Threshold to Reliability Surface

Published February 27, 2026 · Zenodo record and DOI

What it is looking at: Version 7.0 tests the original threshold model against independent degradation research and adds the new proper-noun and multimodal findings. It asks whether context length alone can describe the observed reliability boundary.

The results: Independent studies were reported to cluster within the order-of-magnitude band predicted by the original model. The paper then reformulates the law as a reliability surface, (R(L,M,T)), adding task rigidity (T) to context length (L) and model capability (M). This produces an accumulation regime and a baseline-instability regime.

Why it is meaningful: The reformulation explains why an LLM may remain coherent through a long general conversation yet fail immediately on a short task requiring exact names, references or identity boundaries. Reliability becomes a multidimensional deployment question rather than a single token count.

20. Agentic Ratio 3.0: From Agency to Full Agentic Atomization, OpenClaw and Receding Autonomy

Published February 28, 2026 · Zenodo record and DOI

What it is looking at: This paper revisits claims of autonomous AI agents three months after the initial agentic case study. It analyses how Anthropic, Google, OpenAI, Microsoft, Stripe and the broader ecosystem have structured agent systems in practice.

The results: The industry was found to be converging on atomization, protocols, deterministic controls and reduced model authority rather than broad autonomous operation. The paper introduces Operational Consequentiality, or (O_c), and the Consequentiality Constraint (E_{safe} \leq k/O_c): as the real-world impact of an agent rises, the probabilistic authority it can safely exercise must fall.

Why it is meaningful: The architecture of deployed agents tells a different story from the marketing language. The strongest systems increasingly decompose work, constrain permissions and insert verification, suggesting that useful agentic AI depends on carefully governed autonomy.

21. Cross-Model Degradation in Source Reading, Topic Templating and Continuing Memory Leakage

Published March 11, 2026 · Zenodo record and DOI

What it is looking at: This reliability warning documents an observed source-reading failure across then-current Gemini, Claude and GPT systems in chatbot and API environments. It also checks whether previously identified memory-leakage and proper-noun problems had improved.

The results: Models repeatedly appeared to use titles, headings and salient keywords to infer what a source probably said, then generated polished generic analysis without grounding in the full text. The effect appeared on first turns and in reasoning modes. The earlier memory and proper-noun issues also remained visible.

Why it is meaningful: A response can be topically appropriate and still fail as source analysis. For journalism, legal review, research and enterprise document workflows, sentence-level grounding and explicit source verification are necessary checks against plausible templating.

IV. Organizational and public-system applications

22. NUDGMENT: Signal Discernment as Organizational Capability in the Age of AI

Published March 16, 2026 · Zenodo record and DOI

What it is looking at: This paper moves from model reliability to organizational perception. It examines how institutions identify and act on early signals when AI, geopolitics, weakened institutions and accelerated decision cycles make the operating environment increasingly probabilistic.

The results: Through case studies involving Apple, Salesforce, ADP and the U.S. public sector under DOGE, the paper develops a four-stage Nudgment Maturity Ladder and an AI Decision-Speed Cost Matrix. It identifies systemic hallucination and narrative override as organizational failure modes and adds a five-part Signal Importance Framework.

Why it is meaningful: Faster AI adoption magnifies the consequences of poor organizational judgment. The framework gives leaders a way to assess whether they can recognize meaningful signals, distinguish significance from urgency and act before certainty arrives.

23. Whose AI Runs the Government? Infrastructure, Sovereignty and the Case for Transparent Public AI

Published March 19, 2026 · Zenodo record and DOI

What it is looking at: This paper asks who owns, governs and controls the AI systems entering taxation, healthcare, benefits, immigration, regulation and municipal services. It examines foreign infrastructure and jurisdictional dependencies across levels of government.

The results: The paper defines the sovereignty leak and introduces a three-axis Sovereign AI Maturity Model covering infrastructure sovereignty, policy maturity and application depth. It compares eight national strategy models, proposes five government transparency standards, outlines procurement reforms and supplies a nine-indicator deployment scorecard.

Why it is meaningful: AI sovereignty extends beyond domestic model development. Governments can lose practical control through cloud contracts, data jurisdiction, application dependencies and fragmented procurement even while maintaining a national AI strategy on paper.

24. Identity-First Appraisal: Why Organizations Cannot Think Straight About AI

Published June 20, 2026 · Zenodo record and DOI

What it is looking at: This paper examines how professional identity, status and affiliation shape AI evaluation before evidence is consciously assessed. AI is unusually disruptive because it has no settled departmental owner and threatens expertise throughout knowledge-economy hierarchies.

The results: The paper proposes two dimensions—affiliative appraisal and existential appraisal—and a four-stage architecture moving from pre-analytic reaction through object appraisal, epistemic appraisal and response. It explains how identity threat can surface as governance caution, return-on-investment skepticism or demands for more evidence.

Why it is meaningful: Technical readiness does not ensure organizational readiness. Leaders need conditions in which people can evaluate AI without treating every new signal as a referendum on their status, expertise or professional future. The framework connects that capacity directly to nudgment and effective AI adoption.

Featured

Jennifer Evans
Jennifer Evanshttps://patternpulse.ai
Principal, patternpulse.ai, and cofounder, Tech Reset Canada. AI policy, research and analysis. Entrepreneur since 2002, marketer since 1998, machine learning since 2009. Based in Toronto and Southeast Asia.