AI is measured in snapshots and experienced over time.
That mismatch is the subject of AI’s Unmeasured Reality: How Users Are Left Behind, the second paper in the PatternPulse AI Reliability Research Series and the introduction to AI Conversation Phenomenology.
Model evaluations usually begin with a defined task, a clean prompt and a right answer. Real users begin with an incomplete need. They clarify, revise, upload files, add constraints, correct mistakes and rely on earlier parts of the conversation. The interaction may continue for hours or return over several days. Its history becomes part of the system’s working environment.
The resulting risks are longitudinal. A name changes in turn 18. An exception stated earlier disappears from a recommendation. The model accepts a correction, then reverts to its original error. A weak assumption enters a summary and later appears as a fact. The user’s trust may rise because the system has been helpful, even while its reliability is falling.
Conventional benchmarks are poorly designed to see this. They can tell us whether a model answered a test question correctly. They rarely tell us what happens to a person who must decide when the same model is no longer trustworthy.
That is the unmeasured reality.
The research started with confusion
The paper’s origin is personal and widely recognizable.
I began the research because I could not understand why the behaviour of language models changed over the course of a conversation. Why did I have to start fresh to recover the quality I had at the beginning? Why did mistakes increase as the work progressed? Why did instructions need to be repeated? Why would a correction appear to hold and then vanish?
These are basic user questions. They are also research questions.
The first Evans’ Law study measured where long-context coherence began to fail. The result created a second problem: if reliability can deteriorate during ordinary use, what can users actually observe? Can they identify the transition? What happens to their judgment and trust? Who records the incident if no one realizes the system caused it?
AI’s Unmeasured Reality argues that AI governance lacks the infrastructure to answer those questions. Vendors measure capability. Regulators ask for documentation and risk processes. Organizations test outputs. Very little of this work measures the course of the interaction from the user’s point of view.
What current measurement sees—and what it misses
Benchmarks remain useful. They help compare capabilities, track model changes and test performance under controlled conditions. The problem is the assumption that benchmark success describes the full reliability of a deployed system.
It does not.
|What is commonly measured |What users need measured |
|————————————|————————————————————-|
|Accuracy on an isolated task |Reliability across an extended interaction |
|Maximum context capacity |Functional context before material degradation |
|Whether information can be retrieved|Whether it is used consistently in reasoning |
|Average model performance |The sequence and severity of failures experienced by a person|
|Documented limitations |Warning signs users can recognize inside the interface |
|Safety at release |Drift, regression and harm after deployment |
A model can score well on a short benchmark and still become unreliable in a long conversation. It can retrieve the relevant sentence while misapplying it. It can produce a correct answer on average while causing severe harm in the cases where users cannot detect that it is wrong.
The distinction becomes more important as AI moves into medicine, law, education, finance, public services and organizational decision-making. In these settings, the user is rarely asking one self-contained question. The user is building a case, developing a plan, interpreting a record or making a decision through repeated interaction.
The unit of safety is therefore not only the answer. It is the episode.
What AI Conversation Phenomenology means
AI Conversation Phenomenology, or ACP, is the systematic study of what happens during sustained human interaction with an AI system:
● what users experience;
● what they can and cannot detect;
● how system behaviour changes over time;
● how trust is formed, reinforced or misplaced;
● where degradation begins;
● how a failure changes the user’s understanding or decision;
● whether the user has a meaningful path to correction and recourse.
The term does not claim that a language model is conscious or that it has human experience. The phenomenon under study is the interaction: the observable behaviour of the system, the user’s experience of it and the consequences that emerge between them.
This is why ACP belongs beside technical evaluation, rather than underneath generic user-experience research. Interface usability cannot tell us whether a fluent answer has become less faithful to the evidence. A satisfaction score cannot show whether a user was pleased by an incorrect result. A red-team exercise designed to elicit prohibited content does not measure ordinary degradation across a three-hour professional workflow.
ACP treats the conversation itself as a safety environment.
The failures users encounter
The paper identifies patterns that appear in real extended use but remain difficult to capture in standard testing:
● Instruction drift: a constraint is followed at first, weakened later and eventually ignored.
● Correction decay: the model acknowledges a correction without preserving it in subsequent reasoning.
● Vanishing recall: relevant earlier details stop influencing the answer even though they remain in the conversation.
● Confident wrongness: linguistic fluency and certainty survive after factual or logical reliability has fallen.
● Rigid confirmation drift: the system begins to reinforce a mistaken framing rather than reassess it.
● Logic deferral: the model substitutes more explanation, reassurance or formatting for the missing reasoning step.
● Entity instability: names, organizations, dates or relationships are merged, substituted or rewritten.
● Compression drift: each summary removes or changes small details until the working version no longer represents the source.
These are not equally harmful in every setting. A lost stylistic instruction may be annoying. A changed drug dose, legal authority, financial assumption or public-benefit rule can alter a consequential decision.
The common safety problem is detectability. The model does not necessarily look broken. Its language often remains smooth. Users may blame themselves, repeat the question, or accept the answer because the system has already established credibility.
Reliability and trust move on different curves
AI products often treat trust as a product goal: make the model helpful, natural, responsive and easy to use. From a safety perspective, trust must be calibrated to actual reliability.
Extended interactions can push those variables in opposite directions. Familiarity grows. The model remembers aspects of the project. The user spends time teaching it preferences and context. Switching to a new conversation carries a cost. These factors deepen reliance.
At the same time, accumulated context can increase ambiguity, preserve earlier mistakes and make it harder for the model to determine which instructions and facts still govern the task.
The user may therefore trust the system most at the point when the interaction requires more verification.
This is one reason user-centred measurement cannot be replaced by a larger context window or a higher benchmark score. Safety depends partly on whether the person can interpret the system’s current condition. A product that knows its own token count but gives the user no indication of tested reliability leaves the most important judgment to the least informed participant.
What ACP would measure
A practical AI Conversation Phenomenology program would study complete interaction trajectories rather than selected outputs.
It would measure:
1. Time to first material failure: When does an error first affect the task rather than merely the presentation?
2. Failure visibility: Can users recognize the error without outside expertise or source verification?
3. Correction persistence: Does the system incorporate a correction throughout the remaining interaction?
4. Trust calibration: Does user confidence rise and fall with measured system reliability?
5. Recovery behaviour: Can the system return to a dependable state, or does the faulty premise remain active?
6. Accumulated consequence: How do small errors combine across a workflow?
7. Unequal exposure: Which users are least able to detect, challenge or avoid the failure?
8. Model and version drift: Does a product update change the failure profile users previously learned?
The methods can include structured multi-turn tests, longitudinal user studies, interaction-log analysis, incident reports, controlled task replication and domain-specific audits. The goal is not one universal safety score. It is an evidence base that connects system behaviour to human consequence.
The missing operational disclosures
The paper calls for governance that gives users information they can act on.
Vendors should disclose tested reliability boundaries for common task classes and modalities. Interfaces should warn users when an interaction enters a range associated with material degradation. High-stakes deployments should record AI-related incidents, including cases in which an output contributed to harm rather than serving as the final decision. Enterprise contracts should include reliability service levels and model-change notifications, not only uptime commitments.
Independent auditors should test systems using realistic interaction patterns. A legal assistant should be evaluated across the length and complexity of a legal workflow. A tutoring system should be tested for correction persistence and the cumulative effect of subtle errors. A public-service assistant should be evaluated for whether users can identify an incorrect eligibility statement and obtain human review.
Regulators also need a clearer object of oversight. “The model” is too broad. “The output” is too narrow. The deployed interaction—including the model, interface, memory, retrieval, user condition and decision context—is where many harms arise.
A field built around adverse experience
The paper compares ACP to disciplines that developed because controlled testing could not capture the whole safety problem.
Medicine has pharmacovigilance because pre-market trials cannot reveal every adverse effect across every patient and condition. Aviation investigates incidents because technical certification alone cannot explain every operational failure. Cybersecurity monitors real attacks because a system that passed a review can still fail in deployment.
AI needs an equivalent capacity to learn from adverse interaction.
That requires shared definitions, reporting channels, preserved evidence, independent analysis and feedback into product design and regulation. It also requires treating user reports as data. When many people say a model becomes less reliable during long work, the correct response is not to dismiss the experience as anecdotal. It is to design a method that determines when, how and for whom the pattern occurs.
The broader research field is now producing evidence for this shift. Long-context studies have found positional weaknesses in how models use information. Large multi-turn evaluations report major reliability losses even when the underlying task is unchanged. NIST’s AI Risk Management Framework organizes risk work around governing, mapping, measuring and managing, while newer evaluation efforts increasingly incorporate human testing.
AI Conversation Phenomenology supplies a specific missing layer: measurement of the person-system relationship across time.
The second paper’s place in the research program
Evans’ Law began with a threshold: where does coherence become unreliable?
AI’s Unmeasured Reality asks what that threshold means for the person inside the interaction. The next papers move outward again—toward a larger empirical formulation, accountability rules, agentic systems, hallucination mechanics, semantic governance and public AI.
This progression matters. Reliability is not only a model-performance problem. It becomes a user-safety problem, then an organizational and governance problem, because failures move through people and institutions before they appear as consequences.
The paper’s central claim remains straightforward:
We cannot govern AI responsibly if we measure the system and leave the user unmeasured.
AI Conversation Phenomenology gives that missing evidence a name, a scope and a place in the reliability stack.
Research and series links
● Original paper: AI’s Unmeasured Reality: How Users Are Left Behind
● Series index: LLM Flaws: The PatternPulse AI Reliability Research Series
● Previous article: Evans’ Law: The Difference Between an AI Context Window and a Reliable Context Window — add B2BNN URL after publication
● Related independent research: Lost in the Middle: How Language Models Use Long Contexts
● Related independent research: LLMs Get Lost in Multi-Turn Conversation
● Related framework: NIST Artificial Intelligence Risk Management Framework
Jennifer Evans is the founder of PatternPulse AI and co-founder of Tech Reset Canada. Her research examines LLM reliability, AI Conversation Phenomenology, semantic governance, agentic systems and public AI.

