Thursday, July 30, 2026
spot_img

AI’s Accountability Gap: When AI Fails, Who Has to Report It?

An AI system can influence a medical, legal, financial or educational decision without creating a public record when it fails.

That absence is the accountability gap.

AI’s Accountability Gap: A Policy Blueprint for Policymakers, the fourth paper in the PatternPulse AI Reliability Research Series, initially commissioned by the office of a US state lawmaker, takes the measurable failures documented in the first three papers and asks what governments should do with the evidence.

Evans’ Law provides an estimate of where long-context coherence may begin to degrade. AI Conversation Phenomenology supplies a method for observing what users experience across an interaction. The expanded empirical work documents recurring signatures, including instruction loss, hallucination, compression drift and conversational collapse.

The policy problem is that these events can occur in consequential settings while remaining outside any common reporting system.

A patient may receive incorrect health guidance. A lawyer may submit a fabricated authority. A student may be taught an error through repeated tutoring sessions. An employee may spend hours verifying a system that was sold as an efficiency tool. An institution may discover the problem, resolve the individual case and generate no evidence that regulators, researchers or other users can see.

The same failure can then recur elsewhere as a new anecdote.

The blueprint proposes a different approach: treat AI failure as an observable public-safety event that can be reported, measured, disclosed and assigned to responsible actors.

Other high-consequence systems learn from failure

Medicine, aviation, automotive safety and cybersecurity all developed mechanisms for learning after deployment.

Clinical trials cannot reveal every adverse drug reaction, so healthcare systems collect post-market reports. Aircraft certification does not eliminate operational incidents, so aviation authorities investigate near misses and crashes. Vehicle testing does not expose every defect, so manufacturers and regulators track complaints, recalls and injury patterns. Cybersecurity programs gather vulnerability and incident data because systems change after release and attackers find conditions that designers did not anticipate.

These systems differ in law and method. They share a basic premise:

A failure that happens in use is evidence.

AI lacks an equivalent, broadly required infrastructure for adverse interaction. Vendors publish selected evaluations and system cards. Organizations run internal reviews. Users file support tickets or post screenshots. Insurers, courts and regulators may encounter the consequences later.

The evidence remains fragmented. Terminology varies. Model versions change. The product provider, deployer and professional user may each hold one part of the incident. No participant has a complete view, and no common mechanism necessarily assembles one.

The result is a governance system that reacts to visible scandals while missing the recurring lower-level events that reveal a pattern.

What the accountability gap contains

The paper identifies several connected absences:

● no standard definition of a reportable AI adverse event;

● no common registry for incidents and near misses;

● no mandatory disclosure of tested reliability limits;

● no operational warning when a user crosses a known degradation range;

● limited independent verification of vendor claims;

● weak preservation of model, context and configuration evidence;

● unclear responsibility when several organizations contributed to the deployment;

● few practical tools that help users recognize a degrading interaction;

● uncertain liability when a system is used beyond a known or reasonably testable limit.

These are infrastructure problems. Ethics principles alone cannot produce incident data. A general transparency commitment does not tell a user that a specific model version has entered an untested operating range. A broad instruction to maintain human oversight does not establish what the human must check or how much evidence they need.

The blueprint organizes the response around three enforceable pillars.

Pillar one: adverse-event registries

The first pillar is a reporting system for AI incidents, modelled on the registries used in healthcare and other safety-critical sectors.

The purpose is pattern detection.

One incorrect response may reflect a prompt problem, missing information, a deployment error or model degradation. Hundreds of reports showing the same failure signature in the same model version, task class or context range create a regulatory signal. A near miss can be valuable for the same reason: it shows where harm almost occurred and which control prevented it.

A useful AI adverse-event report would capture:

1. the model, product and available version identifier;

2. the organization or service through which it was deployed;

3. the task and domain;

4. the approximate interaction length and number of turns;

5. the modalities, files, retrieval systems and tools involved;

6. the instruction or fact that failed;

7. the observable failure signature;

8. whether the user recognized the problem;

9. whether a correction worked and persisted;

10. the actual or potential consequence;

11. the action taken by the vendor or deployer.

The report should preserve privacy and protect legitimate confidential information. Those requirements affect the design of the registry. They do not remove the need to collect the safety signal.

Reporting also needs a threshold that captures material events without turning every poor sentence into a regulatory filing. A practical definition would focus on an output, action or interaction failure that caused harm, materially changed a consequential decision or created a credible risk that required intervention.

The system should accept evidence from several points in the deployment chain: vendors, organizations, regulated professionals, auditors and users. Different participants see different parts of the same event.

The registry then turns distributed experience into a shared evidence base.

Pillar two: independent verification and reliability disclosure

Incident data describes what happened in the field. Independent testing establishes the operating limits that should have been disclosed before deployment.

The second pillar applies AI Conversation Phenomenology and Evans’ Law as verification tools. Evaluators test complete interaction trajectories across the task lengths, modalities and conditions users will actually encounter.

A meaningful reliability disclosure would identify:

● the model and configuration tested;

● the task classes included;

● text-only and multimodal results reported separately;

● the tested range across context length and conversational turns;

● the point of first material degradation;

● the failure signatures observed;

● correction and recovery behaviour;

● known conditions outside the tested range;

● the date of evaluation and the update that triggers retesting.

This creates an operational label rather than a general claim of safety.

A context window of one million tokens may remain the technical maximum. The disclosure could state that coherent performance was independently validated to a much smaller range for a specific legal, medical, financial or educational task. Users and deployers would then know where the evidence ends.

The warning inside the product should reflect that boundary. When an interaction approaches or passes the validated range, the system can tell the user that reliability has not been established for the remaining session. Higher-consequence uses can require a stronger response: save verified state, begin a new session, reintroduce authoritative facts and obtain human review before continuing.

Independent verification also creates a common basis for procurement and enforcement. A buyer can compare systems on dependable operating range. A regulator can examine whether the provider disclosed what it knew. A court or insurer can distinguish an unforeseeable event from a deployment that continued past a documented limit.

Pillar three: practical user education

The third pillar is user education designed around observable failure.

Generic AI literacy often tells people that models can make mistakes and that important facts should be checked. Those statements are true and insufficient. Users need to recognize the conditions in which verification demand is rising.

Practical education would teach people to watch for:

● an earlier instruction disappearing;

● a correction holding for one turn and then decaying;

● a name, date, amount or relationship changing;

● a summary becoming less faithful after each compression;

● unsupported specificity entering an answer;

● the model repeating completed work;

● a polished response that cannot show how its claims connect to the source;

● a long session in which the cost of checking has begun to exceed the value of continuing.

Users also need a response protocol:

1. stop treating the current conversation as an authoritative record;

2. preserve the source material and verified decisions outside the chat;

3. check the disputed claim against a primary source;

4. begin a new interaction with a verified state summary;

5. use human or domain-expert review for consequential decisions;

6. report material incidents through the available channel.

Education gives people an immediate form of protection while reporting and disclosure systems develop. It does not transfer the safety burden from vendors and deployers to individual users.

A person cannot compensate for an undisclosed reliability limit they have no way to measure.

The $169 billion estimate

The paper estimates that AI-related failures and verification burdens could cost the United States approximately US$169 billion per year across:

● healthcare errors;

● legal-system burdens;

● educational remediation;

● financial losses;

● human verification work;

● enterprise incidents.

The number should be presented for what it is: a preliminary modelled estimate from the paper.

It is not a settled national account of observed AI harm. Many AI incidents are not identified, causality is difficult to isolate and the categories can contain substantial uncertainty. The estimate brings dispersed forms of cost into one frame and shows why measurement matters. It should not be cited as though a national reporting system has already counted US$169 billion in confirmed losses.

The verification category is particularly important for business. AI can produce an apparent productivity gain while transferring work into checking, correction, rework and incident response. Those costs rarely appear in capability benchmarks. They are absorbed by employees, customers, professionals and public institutions.

An adverse-event system would improve the estimate over time. Reliable public data could replace broad assumptions with observed frequency, severity, domain and model-version evidence.

Accountability needs a visible operating limit

Liability becomes difficult when no one has defined the conditions under which a system was dependable.

A vendor can point to the deployer’s implementation. The deployer can point to the professional user. The user can point to the model’s fluent answer. Each participant can claim that another actor made the final decision.

Reliability disclosure creates facts that clarify responsibility.

The policy question can then examine concrete conduct:

● Did the vendor test and disclose material limitations?

● Did the product warn the user when it moved beyond the validated range?

● Did the deploying organization operate the system on a task that had not been tested?

● Did the institution preserve and report a known incident?

● Did a regulated professional apply the required verification?

● Did a model update materially change the failure profile without notice?

This does not predetermine the legal answer. It gives regulators, courts and insurers evidence for assigning it.

The framework also recognizes that responsibility should follow control. A user cannot be expected to know an undisclosed model route or hidden system instruction. A vendor cannot control every downstream use. A deploying organization chooses the workflow, permissions, review process and consequences attached to the output. A workable liability system distinguishes these roles.

Existing regulators can begin

The blueprint does not require every jurisdiction to create a single new AI regulator before action begins.

Existing authorities already oversee consequential conduct in health, financial services, consumer protection, education, professional services, employment and public administration. They can adapt familiar instruments:

● incident-reporting requirements;

● product and service disclosures;

● professional standards of care;

● procurement conditions;

● audit and record-preservation rules;

● complaint and redress mechanisms;

● penalties for misleading reliability claims;

● restrictions on use outside validated conditions.

This sectoral route also keeps the consequence in view. A hallucinated medical recommendation and a poor marketing caption may arise from related model behaviour. Their reporting thresholds, review requirements and remedies should reflect different stakes.

Cross-sector coordination remains necessary. A common incident vocabulary and reliability-disclosure format would allow patterns to travel across regulators. Central technical expertise can support sector bodies that understand the domain but lack model-evaluation capacity.

New legislation may still be required for coverage, enforcement or jurisdictional gaps. The immediate policy work can start with authorities and safety mechanisms that already exist.

What organizations can do before regulation

Enterprises do not need to wait for a legal requirement to build the core evidence.

They can begin by:

1. creating an internal AI incident and near-miss registry;

2. recording model and configuration details with each consequential output;

3. classifying deployments by task, modality and consequence;

4. testing first-failure points under realistic interaction length;

5. placing context and review limits into workflow design;

6. requiring vendors to disclose tested operating ranges;

7. preserving user reports as reliability evidence;

8. retesting after model, routing, memory or retrieval changes;

9. defining who can stop a deployment when a pattern appears;

10. providing users with a clear route to human review and correction.

This work supports risk management, procurement, insurance and product quality. It also prepares the organization for a regulatory environment built around evidence.

The most mature governance program will know more than how many AI tools it uses. It will know how they fail, how often, under which conditions, who notices and what happens next.

What the blueprint does not promise

An adverse-event registry will not prevent every failure. Reporting systems can suffer from under-reporting, inconsistent evidence and delays. Independent verification can age quickly as models change. A warning can be ignored. User education reaches people unevenly.

These are design constraints, not reasons to leave the evidence uncollected.

The framework also does not claim that one Evans’ Law threshold can govern every task. Reliability testing must remain model-, version-, modality- and use-case-specific. Later work in the PatternPulse program expands the original threshold into a broader reliability surface that includes task rigidity.

The policy principle survives that evolution:

When a system has a measurable operating limit, the limit should be tested, disclosed and connected to responsibility.

The fourth paper’s place in the series

The series begins with a technical observation and moves outward.

Evans’ Law identifies a long-context reliability boundary. AI Conversation Phenomenology examines the user’s experience of crossing it. The When, Where and How of LLM Failures, Measured expands the evidence across models and modalities.

AI’s Accountability Gap turns those findings into public infrastructure.

The next paper tests another escalation: what happens when the same probabilistic system is wrapped in tools, planning loops and claims of persistent agency. Agentic systems raise the consequence of failure because generated language can become an action, and one degraded step can become the premise for many more.

That escalation makes the accountability framework more urgent. A society cannot govern increasingly active AI systems through isolated support tickets and voluntary postmortems.

It needs a record of harm, independent evidence of operating limits, practical warnings and a way to determine who was responsible for the conditions of use.

The accountability gap is therefore an evidence gap first.

Build the evidence, and governance gains something it can act on.

Research and series links

● Original paper: AI’s Accountability Gap: A Policy Blueprint for Policymakers

● Series index: LLM Flaws: The PatternPulse AI Reliability Research Series

● Previous article: When LLMs Fail: The Reliability Boundary Is Measurable — add B2BNN URL after publication

● Earlier article: AI’s Unmeasured Reality: The Safety Evidence Benchmarks Leave Out — add B2BNN URL after publication

● Related PatternPulse policy analysis: Canada’s AI Strategy Needs a Regulatory Engine

● Related framework: NIST Artificial Intelligence Risk Management Framework

Jennifer Evans is the founder of PatternPulse AI and co-founder of Tech Reset Canada. Her research examines LLM reliability, AI Conversation Phenomenology, semantic governance, agentic systems and public AI.

Featured

Agentic Ratio Validation: Harness Quality Sets the Safe Range of Agency

From Agent Harness Engineering, https://openreview.net/forum?id=3hXEPbG0dh, May 2026 In the...

Agentic Update: Containment and Cost, What July’s Agent Releases Have in Common

Agentic development focus has shifted again, from its action...
Jennifer Evans
Jennifer Evanshttps://patternpulse.ai
Principal, patternpulse.ai, and cofounder, Tech Reset Canada. AI policy, research and analysis. Entrepreneur since 2002, marketer since 1998, machine learning since 2009. Based in Toronto and Southeast Asia.