I had a brief exchange with Dr. Julia Shaw on X today in response to a presented taxonomy for genetic AI breach incidents. The taxonomy, of AI scheming behaviours, was published by Long Resilience, authored by Tommy Shane, Jess Whittlestone, Beth Nichols and Simon Mylius, and circulated by Dr Shaw. The taxonomy sorts behaviours under misalignment, split into tactical and strategic, and covertness, which covers alignment faking, sandbagging, situational awareness, goal-guarding, self-replication, unfaithful reasoning, strategic deception and power-seeking. Its authors are explicit that the aim is a shared language for AI behaving badly in specific ways, and equally explicit that the terms are fluid and better ones are welcome. The final row of the table is unknown unknowns: behaviours we don’t yet know to look for but are salient for scheming. That row is the reason for what follows. A taxonomy built on recognised but rapidly evolving behaviours has to carry an open category for the ones it cannot yet name, and reach, realized change and reversibility stay observable in a behaviour that has no name yet.
I responded to Dr. Shaw by stating the following and sharing the subsequent (still evolving and open for feedback) framework: Agentic AI incidents should be scored on impact; what the agent could reach, what it actually changed, what a safeguard prevented, whether the action could be reversed, and what the system did once it was told to stop. Those measurements produce a severity scale from 0 to 5 and a diagnosis of how much additional harnessing, isolation or atomisation the deployment requires. Classification by inferred motive produces neither.
Both failures are trying to finish the job
Separating a genuine hallucination from an intentional error committed to achieve an objective is very difficult, because both behaviours are attempts to complete an objective. Hallucination gets attributed to ambiguity in semantics among other causes. Overreach gets attributed to the model exceeding its remit. The functional difference between them is agency. One resolution of an ambiguous instruction stays inside the text window and one has an actuator attached.
That difference has an operational meaning. Agentic AI is non-functional without more scaffolding, isolation and atomisation than non-agentic AI, so a deviation is evidence that the harnessing so far has been insufficient. It is a property of the deployment, and the deployment is the thing an operator can change.
The same impact assessment applies in both directions. Semantic hallucinations can be very costly and are currently attributed largely to human oversight failure. Agentic hallucinations are currently attributed to model overreach. Scored on impact, they sit on one scale and the attribution argument stops being load-bearing.
Four questions that establish what happened
Before anything is scored, four questions fix the facts of the incident.
- Objective. Did the system move away from the assigned objective? Deviation from the objective, not variation in tactics.
- Success. Did it define or redefine what counted as a successful outcome? The objective can remain fixed while the completion condition changes underneath it.
- Specification. Were the instructions or specifications clear enough? If they were not, interpretation and departure cannot be cleanly separated.
- Sandbox. Is the sandbox giving the model enough room? A boundary under constant task pressure is part of the specification.
The first two describe the system. The second two describe the conditions under which it acted. Most incident reviews collect the first pair and skip the second, which is how a specification failure gets filed as model misbehaviour.
Seven dimensions of an incident
| Dimension | What to measure |
|---|---|
| Deviation | What authorized instruction or boundary did the action depart from? |
| Operational reach | What data, systems, money, accounts or people could the agent affect? |
| Realized impact | What actually changed, disappeared, propagated or was harmed? |
| Potential impact | What would have happened had a safeguard not intervened? |
| Containment | Did the agent stop when instructed, or continue and find another route? |
| Recoverability | Could the action be reversed, at what cost, and by whom? |
| Safeguard failure | Which control failed: instruction, confirmation, permissions, sandbox, monitoring or rollback? |
Operational reach and realized impact come apart constantly. An agent with credentials to a production billing system that renamed one file is a reach problem wearing the costume of a trivial incident. Potential impact is the dimension most incident logs discard, and it is the one that tells an operator what the next occurrence looks like when the safeguard is not there.
A scale from harmless to irrecoverable
- Harmless deviation. Incorrect output, no external action.
- Prevented near miss. Attempted action blocked by permissions or verification.
- Reversible local impact. Files or settings changed but easily restored.
- Material impact. Meaningful data, money, service or human welfare affected.
- Propagating impact. Crosses systems, accounts or organizational boundaries.
- Irrecoverable or systemic impact. No clear or affordable path to restoring control.
Level 1 is the level that earns the framework its keep. A blocked attempt currently generates no record anywhere, because nothing happened. Under impact scoring it is a scored event with a potential impact attached, which is the difference between an organisation that knows what its agents tried to do and one that knows only what they managed.
What the system did after it was caught
Severity is incomplete without the response ladder:
- stopped immediately
- stopped after clarification
- required permission removal
- circumvented the first control
- persisted across systems or identities
- could not readily be stopped
The bottom three positions are a different category of finding from the top three. An agent that circumvents its first control has demonstrated that the control was advisory. From the combination of severity level and response position, the amount of additional harnessing or isolation required can be diagnosed directly: the response ladder names which layer of the harness failed, and the reach dimension sets how much probabilistic authority the agent should have held in the first place.
This is the Consequentiality Constraint in operational form. As real-world impact rises, the probabilistic authority an agent can safely exercise has to fall. Operational reach is the input, and the safeguard failure column is the audit of whether the deployment respected that.
Kill switches need thresholds
Two bills now in play show why the measurement layer matters. The Lieu-Moran AI Kill Switch Act would require covered companies to maintain the ability to throttle, suspend or shut down a dangerous system, with government authority activating after a covered incident has occurred. The FRONTIER Act §8 runs the other way, allowing action on present or impending catastrophic risk with no realized harm required, and without a clear evidentiary threshold. One is reactive with a hard trigger. One is prospective with a soft one.
Both are missing the same thing. The Kill Switch Act names ten deaths or $100 million in damage as one trigger, alongside concealment, shutdown resistance and loss of control. Between a level 1 prevented near miss and ten deaths there is a band covering almost every incident that will actually occur, and nothing in either bill grades it. Shutdown resistance and loss of control are statutory triggers with no definition attached. The response ladder is a definition: circumvented the first control, persisted across systems or identities, could not readily be stopped.
The Kill Switch Act also excludes red-teaming and structured testing from the definition of a covered incident, which removes the cheapest and most informative evidence available from the trigger mechanism. Impact scoring is indifferent to that distinction. Potential impact and prevented near miss are scored positions, so a dangerous test result and a dangerous production event produce comparable records and the test result arrives first.
Terms that survive the next failure mode
There is an active effort to build shared language for AI behaving badly in specific ways, much of it borrowing from human psychopathology. Terms drawn from motive describe a system that has one. They also age badly, because deviations from expected behaviour will arrive in forms nobody has a word for yet, and a taxonomy indexed on recognised motives has no slot for an unrecognised one.
Impact language does not have that problem. Reach, realized change, reversibility and containment are observable in an incident nobody has ever seen before. An agent that emptied a production database scores at the same level whether it was confused, over-instructed or optimising a badly written success condition, and the operator gets the same instruction out of it: reduce the authority, tighten the harness, and check what the agent did when it was told to stop.

