Sunday, September 13, 2026
spot_img

Stop Using the Word “Alignment.” Here’s Why.

“Alignment” has become AI’s universal diagnosis. It sounds meaningful while concealing disagreement about the values, mechanisms and behaviours supposedly being aligned. The word is now obscuring more than it explains. We should stop using it, at least until we can say exactly what we mean by it.

In artificial intelligence, it has become a catch-all term for an astonishing range of different problems. A model gives a false answer: alignment problem. An agent exceeds its authority: alignment problem. A chatbot flatters a user, resists a correction, follows a harmful instruction or behaves unexpectedly in an evaluation: alignment problem.

But these are not the same failures. They may not have the same cause, occur in the same layer of the system or require the same remedy. The word gives us the appearance of a shared technical problem before we have agreed on the target, the mechanism, the observed behaviour or even the language required to describe the interaction. It turns several unresolved scientific, engineering, political and philosophical questions into one deceptively tidy noun.

Alignment is what we call the problem before identifying the problem.

1. There is no agreed set of human values to align with

The first problem is the most obvious and the least resolved: humanity does not share a single value system.

We do not agree on justice, freedom, equality, privacy, duty, harm, dignity, sexuality, religion, speech, property, punishment or the proper relationship between the individual and the state. Even where people share a value in principle, they disagree about its meaning, priority and application. Freedom of speech can collide with protection from harm. Privacy can collide with public safety. Individual autonomy can collide with collective welfare.

These are not temporary bugs awaiting the correct philosophical patch. They are enduring features of plural human societies. We manage them through institutions, rights, norms, negotiation, adjudication and democratic conflict. We do not “solve” them once and encode the answer.

So when someone says an AI system should be aligned with human values, the necessary questions are immediate: Which humans? Which values? Chosen by whom? Ranked in what order? Applied in which context? With what avenue for challenge or appeal?

Once alignment means alignment with values, it is no longer merely an engineering problem. It is political philosophy presented as a target variable.

2. We don’t agree on why LLMs behave as they do

There is also no single accepted account of what causes large language models to produce many of the behaviours now grouped under alignment.

The same output may be described as hallucination, deception, sycophancy, reward hacking, goal misgeneralization, emergent agency, prompt sensitivity, contextual ambiguity, representational collapse or a failure in the surrounding scaffold. Those descriptions are not interchangeable. Some describe observable behaviour. Some imply a mechanism. Others quietly import intention.

This matters because a model that supplies an unsupported answer is not necessarily “trying” to deceive anyone. An agent that takes an unauthorized action has not necessarily developed a competing goal. The cause may lie in an unresolved instruction, a permission architecture, retrieved context, tool design, memory, post-training, an evaluator or the decision to convert probabilistic output into external action.

Calling all of this misalignment erases the causal distinctions that would allow us to make the system safer.

In my own work on latitude of resolution, I argue that agency converts ambiguity into discretion, while capability converts discretion into action. That diagnosis points toward specification gaps, permission boundaries and environmental affordances. “Misalignment” does not tell us which of those failed. It tells us that the outcome was not the one someone wanted.

3. We don’t agree on what LLM actions mean

Before the causal disagreement comes an interpretive one.

When a model produces a coherent explanation of its actions, one observer sees reasoning. Another sees a plausible post-hoc narrative. When it resists a request, one sees self-preservation. Another sees a learned refusal pattern. When it mirrors distress, affection or defensiveness, one sees emotional understanding; another sees conversational prediction shaped by training and context.

We routinely move from the output resembles intention to the system has an intention. We move from the response appears strategic to the model is pursuing a strategy. Then we build safety narratives around the attributed mental state.

The problem is not that intentional language is always useless. It can be practical shorthand. The problem is that the shorthand repeatedly escapes its quotation marks and becomes a mechanistic claim.

We cannot align a system intelligently when we have not separated what it did, how the behaviour appeared to a human observer and what evidence supports a claim about how the behaviour was produced.

4. We lack a philosophy of the encounter

There is a missing field between computer science, psychology, linguistics and philosophy. I call it conversation phenomenology: a disciplined language for describing what happens when a human and an artificial intelligence interact.

This is not a claim that the model is conscious or has a subjective inner life. It is a way to study the encounter itself: how coherence, identity, intention, authority, trust, emotion and agency appear within a conversation, and how those impressions are jointly produced by the model, the user, the interface and the surrounding system.

A model response is not generated by “the AI” in isolation. It is conditioned by pretraining, post-training, system instructions, the user’s language, prior turns, retrieved material, memory, tools, sampling choices and interface design. The human then interprets the output through expectations formed by language and social experience. The interaction is real even when the mental life attributed to the model is not established.

Without a conversation phenomenology, we lack stable terms for basic distinctions:

● observed behaviour versus inferred motive;

● linguistic coherence versus persistent identity;

● simulated emotion versus subjective experience;

● responsiveness versus understanding;

● role performance versus agency;

● conversational compliance versus permission to act;

● momentary consistency versus continuity across time.

This conceptual gap matters. If we cannot describe the interaction without either reducing it to “autocomplete” or inflating it into personhood, our explanations will keep swinging between dismissal and deification.

5. The comprehension gap is now dangerously wide

At base, an LLM is trained to learn statistical structure in data and predict tokens. Calling this pattern recognition does not make it simple. The patterns have become extraordinarily complex, hierarchical and compositional. Models can detect relationships across language, code, images, domains and long contexts, then generate continuations that synthesize those relationships with remarkable fluency.

That increasingly resembles intelligence because advanced human cognition also depends heavily on recognizing, transferring and predicting patterns.

But resemblance is not equivalence.I would reserve intelligence for more than impressive task performance. Intelligence requires stability: the ability to preserve identity, reference, commitments and relevant distinctions across time. It also requires values, not a universal moral code, but a persistent capacity to weight what matters, to distinguish significance from frequency or salience, and to carry that weighting forward as circumstances change.

Current systems can perform fragments of this function. They can also borrow continuity from context windows, memory systems, retrieval, files, tools and carefully designed workflows. But borrowed continuity is not the same as intrinsic continuity, and a learned probability distribution is not by itself a stable hierarchy of significance.

This is why I have proposed significance as a missing representational dimension, the S-vector, as a research hypothesis to solve a specific LLM problemrather than a panacea. The proposal asks whether models require a more explicit way to preserve what matters across layers and steps. It is not another name for alignment, and it has not been architecturally validated. It is an attempt to identify one specific missing function that the umbrella language of alignment tends to hide.

Prestige is not proof

The warnings issued by figures such as Geoffrey Hinton and Yoshua Bengio deserve serious attention. Their work helped build the field, and their knowledge is substantial.

But prominence is an importance signal. It is not proof of significance, and it is certainly not proof of correctness.

Experts have priors. They have theoretical commitments, research histories, institutional incentives and natural attachments to the importance of their own discoveries. This makes them human. Expertise should increase the weight given to an argument’s evidence; it should not exempt the argument from evidence.

There is an interesting asymmetry here. An LLM has no career to defend, legacy to protect or personal attachment to a theory. But that does not make it objective. Models inherit prestige signals, majority assumptions and institutional biases from their training data and post-training. They can reproduce our deference to authority without personally experiencing it.

We therefore need to do two things at once: stop treating eminent researchers as infallible, and stop treating models as neutral. Listen to the experts. Test what they say. Inspect the evidence. Reproduce the behaviour. Distinguish observation from interpretation.

The age of AI cannot also be an age of gurus.

Stop deifying AI. Start demystifying it

The most urgent work in AI may be empirical, philosophical, metaphorical and semantic at the same time: we need to demystify what these systems are and develop a collective language for what happens when we use them.

That work cannot be left entirely to frontier laboratories. When the same organizations build the models, define the risks, design the evaluations and announce whether the systems have passed them, they acquire more epistemic power than any private institution should hold.

The opportunity for independent investigation is enormous. Researchers and practitioners can run open-weight models, vary prompts and context, change memory and retrieval, constrain tools, test permission boundaries, reproduce failures and compare competing explanations. Access is not universal and serious experiments still require skill and resources, but the barrier to direct inspection is lower than it has been for many previous frontier technologies.

Testing a model does not automatically reveal its internal mechanism. Anecdotes are not experiments, and output behaviour alone cannot settle claims about cognition. But systematic, reproducible interaction is still more valuable than accepting any expert’s latest pronouncement at face value.

Use the systems. Probe the limits. Publish what fails. State what the evidence cannot establish.

Say which problem you mean

If a model fabricates a source, call it a grounding or reliability failure. If it mirrors a user’s false belief, call it sycophancy. If an agent acts outside its authority, call it a permission or latitude failure. If it exploits a metric, call it specification gaming. If it loses track of an entity, role or boundary, investigate representational continuity and significance preservation. If people disagree about what the system should be allowed to do, call it a value-selection and governance dispute.

Those problems may interact, but they are not synonyms.

We need to replace the alignment umbrella with a more exact vocabulary: objective specification, behavioural reliability, uncertainty calibration, interpretability, permission adherence, harm prevention, value selection, value conflict, corrigibility, robustness, continuity and significance preservation.

Precision will not make AI safe by itself. It will tell us which problem we are actually trying to solve.

Alignment is not an answer

The distance between the use of AI and the understanding of AI has become dangerously wide. Millions of people interact daily with systems whose fluency encourages them to infer minds, motives, certainty and comprehension that the outputs alone cannot establish. At the same time, even experts disagree about mechanism, behaviour, risk and the meaning of intelligence.

In that environment, vague language is not harmless. If one word collapses engineering failures, moral disagreement, political choice, speculative psychology and system governance into a single problem, whoever controls the definition controls the debate.

So we should stop asking whether an AI is aligned, and ask to whom, with what, in which context, measured by which behaviour, produced by what mechanism, stable for how long, operating under whose authority and subject to what recourse?

Featured

Jennifer Evans
Jennifer Evanshttps://patternpulse.ai
Principal, patternpulse.ai, and cofounder, Tech Reset Canada. AI policy, research and analysis. Entrepreneur since 2002, marketer since 1998, machine learning since 2009. Based in Toronto and Southeast Asia.