Sunday, July 12, 2026
spot_img

The Operational Layer Goes Dark

Enterprise AI is moving into surfaces no one can inspect, and Canada has no sovereign floor beneath it

By Jen Evans and Laszlo Lakatos-Hayward

Executive Summary: Agentic AI is shifting from answering questions to carrying out work, and the most detailed adoption record published to date shows that change accelerating: more than fivefold growth in six months, with the steepest uptake among people who do not write software. That record measures usage and leaves correctness blank, yet the climb is being read as a verdict on reliability that the evidence does not render. As delegated work moves into surfaces an operator cannot inspect, such as an AI assistant that lives inside a companyโ€™s Slack and follows its threads, the binding form of dependency becomes organizational memory rather than the model itself. For a single firm that is a lock-in problem. For Canada it is a sovereignty problem, because the surface this work runs on is foreign-operated and there is no domestic equivalent built to host or audit it. The corrective is to rent the best model available while owning the context layer outright, holding institutional memory somewhere inspectable, portable, and model-neutral, and to build that floor before the context layer becomes another foreign-owned platform dependency



Two documents surfaced in the past week, and read together they describe a single shift that the AI industry would rather discuss one half at a time. The first is an economics paper from OpenAI showing that agentic AI usage grew more than fivefold in the first half of 2026, with the fastest growth among people who don’t write software. The second is a widely circulated new feature (or warning) that tagging Claude into a company’s Slack hands Anthropic something more than a user might suspect and more valuable than a subscription: what amounts to working memory of how the organization actually operates.

The first document is about adoption. The second is about capture. The thing that connects them is visibility. Work is moving into a layer the operator can no longer see into, at the exact moment that layer is being trusted with larger and more consequential tasks. For an individual firm that is a reliability and lock-in problem. For Canada it is an AI sovereignty problem, because the surface this work runs on is foreign-operated, and there is no domestic equivalent built to host or audit it.

The measured shift from asking to delegating

The OpenAI paper, The Shift to Agentic AI: Evidence from Codex, is the most detailed adoption record published to date, and its headline figures are striking. Weekly active usage of the company’s agentic tool grew more than fivefold between January and June 2026. More than ten percent of users now run three or more agents at the same time in a given week. Just over a quarter of active users invoke reusable skills that encode instructions for complex workflows.

The figure that matters most for what follows is task length. In December 2025, only about two percent of users sent a single request that would have taken an experienced human eight hours or more to complete. By May 2026 that share had risen to roughly twenty-six percent, a more than tenfold increase in half a year. Output rose to match: inside OpenAI, the paper reports that the median worker in a legal role generated thirteen times more monthly output than seven months earlier, and the median researcher more than fifty times as much.

The paper’s own framing is that this is a change in kind, not degree. Conversational tools are used to ask. Agentic tools are used to delegate work that the system then carries out by inspecting files, running commands, and modifying artifacts. The unit of work shifts from the question to the handed-off task. That shift is real, it is fast, and it is happening across seniority levels and well beyond engineering. None of that is in dispute here.

What the adoption record does not measure

Adoption does not equate accuracy, and the paper is careful to measure only the first. Its classifiers count tokens, runtime, concurrency, and task type, and they do so without researchers verifying or reading the underlying messages. There is no estimate of how often the delegated work is right. This is not an oversight to be criticized; it is a scope the authors state plainly. What needs attention is how the result is being read. A fivefold rise in usage and a tenfold rise in task complexity are being treated, in commentary and in procurement conversations alike, as if they were evidence that the tools had become dependable enough to trust with that work. The paper makes no such claim.

It does do something more useful. It names verification as the central open question. The authors write that agentic systems create new opportunities for delegation while introducing new requirements for supervision, verification, and coordination, and that the value of the technology depends on whether organizations can redesign review processes around delegation and verification. In other words, the people closest to the data are telling readers that the hard part is checking the work, and that the checking has not been solved. The denominator in the reliability equation, the rate at which delegated work is actually correct, is the one quantity the record leaves blank.

Why longer tasks widen the margin for error

There is a structural reason to expect the error surface to grow rather than shrink as delegation lengthens, and it follows a pattern I have documented and articulated as Evans’ Law: the longer a model reasons, the greater the likelihood that the response will be incorrect, up to the point where an incorrect answer becomes more likely than a correct one. The law is about predictability. A long chain of reasoning and tool calls has more independent places to diverge than a single exchange, and the paper’s own complexity curve shows those chains lengthening quickly.

Agentic AI is a complicated and proliferating beast, but the one thing that is true across almost all deployment is how little of it actually is AI. It is either highly scaffolded or broken into such tiny, discreet steps with isolated AI that the agency is all but removed from the AI itself. It is simply an ingredient But breaking work into scaffolded steps does not retire the need to verify it. It relocates that need across hops the operator does not watch. When a chatbot returned a wrong answer, a person read the wrong answer and could discard it. When an agent files the wrong record across eight tool calls, the error is wrapped in confident, completed-looking action and buried inside a workflow no one observed. The cost of catching each error rises precisely because the visibility of each step falls. More autonomy means more surface area for divergence, and less occasion to notice it.

The cliff between code and everything else

The paper supplies the sharpest version of the reliability concern through its own logic. Codex began as a software tool, the authors explain, because software is a domain where outputs are useful, economically important, and comparatively easy to verify. They restate the point in the conclusion: software work is digital, produces verifiable artifacts, and breaks into modular subtasks. Code either compiles and passes its tests or it does not. That property is what makes supervised agentic coding workable, and it is genuinely a strength.

Then the same paper shows that the fastest growth is happening outside software, in legal, recruiting, sales, communications, and other knowledge work where none of those verifiability properties hold. A drafted contract clause, a summarized hiring decision, a routed customer promise: these have no compiler and no test suite. The scaffolding that lets an organization trust a coding agent does not exist for the work that is adopting agents most rapidly. The verification floor and the adoption frontier are moving in opposite directions.

The context layer is where sovereignty is decided

This is where the second document comes into play. The argument, made by Princeton’s Arvind Narayanan, is that once an AI assistant lives inside a company’s Slack and follows its threads, the binding form of lock-in stops being the model and becomes the organizational memory: the exception paths, the unwritten owners, the decisions a thread reached, the institutional knowledge that something was tried and failed. Models can be swapped. Agents can be copied. The accumulated understanding of how a specific organization works is far harder to move, and harder still when it sits inside the same vendor that supplies both the intelligence and the surface the work runs on.

For a single firm that is a commercial risk. For a country it is a governance one, and the two failure modes share a root. An organization cannot verify what it cannot see, and it cannot see work that is buried inside another party’s agent layer. The tacit context the agent must reconstruct in order to act, who owns a decision, what a conversation actually settled, whether a given course was already ruled out, is exactly the knowledge it will sometimes reconstruct wrong, with confidence, and then act upon. Dependency and invisibility compound. A surface that is both indispensable and unauditable is a sovereignty exposure regardless of the vendor’s intentions.

Government already runs on this surface

None of this is theoretical for the public sector, because the operational layer in question is already inside government. Slack operates dedicated government tiers authorized to FedRAMP standards and running on United States cloud infrastructure operated by United States personnel, with the Department of Defense, Veterans Affairs, and the General Services Administration among its named users. Cross-organization collaboration on that tier is restricted to other government instances. The agent now moving into messaging surfaces is moving into the same place where allied governments already coordinate their work.

Canada’s position is weaker, and the weakness is instructive. The federal posture has been informal adoption rather than sovereign provision. The Treasury Board pressed years ago for a risk-based approach to unblocking collaboration tools, even as individual departments continued to block them, and a cross-government community spanning more than a dozen federal departments along with provincial and municipal staff grew up on Slack as a workaround rather than as sanctioned, inspectable infrastructure. There is no Canadian equivalent of a sovereign-grade, auditable government instance built for this purpose. The task class the agentic layer is expanding into, drafting and summarizing internal messages, is precisely the surface on which Canadian coordination already informally depends.

A caveat is owed on currency. The clearest public record of federal collaboration-tool policy dates to the Treasury Board’s earlier guidance, and departmental practice varies and is not fully documented in public. The accurate claim is informal adoption plus the absence of a sovereign substitute, not a current sanctioned deployment at scale. That distinction is worth confirming against present departmental policy before it is relied on, and it does not soften the underlying point: where work moves into an unauditable foreign-operated surface, sovereignty is lost regardless of how that adoption arrived.

The strongest case for the other view

The optimistic reading is that this is efficiency. Reusable skills and codified workflows do reduce variance on repeated tasks, and the paper shows organizations adopting them fastest exactly where shared context makes them most valuable. Supervised agentic coding keeps a human in the validation loop by design. Some classes of error are more catchable inside a structured tool call, which leaves a trace, than inside free-flowing text, which does not. Standardization through skills is a plausible path toward making AI use legible inside a firm rather than less so.

The limit of that case is drawn by the same evidence that supports it. Each of these mitigations is strongest in software and in high-context, well-resourced environments with mature review processes. The paper’s own signal is that adoption is growing fastest among non-developers and outside the settings where those review processes exist. In government, the mitigations run into a harder wall still: a verification process cannot give a country sovereignty over a surface it does not host and cannot inspect. The tools that make agentic work legible inside a firm do not make a foreign-operated layer accountable to a domestic public.

What a sovereign posture would require

The corrective is not to refuse the intelligence. It is to refuse the bundle. The sound architecture, for firms and governments alike, is to rent the best model available from whichever vendor leads in a given quarter while owning the context layer outright. Company and government memory should be inspectable, permissioned, portable, and model-neutral, held somewhere other than inside the same vendor that supplies the model and the workflow surface. Separating the memory from the model is what preserves the ability to leave, to compare, and to check.

For Canadian policy specifically, four requirements follow.


Procurement for agentic systems in the public sector should mandate that organizational memory and audit logs remain in Canadian-controlled, model-neutral storage, separable from any single vendor.

  • – Agent actions in messaging and workflow surfaces should be logged in a form a department can inspect and export independently of the platform, so that verification does not depend on the vendor’s goodwill.
  • Reliability for non-software tasks should be measured before deployment, not assumed from adoption figures, with the correctness rate treated as a procurement gate rather than a footnote.
  • the country needs a sovereign-grade collaboration and agent-hosting capability, or a contractually enforceable equivalent, so that the surface government runs on is one it can audit and, if necessary, leave.

Collective bargaining is part of the mechanism, not a side note. Federal IT workers’ representatives are already negotiating clauses governing how AI is introduced into public-sector work, and those clauses are one of the few existing levers that can require inspectability and human review as conditions of deployment rather than aspirations after it.

Adoption vs quality as metric

The operational layer of knowledge work is going dark faster than the means to inspect it, and for Canada faster than any capacity to host or audit it at home. The most authoritative record of that shift measures usage climbing steeply and concedes, in its own words, that whether the work is correct is the question still to be answered. Adoption is being read as a verdict on reliability that no one has actually rendered.

Sovereignty over the context layer is the precondition for ever rendering it. A country that does not control where its institutional memory lives cannot inspect the work done against it, cannot measure how often that work is right, and cannot leave when leaving is warranted. The denominator is missing from the record because the architecture currently in use is built to keep it missing. Building an architecture that exposes it is the work now in front of Canadian firms and Canadian policy, and it is worth beginning before the layer closes over entirely.

Featured

The Third Position: Collective Procurement and the NATO Maven System

Part of the Canadian AI sovereignty series My recent analysis...

Data Adjacency: How Canada Is Now Exposed to AI Systems It Never Procured

Part of the Canadian AI Sovereignty Series The implications of...

Models May Be Attempting Ambiguity Resolution, a Small Subset of Intelligence

On July 6, Anthropic published interpretability research identifying what...

Quick Take: Is a flock of AI unicorns a bubble?

As of June 2026, there are currently 1,778 startup...