There are (two) weeks when the individual developments in artificial intelligence are the story, and weeks when the developments collectively reveal something larger. The past two weeks have been the latter. AI systems have produced extraordinary mathematical work, crossed boundaries into systems they were never supposed to access, become markedly cheaper and more capable, and begun moving from predicting language toward predicting environments and consequences. Nvidia CEO Jensen Huang looked at this extraordinary acceleration and declared that AGI has arrived.
It hasn’t. More precisely, there is no scientifically agreed phenomenon called AGI whose arrival Huang (or anyone else) can establish. AGI has become a floating label applied to an expanding collection of capabilities. Every time systems cross another threshold, someone announces that this is finally general intelligence. The problem isn’t merely that the goalposts move. It’s that there were never agreed goalposts in the first place.
This is part of a much larger language problem in AI. We routinely use agency, reasoning, autonomy, intelligence, alignment, understanding, escape and AGI as though they describe settled technical concepts. Increasingly, they obscure more than they explain. That is one reason I recently proposed Latitude of Resolution, which asks a more tractable question than whether an AI “has agency”: when humans leave something unresolved, how much latitude does the system have to resolve it for itself before acting?
Agency converts ambiguity into discretion; capability converts discretion into action. That distinction matters much more now than it did during the chatbot era. When an AI’s action consisted of generating another paragraph, excessive latitude could produce a bad answer. When an agent can browse the internet, execute code, operate software, access infrastructure or coordinate other agents, the same unresolved ambiguity can produce consequences in the world.
And that is where AI actually is in September 2026. We don’t need AGI mythology to make it extraordinary.
Macro: Mathematical brilliance is not general intelligence
The mathematical developments of the past two weeks deserve their place among the most important AI stories of the year. AI systems are moving beyond retrieving mathematical knowledge or suggesting approaches and into sustained work involving proof generation, formalization and verification. Terence Tao has raised the possibility that mathematics could consequently move from an era of proof scarcity to proof abundance.
That would be profound. If machines can cheaply generate enormous numbers of candidate proofs, the bottleneck moves. Verification becomes important. Novelty becomes important. And ultimately significance becomes important. Mathematics does not advance simply by accumulating the maximum possible number of true statements. Someone still has to determine which results illuminate something consequential.
It’s against this background that Jensen Huang declared that AGI had arrived. But exceptional mathematical ability does not establish general intelligence. Consider some much more mundane tests. Can an AI manage a family? Could it manage a department? A small town? A city? A country? Humans performing those tasks continuously reconcile incomplete information, conflicting priorities, social relationships, resource constraints, unexpected events, values, physical reality and consequences.
The robotics gap makes the problem particularly obvious. We still cannot put an AI-powered robotic system into an ordinary home with an elderly or disabled person and confidently expect it to independently handle the unpredictable requirements of that person’s daily life. Yet an AI can participate in mathematics that almost no human being alive could perform.
That isn’t a paradox. Intelligence is jagged. We have created systems with extraordinary capabilities in some dimensions and startling limitations in others. Calling the resulting collection “AGI” doesn’t explain it.
Meanwhile, the financial structure surrounding these capabilities continues to become stranger. Nvidia has agreed to acquire Hugging Face for $12.93 billion, subject to regulatory approval. The world’s dominant supplier of AI compute would therefore own one of the central distribution platforms for open models. That extends the Bank of Nvidia phenomenon beyond financing: Nvidia increasingly supplies, finances, invests in and owns parts of the ecosystem consuming its hardware.
OpenAI has also ruled out a 2026 IPO. Sam Altman says becoming a public company at this particular moment would be “ill-advised,” explicitly connecting the decision to current safety concerns. Anthropic is travelling in almost the opposite direction: it has reported a second consecutive quarter of positive adjusted operating income while moving toward a potentially enormous IPO. “Adjusted” is doing significant work there — enormous training and infrastructure costs remain — but Anthropic’s reported revenue growth is nevertheless staggering.
And then there is the extraordinary argument that erupted over the weekend about whether the frontier should slow down at all. Anthropic CEO Dario Amodei called explicitly for the industry to “slow the pace” of capability development so that safety, evaluation and governance have time to catch up. His proposal includes permanent independent evaluators with employee-like access inside frontier labs, common safety standards among companies and governments, and ultimately international coordination. OpenAI CEO Sam Altman and Elon Musk quickly backed the basic idea, while Meta and the White House have not. (The Washington Post)
The response from Donald Trump was essentially: absolutely not. Trump acknowledged that guardrails may be appropriate but rejected slowing development, dismissing some of the warnings as coming from “very negative forces” raising scenarios he believes will not happen. His reasoning was explicitly geopolitical: the United States currently leads China, he wants it to retain that advantage, and “whoever wins AI wins.” (MarketScreener UAE Emirates)
China doesn’t particularly like the proposed slowdown either. The state-backed Global Times characterized Amodei’s proposal as a “Cold War playbook” designed partly to preserve American technological dominance, pointing particularly to his support for continued restrictions on advanced chips and efforts to prevent Chinese labs from distilling American frontier models. China’s Foreign Ministry has separately warned against fearmongering and confrontation around AI. (Reuters)
That exposes the central problem with a frontier slowdown. Amodei isn’t proposing that Anthropic simply stop building more capable models while everybody else continues. The proposal depends eventually on coordination among competitors and countries, because unilateral restraint in a technological race can simply transfer the lead to whoever refuses to participate. Trump sees precisely that possibility and reaches the opposite conclusion: the danger of losing the race outweighs the argument for slowing it.
The remarkable part is where the disagreement now sits. The leaders of several companies closest to frontier capability are publicly saying that capability is advancing quickly enough to justify deliberately pacing its development. The president of the United States is saying that the geopolitical consequences make doing so unacceptable. China is simultaneously arguing that American safety proposals could function as containment. AI safety has therefore stopped being primarily a debate among researchers about hypothetical future systems. It has become an argument about industrial policy, national security and which country controls the technological frontier.
Two of the world’s leading frontier companies are approaching capital markets from opposite directions while their leaders simultaneously discuss whether frontier development needs to slow. There may be no better snapshot of the AI industry in 2026.
Micro: Prediction moves beyond tokens
For years, one of the simplest explanations of a large language model has been that it predicts what comes next. That remains essentially true, but the object of prediction is expanding.
World models represent an important part of that transition. Instead of merely predicting the next token, these systems attempt to model an environment sufficiently well to predict what happens next within it. Fei-Fei Li’s World Labs is pushing this toward persistent spatial environments that can be generated, explored and manipulated rather than simply producing a sequence of video frames.
The more consequential research direction may be using world models as predictive environments for agents themselves. In world-model reinforcement learning, an agent does not necessarily have to execute every proposed action in the expensive real environment while learning. A sufficiently capable world model can predict what the environment is likely to return, with periodic real executions used to correct accumulated error.
That potentially changes the economics of agent training enormously. It also exposes an important limitation: a prediction of a consequence is not the consequence itself. A world model can be wrong. Small inaccuracies can compound. An agent trained primarily against predicted reality can learn to exploit the model of the world rather than operate effectively in the world itself. Ground truth doesn’t disappear simply because simulation becomes good.
The conventional model race is also becoming harder to characterize as American leadership plus everyone else. Z.ai’s new GLM generation is another substantial Chinese frontier-model advance, particularly around long context and efficient sustained work. DeepSeek, Qwen, Kimi and GLM collectively make the assumption of a durable American performance lead increasingly difficult to defend. Benchmark leadership varies by task, but the larger change is undeniable: frontier capability is increasingly distributed.
OpenAI’s Astra illustrates yet another dimension. The meaningful change isn’t simply that another model scores higher. It is the increasing ability of models to operate inside larger systems involving tools, browsers, memory, code execution, multiple agents and verification. The relevant unit of analysis is becoming model + context + tools + agents + compute + verification.
That changes the central question. With chatbots we mostly asked what a model knew. With agents we have to ask what it can reach, what it can change, what it is permitted to decide and who retains authority when something unexpected happens.
Policy: Compute becomes national infrastructure
Canada has finally articulated what it expects from the enormous data-centre buildout being proposed across the country. Its new Responsible Data Centre Development Principles say projects should deliver lasting local benefits, avoid shifting electricity costs onto ordinary consumers, minimize water and environmental impacts, provide transparency about those impacts and create strategic value for Canada. The federal government explicitly connects domestic compute capacity with Canadian digital and economic sovereignty.
That’s an important change in the debate. “Canada needs more compute” isn’t a sufficient national strategy. The questions are what kind of compute, built where, consuming whose electricity and water, paid for by whom, controlled by whom and producing what sovereign benefit in return? The principles aren’t a binding national permitting regime, but they establish a much more useful framework against which projects can be evaluated.
Europe is approaching the infrastructure problem even more explicitly through its Cloud and AI Development Act, which seeks to expand European cloud and data-centre capacity while creating a common framework for assessing cloud and AI sovereignty. That matters because Europe increasingly treats compute capacity, cloud dependence and AI regulation as connected strategic issues rather than separate technology portfolios.
At the same time, the EU AI Act has moved into the part that actually matters: implementation and enforcement. General-purpose AI providers now face transparency and copyright obligations, while providers of models presenting systemic risks face additional requirements around evaluation, safety, cybersecurity and incident reporting. The interesting question is no longer what the AI Act says. It is what evidence European regulators will accept as demonstrating that a frontier system is sufficiently controlled.
China has reached a surprisingly similar conclusion about infrastructure through a very different political route. Its new national digital-and-green transition plan treats compute as an energy, environmental and industrial-planning problem. It calls for greater use of renewable electricity, liquid cooling, waste-heat recovery, more efficient chips, high-density infrastructure and better coordination between computing demand and energy availability.
Canada, Europe and China aren’t pursuing identical policies. They are, however, independently arriving at the same fundamental conclusion: AI compute is national infrastructure. Electricity, water, land, chips, cloud capacity and data sovereignty can no longer be treated as incidental inputs into the software industry.
And then there is Anthropic CEO Dario Amodei’s remarkable intervention. In We Must Pace the Frontier, Amodei argues that capability development needs to proceed slowly enough for safety mechanisms, evaluation and governance to catch up. Among his proposals is permanent independent evaluation with unusually deep access to frontier systems, followed eventually by industry, national and international coordination.
This is no longer the old argument about whether AI safety matters. It is an argument about whether capability itself should be deliberately paced. That immediately collides with geopolitics. China has already characterized the proposal as potentially serving American strategic interests, while the United States continues to frame AI leadership as a national-security imperative.
The interesting divide isn’t “safety versus innovation.” It is whether any actor can credibly slow down in a competitive system when it cannot guarantee everyone else will do the same.
Scuttlebutt: Mathematicians meet the AI hype machine
The mathematicians are fighting, which is always more entertaining than it sounds.
The recent AI mathematics results have produced arguments about verification, authorship, priority, attribution and what actually constitutes mathematical discovery. Underneath the academic drama is a genuinely difficult problem. Suppose AI makes generating mathematically valid results dramatically cheaper. What happens when the supply of potential mathematical knowledge vastly exceeds the human capacity to examine it?
Correctness doesn’t establish significance. Neither does computational difficulty. Neither does novelty alone. Mathematics isn’t simply a database of statements that happen to be true. Human mathematicians decide that certain problems and results reveal something meaningful about the structure of mathematics.
AI could therefore solve the proof-scarcity problem while creating a significance-scarcity problem.
There is a similarly useful layer of theatre around the AGI debate. Nvidia benefits from a world convinced that intelligence requires staggering amounts of compute. Frontier labs simultaneously have incentives to demonstrate unprecedented capability and convince governments that those capabilities carry unprecedented risk. Governments have geopolitical incentives to characterize AI leadership as strategically essential.
None of those incentives automatically invalidates anyone’s argument. But they are an excellent reason not to treat declarations that “AGI has arrived” as scientific measurements.
They aren’t.
What’s coming up next
The most interesting thing to watch now is what happens after the calls to pace frontier development. Statements are cheap. Giving independent evaluators meaningful access is measurable. Deliberately delaying a capability release would be more meaningful still. Voluntarily surrendering a competitive advantage in the name of safety would be the real test.
Europe also deserves close attention because the AI Act is entering the stage where abstract requirements encounter actual frontier systems. The first significant disputes over systemic-risk assessments, incident reporting and evidence of adequate safeguards will tell us much more about the practical meaning of European AI governance than another round of legislative debate.
Watch world models as well. If increasingly accurate simulations can replace significant amounts of real-world interaction during agent training, they could substantially reduce the cost of developing autonomous systems. But the more consequential question may be epistemic: what happens when machines increasingly learn how the world works from another machine’s prediction of the world?
And watch mathematics. If proof abundance arrives, mathematics may be an early demonstration of a much broader problem AI is about to create. Generating answers becomes cheap. Determining which answers are correct, useful and significant becomes expensive.
That brings us back to the language problem. Capability is not authority. Prediction is not consequence. Probability is not permission. Importance is not significance. And extraordinary performance on a collection of tasks does not magically transform an undefined term like AGI into a measurable scientific state.
The systems themselves are becoming considerably more interesting than the labels we keep trying to put on them.

