Saturday, August 1, 2026
spot_img

The Maturing AI Market: Most Tasks Don’t Need Frontier Reasoning

The implication for home use, private and public sectors is greater affordability

The AI market is maturing, and increasingly people use the cheapest model that clears the bar, because most of what they ask for never required frontier reasoning in the first place. Four independent data sources show the same distribution. OpenAI’s message logs, Anthropic’s Economic Index, QuestMobile’s Chinese app rankings, and OpenRouter’s token volumes demonstrate demand concentrating on explanation, editing, summarizing, planning and recommendations. Frontier-grade reasoning serves a minority of volume and a majority of dollars, and the gap between those two facts is the growing shape of model consumption.

Screenshot

What the message logs contain

It’s some fascinating data, and not necessarily evident if you’re just reading industry chatter. OpenAI’s study with the National Bureau of Economic Research classified more than a million and a half consumer messages and found roughly 80 percent falling into three buckets: practical guidance, seeking information, and writing help. Sorted by what the user wants back:

What people ask for, by subject (OpenAI/NBER, 1.5M+ consumer messages)

  • Practical guidance, seeking information, and writing help together cover nearly 80 percent of conversations
  • Writing has fallen from 36 percent to 24; seeking information has risen from 14 to 24
  • Programming: 4.2 percent
  • Relationships and personal reflection: 1.9 percent. Games and role play: 0.4 percent

What people want back

  • Asking (information or advice): 49 percent. Doing (produce an output): 40 percent. Expressing: 11 percent
  • Among work messages, Doing rises to 56 percent, and nearly three-quarters of those are writing tasks
  • Two-thirds of writing messages ask the model to modify text the user supplied rather than generate new text

What comes out the other end (Anthropic Economic Index, June 2026)

  • 93 percent of conversations produce an identifiable output
  • Explanations 17 percent, documents and reports 15 percent, guidance 11 percent
  • Conversational outputs and written deliverables take about a third each; code and technical work about a sixth
  • Personal use runs about 35 percent on weekdays and just under 50 percent on weekends

That is the bulk of the demand curve: ordinary knowledge work and ordinary domestic life.

The reasoning those tasks actually consume

The research on where extended reasoning pays is unusually clear, and it has been clear since 2024. Sprague and colleagues at UT Austin ran a meta-analysis across more than 100 papers plus their own evaluations on 20 datasets and 14 models. Chain-of-thought delivered average gains of 14.2 percent on symbolic reasoning, 12.3 percent on math, and 6.9 percent on logic. Across every other category the average gain was 0.7 percent. On MMLU, answering directly matched answering with reasoning unless the question or the answer contained an equals sign.

Longer thinking also carries a downside that grows with length. The overthinking literature documents an inverted-U curve where accuracy rises with reasoning tokens, peaks, and then falls as models abandon correct answers they had already reached. Easy problems cross that threshold earliest, around 2,000 tokens against roughly 8,000 for hard ones. Work on inverse scaling in test-time compute found models becoming more distracted by irrelevant detail, drifting toward spurious correlations, and overfitting to problem framings as reasoning length grew.

Users have converged on this empirically without reading any of it. In Anthropic’s sample, extended thinking is switched on in 31 to 34 percent of work conversations. Opus, the most capable tier, serves 10 percent of chat and Cowork conversations. Nine out of ten conversations run on something smaller, and satisfaction does not collapse.

Doubao leads on distribution

China’s AI-native apps reached 499 million monthly active users in May 2026, up 85.4 percent year over year, with average monthly usage of 92.7 sessions and 183 minutes per person. ByteDance’s Doubao holds 382 million of those users. Alibaba’s Qwen holds 167 million. DeepSeek holds 130 million. In the first quarter, Doubao users averaged 54.8 sessions a month against DeepSeek’s 41.7 and Qwen’s 19.8.

Doubao is not the strongest model in that list, and it is not close. It is the one wired into ByteDance’s distribution, priced to disappear, and used most often per user. Moonshot’s Kimi ranked third in Chinese monthly actives before DeepSeek’s R1 landed and fell to seventh afterward, with weekly actives reported around 4.5 million by late 2025. Kimi K3, released July 16, 2026, is a 2.8-trillion-parameter open-weight model that benchmarks alongside the strongest proprietary systems, and Moonshot’s own strategy treats the consumer app as secondary to developer adoption and weight downloads. Its annualized recurring revenue reportedly doubled from about $100 million in March 2026 to more than $200 million by the end of April, which is a developer number rather than a consumer one.

A market where the “dumbest” assistant has the most users and the smartest open model has almost none is one that has found the reasoning level tasks require.

The frontier earns its price on duration, not on answers

Compute tracks value, and Anthropic’s data shows how. Conversations mapping to higher-wage occupations consume more tokens: marketing managers earn roughly twice what editors do and their conversations use about 2.5 times the tokens. Building an app consumes more than three times the median conversation. A typical explanation consumes about a fifth. Roughly 44 percent of that wage gradient is explained by which artifacts higher-paid people ask for.

Autonomy follows the same line. Rated on a five-point scale, the lowest-autonomy outputs are math, translation and question-answering, where the answer is largely determined by the input. The highest are apps, websites, games and presentations, where the model selects among many possible choices. Claude Code sessions run on Opus 54 percent of the time against 10 percent in chat, and the median chat conversation producing a blog post involves 13 rounds of back and forth while the median Claude Code session producing one contains a single human prompt.

The premium attaches to delegation and duration. It is paid for sustained judgment across hours of work, not for a better answer to a single question, and the market increasingly will pay it only where the work is extended.

The frontier lost the modality where quality is the product

Video and audio generation carry no reasoning dial. There is no thinking-effort setting, no chain of thought to lengthen, and the cost scales with seconds of rendered output rather than tokens of deliberation. If frontier capability were going to win a category outright, this is the one where it should have happened, since the buyer is paying for the artifact itself and quality is visible in the first two seconds.

It happened the other way. OpenAI shut Sora down in spring 2026 while burning a reported $15 million a day in compute against $2.1 million in total lifetime revenue.  The consumer web and app experiences ended on April 26, and the API follows on September 24.  Downloads had already fallen 66 percent from their November 2025 peak.  The model that led on physics and photorealism was retired by its own cost per second.

What survived competes on other axes. Kuaishou’s Kling passed 60 million registered creators and 600 million generated videos by December 2025 and reached roughly $500 million annualized revenue by May 2026. Runway raised at a $5.3 billion valuation on about $300 million annualized, and Google’s Veo reaches YouTube’s two billion users through Shorts.  Price, professional control, and distribution each carved out a defensible position. Absolute output quality carved out none.

The demand side explains why. Monthly actives across AI video platforms passed 124 million in January 2026, while the pure video-generation market was worth under a billion dollars that year.  Supply-side generation volume is growing an order of magnitude faster than the demand-side survey numbers, which places most of that volume in consumer experimentation rather than budgeted production.  A hundred million people making things for fun will not fund frontier inference, and the ones making things for money are buying at Kling’s price.

The text data shows the same split from the other direction. In Anthropic’s classification, more than 80 percent of conversations producing creative writing are personal, dominated by fanfiction, worldbuilding and poetry, while the work-related remainder runs to short-form video scripts, screenwriting and speeches. Creative work at volume is largely unpaid, and the paid slice is short.

Three modalities, three cost structures, one result. Thinking tokens in text, seconds of render in video, and distribution economics underneath both. In each case the model that wins the volume is the one priced for the work people are actually doing, and the model that wins the benchmark collects a premium only where the work runs long enough to justify it. Reasoning level is one instance of a more general rule, which is that fit to the task sets the price, and capability sets only the ceiling.

What the routing market already priced in

OpenRouter is the best read on production deployment, since developers switch models by changing one parameter. In June 2025, US models from OpenAI, Anthropic and Google held around 70 percent of token share. By June 2026 that figure had dropped to roughly 30 percent. Chinese-origin models from DeepSeek, Alibaba, MiniMax, Tencent, Xiaomi and Moonshot passed 45 percent of tokens, up from under 2 percent a year earlier. DeepSeek alone commands 16.3 percent of platform volume, more than any other single provider. Anthropic holds around 12.3 percent of tokens while retaining a far larger share of dollars.

Two lanes, one marketplace. Commodity inference flows to whatever is cheapest and adequate. Premium inference concentrates in agentic and long-context work. Inference volume across the tracked market grew roughly eleven-fold year over year, which means both lanes are growing and only one of them is growing in revenue terms for US labs. These are single-platform figures with shifting denominators, and the direction is more reliable than the decimals.

Usage lags capability, and that cuts both ways

The strongest alternative to reading maturity into this data is that usage describes what people currently trust models to do, which is a lagging indicator of what models *can* do. Anthropic’s April 2026 survey supports this. Close to six in ten respondents expect AI to handle a higher share of their work tasks next year than this year, and over a third expect it to handle most or nearly all of them. The people who delegate the most are the most optimistic, and the agentic surfaces where autonomy and compute concentrate are the fastest-growing part of the market.

Message counts also flatten value. A recipe request and a migration plan each count once. The 4.2 percent of messages that are code may carry more economic weight than the 42 percent that are writing. Anyone reading these numbers as a ceiling on what AI will be asked to do is reading them wrong.

What the numbers do establish is the shape of demand today, and demand today is the thing being priced today.

What this means for buying and building


Benchmark leadership and market position have come apart. A lab can hold the top of the leaderboard and lose the volume, and the Chinese market demonstrates this domestically while OpenRouter demonstrates it globally.

Routing becomes the product. When 90 percent of conversations run acceptably on a mid-tier model and 10 percent need the frontier, the valuable engineering is the classifier that decides which is which. Every serious deployment ends up building or buying one.

Price compression at the frontier follows from the same arithmetic, and the flash, mini and fast tiers appearing across every lab’s lineup are the visible form of it.

For anyone buying, the useful audit is a task inventory rather than a model comparison. Sort your workload by whether the output is determined by the input, which is where cheap models are indistinguishable, or selected from many possible choices, which is where the premium buys something. Default to the cheap tier and escalate on evidence of failure rather than on assumption.

There are also implications for sovereignty: for public procurement the use cases are fairly clear. If the majority of government service tasks are explanation, form-filling, translation, summarization and guidance, and those tasks show negligible gains from frontier reasoning, then the case for domestic, open-weight or sovereign deployment stops depending on parity with the best model in the world. It only has to be good enough for the work, and the work is mostly ordinary. To keep your costs down and reduce data centre demand, don’t use higher reasoning models or settings than your tasks require.


Sources

  • Chatterji et al., “How People Use ChatGPT,” NBER Working Paper 34255 / OpenAI, September 2025
  • Anthropic Economic Index report: “Cadences,” June 26, 2026
  • Anthropic Economic Index report: “Learning curves,” March 2026
  • Sprague et al., “To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning,” arXiv:2409.12183
  • “When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling,” arXiv:2604.10739, April 2026
  • “Inverse Scaling in Test-Time Compute,” 2025
  • QuestMobile, “2026 First-Half AI Application Market Development Insight Report,” July 14, 2026 (via TechNode)
  • QuestMobile Q1 2026 AI Application Insights, March 2026
  • OpenRouter token share analysis, June 2026 (via OfficeChai, DigitalApplied)
  • VentureBeat, Kimi K3 release coverage, July 16, 2026

Featured

DeepSeek’s New V4 Flash: The Post Training Revolution

There's nothing more exhilarating, exhausting and tumultuous to be...

AI’s Accountability Gap: When AI Fails, Who Has to Report It?

An AI system can influence a medical, legal, financial...

Agentic Ratio Validation: Harness Quality Sets the Safe Range of Agency

From Agent Harness Engineering, https://openreview.net/forum?id=3hXEPbG0dh, May 2026 In the...
Jennifer Evans
Jennifer Evanshttps://patternpulse.ai
Principal, patternpulse.ai, and cofounder, Tech Reset Canada. AI policy, research and analysis. Entrepreneur since 2002, marketer since 1998, machine learning since 2009. Based in Toronto and Southeast Asia.