The Meter Is Running. Nobody Built the Architecture to Stop It.

Search for a command to run...

Well said. “The deployment decision preceded the cost model” is probably the line that stood out most to me.
We're seeing a familiar pattern from the early cloud era playing out with AI: teams optimize for capability first, then try to understand consumption after the system is already running at scale.
The interesting part is that the solution isn't necessarily using cheaper models. It's building the architecture to match the right level of intelligence to the right task.
Routing, caching, cost attribution, governance and observability need to become part of the platform—not decisions every engineer has to make independently.
AI consumption isn't the problem. Unmanaged consumption is
The organizations that build that orchestration layer early will have a very different cost curve as agentic workloads scale.
Washington has been trying to answer one question for a decade. Is your crypto token a security or a commodity? That is it. That is the whole question. A single sentence. The kind of thing you would e

On July 10, 2026, the music industry gave AI a dress code. Two labels. "AI Generated" for tracks produced entirely by AI from a text prompt. "AI Assisted" for recordings where human musicians performe

While Washington was busy banning DeepSeek, Stripe was busy buying the pipe it runs through. That tells you everything you need to know about who actually understands the AI landscape right now. The p

I have been told what I could and couldn't do my entire career. By bosses who confused managing me with containing me. By industries that drew invisible borders around who deserved to work with whom.

AndMaverick
10 posts
Token prices fell 80% between 2025 and 2026.
Enterprise AI bills went up.
Read that again. The cost of intelligence got dramatically cheaper. The invoices got larger. That gap is not a pricing problem. It is not a vendor problem. It is an architecture problem. And it is entirely solvable by the organizations willing to stop mistaking consumption for capability.
The word entered the enterprise vocabulary somewhere around early 2026. Tokenmaxxing: treating AI token consumption as a proxy for productivity, the same way an earlier generation mistook lines of code for engineering output.
The numbers behind it are not subtle.
Meta launched an internal leaderboard in April 2026 ranking its 85,000 employees by token consumption, awarding badges like "Token Legend" to the heaviest users. The company's top user burned through 281 billion tokens in a single month. Collectively, Meta's employees consumed roughly 60 trillion tokens in 30 days. The leaderboard was shut down within days. Amazon issued internal guidance telling engineers to stop using AI "just to use AI." Salesforce was simultaneously projecting a $300 million annual AI bill and quietly shopping for model routing solutions to bring it down.
These are not cautionary tales from careless organizations. These are three of the most sophisticated technology companies in the world. And all three of them, at different points in 2026, discovered that AI consumption had quietly outpaced any measurable business outcome it was supposed to produce.
The practice has a structural cause. And that cause is not greed or carelessness. It is the absence of an orchestration layer that could have prevented it.
Here is what happened at the product level.
The UI default in every major consumer-facing AI product in 2026 is the highest-cost model in the lineup. Extended thinking. Reasoning mode. Pro tier. Whatever the label, the effect is the same. To obtain the quality that was once standard operating practice, you now select a mode that bills at a 2 to 5x output multiplier on top of an already elevated base rate.
A task that costs one dollar on a standard model can cost five to twenty dollars on a reasoning model. The reasoning happens in hidden tokens the user never sees but always pays for. Across a 5,000-person enterprise, the gap between every employee defaulting to the highest tier and model selection matched to task complexity runs into the tens of millions annually. None of it correlates automatically with measurable business outcome.
An analysis of 2.4 billion enterprise API calls found that organizations running a tiered model architecture achieved a median blended cost of $2.31 per million tokens. Organizations routing every workload to frontier models paid $18.40 per million tokens. That 87% gap is the direct financial consequence of one architectural decision made at the start of the deployment process and in most cases never explicitly revisited.
The enterprises winning on cost are not using cheaper models across the board. They are using the right model for each task. That distinction sounds obvious. It is almost universally absent.
The deployment decision preceded the cost model.
That is the sentence that appears, in some form, in every enterprise AI cost overrun analysis published in 2026. Teams ship first. Teams measure after. At chatbot scale, that sequence is survivable because volume stays small enough that the arithmetic never bites hard enough to register. At agentic scale, where the same architectural decision executes thousands of times a day, it stops being survivable.
Gartner's March 2026 analysis confirms that agentic AI models require 5 to 30 times more tokens per task than standard chatbots. Enterprises that piloted AI with single-query interfaces and then deployed multi-step agentic workflows at scale experienced cost multiplications they had not modelled. The ROI calculations that justified the agentic deployment assumed chatbot-level token consumption per workflow. The real numbers were an order of magnitude higher.
The failure is one of sequence rather than technology. Ship first, measure after, model the cost never. Then receive the invoice.
This is not a new problem. It is the cloud cost problem from 2018 wearing different clothes. The enterprises that solved cloud waste solved it with governance, visibility, and architectural discipline, not by switching cloud providers or lobbying against pricing. The same framework now applies to AI inference spend.
The organizations that are not solving it are the ones still treating AI as a procurement decision when it is an architectural one.
Here is what intelligent model routing actually looks like in practice.
A routing layer classifies incoming queries by complexity and directs summarization, classification, extraction, and formatting tasks to cost-optimized models while reserving frontier models for genuine reasoning. Complex multi-step problems, long-context synthesis, judgment under real ambiguity. The user sees one interface. Behind it, the system matches intelligence to task automatically, at the governance layer, before the token meter starts running.
Teams implementing this approach consistently reduce AI bills by 40 to 80% within 90 days without impacting the quality of output for the tasks that matter.
The technical components exist. Prompt caching delivers 90% savings on repeated context. Semantic caching identifies similar queries and serves results for near-zero cost. Tiered routing matches cost to complexity. None of it requires switching vendors, negotiating enterprise agreements, or waiting for the labs to change their pricing models.
What it requires is an orchestration architecture. A deliberate system that coordinates models, routing logic, cost governance, and human oversight into a single operating layer, rather than leaving every engineer to independently decide which model to call, at what context length, with what caching strategy, against what budget.
Without that layer, tokenmaxxing is not a behavior problem. It is the inevitable outcome of an unmanaged system operating at scale.
The enterprises that are pulling ahead on AI in 2026 share one characteristic that has nothing to do with which models they chose.
They built the governance layer before they needed it.
They treated AI spend like cloud spend, with real-time visibility, cost attribution at the team or project level, and consumption alerts before they become overruns. They shifted measurement from activity metrics like tokens consumed to outcome metrics like decisions improved, hours eliminated, and revenue protected. And they built the routing architecture that made intelligent model selection automatic rather than accidental.
The organizations still struggling built the opposite. They defaulted to the most capable and most expensive model for every task. They measured AI success by usage rather than outcome. And they are now sitting in front of invoices that cannot be reconciled with the business value they were promised.
The labs did not cause this. The pricing models did not cause this. The absence of an orchestration layer caused this. And that absence is a design choice, one that can be reversed.
The meter has been running. The question is whether anyone will build the architecture to govern it before the next invoice arrives.
Andrew Quillen is the founder of AndMaverick, a global Enterprise AI Orchestration consultancy. The MAO Framework is designed to solve exactly the architecture gap this piece describes. To continue the conversation, visit andmaverick.com.