Alex Karp asked the right question in the most provocative way.
If frontier AI companies create so much value, why do they charge for tokens instead of sharing in the outcome?
The question is commercially interested. Palantir sells an alternative vision of enterprise AI, built around controlled deployment, proprietary data, and operational workflows. But the question landed because many enterprise customers were already asking a less theatrical version of it:
What are we actually paying for?
Handelsblatt reports growing customer resistance to the pricing of OpenAI and Anthropic, alongside frustration about control and results. Its follow-up on cost management makes the operational problem explicit: AI spend is rising quickly, but few companies know precisely what they are paying for.
Vanessa Cann, who contributed to the reporting, describes the transition well. During the pilot phase, consumption was too small to dominate the business case. Once companies scaled useful systems—especially agents, where one user request can trigger many model calls—the hidden cost structure became visible. The new question is no longer whether a pilot works. It is what value a specific workflow creates when operated repeatedly.
This is not an AI rejection cycle.
It is the beginning of an AI unit-economics cycle.
The Token Is a Real Meter and the Wrong Management Unit
A token is not imaginary. It measures part of the work a model performs, gives engineers a way to understand consumption, and gives providers a way to invoice variable usage.
But a token has almost no meaning to a business owner.
A claims executive does not want tokens. She wants a correctly processed claim. A support leader wants a resolved case. An engineering manager wants a safe change merged. A compliance officer wants an investigated exception with an auditable decision.
Tokens sit several layers below those outcomes.
They also fail as a simple comparison unit. One token from a small classifier, one from a frontier reasoning model, and one produced inside a long agent loop are not economically equivalent. Provider price lists now distinguish input, output, cached input, long context, service tiers, tools, and batch processing. The public catalogs from OpenAI, Anthropic, and DeepSeek show how wide the price and product range has become.
Those options are useful for optimization. They do not answer whether the workflow earns its place.
The FinOps Foundation's work on token economics makes the distinction clearly. Token cost is only one layer in a wider stack that includes infrastructure, data, networking, engineering, observability, evaluation, governance, and labor. It also notes that a retrieval pipeline with reasoning and multiple tool calls can consume one or two orders of magnitude more tokens than a direct request to a smaller model.
The provider needs a production meter.
The enterprise needs an economic ledger.
Confusing the two is the root of the rebellion.
Scale Exposes What the Pilot Hid
Pilot economics are forgiving.
The user group is small. The workflow is supervised. Failures are treated as learning. Integration and governance labor are often funded centrally. Consumption may sit inside a trial, a generous subscription, or a budget no one has yet challenged.
Production removes those protections.
Every successful deployment increases volume. Agents add planning, retrieval, tool calls, retries, verification, and memory. More users create more edge cases. Higher autonomy requires better evaluation, monitoring, access control, incident response, and human escalation. A system that looked cheap per interaction can become expensive per completed process.
The evidence is still emerging, but the warning signs are strong. In a May 2026 survey with 75 qualified enterprise respondents, McKinsey found that spend rose nearly fourfold as organizations moved from isolated use cases toward enterprise adoption, while 93 percent reported exceeding their AI budgets. The same analysis cites research finding up to 30-fold variation in token use when an agent executes the same task.
That is a small survey, not a universal market estimate. But it describes a structural problem: agentic demand is nonlinear, while most budgets assume a stable relationship between users, requests, and cost.
This is now a mainstream operating concern. The State of FinOps 2026 report says 98 percent of FinOps teams manage AI spend, up from 31 percent two years earlier, and names AI cost management as the leading skill gap.
The market has moved from Can we build it? to Can we operate it economically?
Karp Is Right About Value, but Outcome Pricing Is Not Magic
Karp's implied alternative is attractive: if a vendor claims to create business value, let it charge for the value instead of the tokens.
Parts of the market are already moving that way.
Intercom now prices its Fin agent by outcomes. A support resolution, procedure handoff, or disqualification costs $0.99; a sales qualification costs $9.99. Zendesk similarly uses automated resolutions as the billing unit for its AI agents.
That is better aligned than charging a customer for every internal inference. It makes the vendor carry more performance risk, and it gives the buyer a unit that resembles work.
It also reveals why outcome pricing is difficult.
An outcome is not a natural object. It is a contract over a definition.
Intercom counts a resolution when the customer confirms success or does not request more help after the last answer. That is a reasonable operational definition, but it is not the same as proving that the customer's underlying problem was solved, that the answer remained correct, or that no downstream cost appeared later.
The difficulty increases outside customer support:
- Who gets credit when AI drafts a proposal but a salesperson closes the deal?
- Is a merged code change an outcome if it creates a production incident a week later?
- Is time saved real value if the organization does not redeploy the capacity?
- How should a vendor be paid for risk avoided, cycle time reduced, or a better decision?
- Which baseline determines the gain: the previous human process, the best alternative model, or doing nothing?
Outcome pricing can align incentives. It cannot eliminate attribution, delayed effects, gaming, or disagreement over the counterfactual.
Tokens answer a supplier question: how much model capacity was consumed?
Outcome pricing answers a commercial question: what event triggers payment?
Neither, on its own, answers the management question: did this workflow create durable net value?
The Enterprise Needs Workflow-Level Unit Economics
The right unit sits between the token and the annual ROI slide.
It is the business event: a case resolved, claim processed, change reviewed, order recovered, investigation completed, or application approved.
For each event, the organization needs a complete cost and outcome record.
| Ledger component | What to include |
|---|---|
| Variable AI cost | Model tokens, tool calls, search, code execution, retrieval, and third-party APIs |
| System cost | Orchestration, storage, observability, evaluation, security, platform allocation, and integration |
| Human cost | Review, escalation, exception handling, correction, training, and change management |
| Risk cost | Expected loss from errors, compliance failure, service interruption, and customer harm |
| Outcome | Completion, acceptance, quality, latency, revenue, capacity released, or risk avoided |
The basic management equation is simple:
Total workflow cost / accepted outcomes = cost per successful outcome
The difficult work is defining “accepted” honestly and measuring the full numerator.

This is why “hours saved” is usually an intermediate metric, not a final one. Saved time becomes economic value only when it changes something observable: more work completed, faster revenue, lower contractor spend, avoided hiring, reduced backlog, better quality, or genuinely redeployed employee capacity.
The same discipline applies to quality. A cheaper model is not cheaper if it doubles review time. A frontier model is not expensive if it prevents a high-consequence error. A human is not inefficient if the automation requires so much monitoring and rework that the total process costs more.
The unit-economics ledger makes those comparisons possible.
Frontier, Smaller, Open, and Human Are Portfolio Choices
The cost backlash will accelerate model substitution, including toward smaller and open-weight systems. That is healthy, but “use a cheaper model” is not a strategy.
Every workflow should be routed to the least expensive execution path that meets its quality, latency, control, and risk requirements.
That portfolio may include:
- Frontier models for complex, ambiguous, or high-consequence work where stronger reasoning materially changes the result.
- Smaller hosted models for bounded, repetitive, high-volume tasks with measurable acceptance criteria.
- Open-weight models where scale, data control, customization, or continuity justifies the infrastructure and operating burden.
- Humans for exceptions, contested decisions, relationship-sensitive work, or tasks where review and failure costs erase the automation advantage.
Open weights move cost; they do not make it disappear. Provider margin may be replaced by accelerator capacity, platform engineering, security, evaluation, and operations. The correct comparison is total cost per accepted outcome, not API price per million tokens.
This also changes the role of the frontier model. It becomes a scarce capability used deliberately—not the default engine behind every summarization, classification, and agent loop.
The Control Plane Becomes the Economic System of Record
I previously argued that AI token costs need an operating model, with workload routing, attribution, budgets, approvals, and exceptions enforced close to runtime.
The follow-on is that the same control plane must connect consumption to outcomes.
It should record, for every production workflow:
- the business event and owner,
- the models, tools, and data used,
- token, infrastructure, and third-party cost,
- latency, retries, fallbacks, and failed loops,
- human review and correction time,
- the evaluation or acceptance result,
- the realized business metric,
- and the model or human alternative available next.
Without this layer, a company can cut tokens while destroying value. It can also celebrate productivity while quietly increasing review labor, operational risk, or vendor dependence.
With it, cost control becomes dynamic. The system can route routine work downward, escalate difficult cases upward, cache stable context, stop runaway loops, enforce budget thresholds, and compare providers using the company's own tasks rather than generic benchmarks.
That is more than FinOps reporting.
It is an operating system for the economics of intelligence.
Procurement Will Move Toward Hybrid Contracts
I do not expect token pricing to disappear.
It remains useful for APIs, infrastructure, experimentation, and any workload where the provider cannot observe the customer's final business result. Pure outcome pricing will work best where the event is narrow, frequent, attributable, and quickly verifiable—support resolution is a good example.
Most enterprise agreements will therefore become hybrids:
- capacity commitments for predictable base demand,
- consumption pricing for variable usage,
- service tiers for latency and availability,
- outcome components where success can be defined,
- and commercial credits or risk sharing when agreed performance is missed.
The buyer's leverage will come from instrumentation and portability. Enterprises should require exportable usage data, workload-level attribution, clear outcome definitions, auditable quality measures, cost ceilings, model-substitution rights, and a tested path to another provider or an open-weight deployment.
The ability to measure and switch is what turns vendor competition into enterprise bargaining power.
The Rebellion Is a Maturity Signal
Enterprises are not discovering that AI has no value.
They are discovering that adoption without economic instrumentation is not a strategy.
Karp is right that a large token bill does not prove a large outcome. He is also right that frontier providers should face harder questions about the value they capture and the control customers retain.
But changing the invoice from tokens to outcomes does not release the enterprise from measurement. It makes measurement more important.
The practical answer is two ledgers:
The provider may meter tokens.
The enterprise must manage successful workflows.
When those ledgers are connected, the cost debate becomes productive. Companies can decide where frontier intelligence earns its premium, where a smaller or open model is sufficient, where a human remains cheaper, and which workflows should not exist at all.
That is not rebellion against AI.
It is the moment AI becomes accountable to the business.
Sources
- Handelsblatt's reports on enterprise pushback against OpenAI and Anthropic and four principles for controlling AI spend frame the immediate market discussion.
- Vanessa Cann's LinkedIn post on the move from pilots to scaled, value-measured workflows provides the practitioner context behind the reporting.
- The FinOps Foundation's State of FinOps 2026, token-economics analysis, and guidance for SaaS model-token costs document the growth of AI spend management and the wider cost stack around tokens.
- McKinsey's enterprise AI FinOps analysis reports the May 2026 survey results and recommends outcome-level attribution, routing, and an AI control plane.
- OpenAI, Anthropic, and DeepSeek publish current model and feature prices in their official OpenAI API pricing, Claude pricing, and DeepSeek API pricing documentation.
- Intercom documents outcome pricing for Fin; Zendesk documents automated-resolution pricing for AI agents.
