AI Token Costs Move into the Enterprise: The Economics of What Should Never Have Run

Image: Depositphotos

At Jackson Hole, Federal Reserve Chairman Kevin Warsh brought AI tokens into the economic-policy conversation. In his keynote remarks, he cited reports that annualized token sales at the two leading AI labs had exceeded $100 billion, up more than 500 percent from a year earlier. He described AI as potentially a new factor of production and asked what it might mean for productivity, capital intensity, labor, market structure and the distribution of economic surplus.

Those questions matter at the level of the economy, but inside an enterprise they turn into operating decisions. Finance and technology leaders can see the price of a model call and the volume of tokens consumed. Those measures do not reveal whether a workflow was allowed to keep moving under an interpretation the institution had actually authorized.

Token Management Is Arriving in the Enterprise

Large organizations are already putting more structure around AI consumption. The Wall Street Journal reported that TIAA uses token limits based on employee role and workload, with additional consumption subject to approval. The company is also trying to distinguish valuable, compute-intensive work from activity that simply consumes more compute. Its longer-term direction is more intelligent routing, including matching a task to an appropriate model.

Token budgets, model routing, context compression, caching and smaller models all matter as organizations try to make AI use economically sustainable. They come into play after a workflow has already been allowed to run. Once an agent begins taking actions, calling tools or sending work into another system, the more consequential issue is whether it was entitled to continue on the interpretation it had formed.

TIAA's approach shows how token discipline is beginning to move into enterprise operations. A usage threshold can flag unusually high consumption and give someone a reason to look more closely, which is useful in its own right. The issue becomes more consequential when an agent is operating across policies, systems and delegated authorities. The institution still needs to know whether a policy term has been resolved as intended, whether the requested action falls within the agent's authority and whether an apparently routine interaction has become an execution decision. Token volume does not settle any of those questions.

Even with better routing, the enterprise still has to decide when an interpretation may begin to drive execution.

 

Click to expand. Figure 1. The Economics of Runtime Governance. Source: Doyle-Spare, The Economics of Agentic AI (2026).

 

From Token Consumption to Token Liability

Token Liability, as defined in my research, represents the inference-cost component of Preventable Computation: token expenditure tied to a trajectory that an institutionally available runtime control could have held before further execution. The term does not treat heavy token use as waste, since a difficult and valuable workflow may legitimately consume substantial compute. It applies to the additional reasoning, tool use or execution that proceeds under an Operational Interpretation the institution did not authorize, where a control was available and would have held the trajectory before the incremental costs arose.

Most FinOps work has focused on serving required computation as efficiently as possible. As an agent moves beyond a single response, the institution also has to decide whether the interpretation driving its next stage of activity is authorized to proceed. That decision matters because an autonomous workflow can retrieve information, reason, call tools, involve another agent and initiate downstream activity. The economic exposure from an unauthorized interpretation can quickly move beyond the original model call.

A workflow that misreads an approval condition may begin assembling an action the institution has not authorized. Its first inference may be inexpensive, but the liability grows as the system continues to reason from that interpretation, retrieves more data, calls tools, generates records, draws other agents into the matter or sends it into a human review queue. A conventional token report captures that consumption after the fact; a runtime control changes the economics by holding the trajectory before those dependent costs accumulate.

From Token Budgets to Governed Execution

A governed token strategy connects consumption decisions to execution authority. Computational proportionality still matters: the firm should use the least costly model and reasoning depth capable of performing an authorized task reliably. Routing also has to reflect consequence and authority. Summarization, customer advice and autonomous execution should not share the same admission logic simply because the same model can perform all three.

By the time a trajectory has called other agents, touched external systems, created records or reached a review queue, much of its cost has already been committed. A checkpoint at every model response would be unnecessary; the practical task is to identify where an Operational Interpretation begins to carry authority to create commitments beyond the initial interaction, because that is where the economic exposure changes. A later control may still contain damage, but it cannot retrieve the computation or operational effort already spent. Firms need to understand how early a divergence could have been recognized with the information and controls available at runtime, along with the systems and people that become involved once the trajectory proceeds. The Governance Cost Stack traces that expanding exposure.

 

Click to expand. Figure 2. The Governance Cost Stack. Source: Doyle-Spare, The Economics of Agentic AI (2026).

 

The numbers in the stack follow the order in which cost accumulates. Once an interpretation is authorized to proceed, the trajectory may draw in more reasoning, tools, systems and human work, followed by audit or remediation if the interpretation proves wrong. Token Liability is the inference-cost portion of that wider exposure: the computation a timely runtime control could have prevented.

As agentic systems move beyond answering questions and begin to act across tools and workflows, token controls will show what was spent. The larger economic decision comes earlier, when the enterprise determines whether the Operational Interpretation behind that activity has authority to create the next commitment.

Research Sources

1. Warsh, K. (2026). “In Our Time.” Keynote remarks at the 2026 Jackson Hole Economic Policy Symposium, Federal Reserve Board, August 28, 2026.

2. Bousquette, I. (2026). “How One Financial Services Company Manages Its AI Token Costs.” The Wall Street Journal, September 10, 2026.

3. Doyle-Spare, M. (2026). “The Economics of Agentic AI: Runtime Governance as a Distinct Determinant of Enterprise AI Cost.” SSRN Working Paper No. 7129939.


About the Author

Maureen Doyle-Spare

Maureen Doyle-Spare is an independent practitioner and researcher in AI governance and enterprise transformation, with more than 25 years at the convergence of technology, operations, and risk in financial services. She is the originator of the Semantic Control Plane (SCP) runtime governance architecture and the connected frameworks of Agentic Workflow Drift, Agentic Workflow Subversion, the Semantic Deviation Index (SDI), the Deterministic Gate, the Agentic Blast Radius, the Semantic Audit Trail, and the Agentic 3 C’s Framework, collectively anchored by the Doyle-Spare Agentic Governance Model (AGM). The supporting research includes the capstone reference architecture on Zenodo (DOI 10.5281/zenodo.20749051) and working papers on SSRN: No. 6459612 (Agentic Workflow Drift and Agentic Workflow Subversion), No. 6531238 (Semantic Deviation Index), No. 6674761 (Agentic 3 C’s Framework), and No. 6926219 (Semantic Layer Integrity Attack).