In the agentic era, tokens are the new currency, and ‘tokenomics’ is the new economic model. But how are organizations to embrace tokenomics?
In his keynote address at Splunk .conf26, Cisco’s President and Chief Product Officer Jeetu Patel outlined a fundamental shift in the AI landscape, moving from interactive chatbots to autonomous agents that function as digital coworkers.
Every model call (prompt + response) is measured in tokens. Agents, unlike humans, run 24/7, call other agents, and hit multiple models/tools per workflow. That’s why agents already consume about 5x more tokens than humans, and token usage is growing faster than human-driven usage.

Jeetu Patel, President and Chief Product Officer, Cisco
This transition has introduced several critical economic and technical vectors:
- Inference as workload: Unlike chatbots, agents require sustained, high-level infrastructure consumption. In February 2026, agent token consumption exceeded human consumption, and it has since increased 14x, with agents now consuming roughly 5x more tokens than humans.
- Compute shift: It is projected that 60% of global AI compute capacity will be consumed by inference rather than training.
- Edge computing: The rise of “edge-side computing” will see trillions of agents running across distributed environments, including every device and workflow.
- Bandwidth demands: Agents are highly resource-intensive, consuming approximately 450% more network bandwidth than humans for the same tasks.
In this context, ‘tokenomics’ is the economic model around tokens used by AI models — how they’re consumed, priced, and optimized across applications and agents. So tokenomics in practice is the “cost discipline layer” of an organization’s AI strategy.
In the agentic era, tokens are the new currency. But how are organizations to embrace tokenomics? In a conversation with Hao Yang, VP AI, Splunk, I drew out a plausible tokenomics playbook.

Hao Yang, VP AI, Splunk
Key pillars of a tokenomics playbook
According to Hao, the crucial elements of an organization’s tokenomics playbook break down into a few clear pillars:
1. Start with value, not cost
Hao believes that tokenomics should not start from a cost-cutting mindset.
“When it comes to tokenomics, as with any business operations, you have to first think about value, because there is no point of talking about costs without talking about value,” he said. “If there is no value, then you shouldn’t spend a penny on it. If there is a lot of value, then you should really figure out how to fund this.”
Implications for an organization’s tokenomics playbook:
- Define clear business outcomes for each AI initiative (e.g., MTTR reduction in SOC, ticket deflection in support, uptime gains in SRE).
- Tie token consumption to those outcomes (tokens per incident resolved, per workflow automated, etc.).
- Use this to decide which AI use cases deserve scale-up funding versus those that remain as experiments.
2. Fiscal discipline and visibility into spend
Hao stressed that, in the current macro environment, budgets are under scrutiny: “From any responsible company’s perspective, that fiscal responsibility is also a real thing… we want to make sure that we have visibility into the spendings.”
For a tokenomics playbook, that means:
- Establish monitoring and reporting for token usage across use cases (SOC, observability, customer service, etc.), teams/business units, and model types (frontier vs open source).
- Set budgets or soft quotas per team/use case, reviewed with finance and product owners.
- Make token spend transparent at exec level (CFO/CEO/board), since AI costs are becoming “pretty real” as deployments scale enterprise-wide.
3. Manage unpredictability from agentic workloads
The move from simple chatbots to autonomous/agentic systems changes the cost profile dramatically.
“Last year, if you are using a chatbot, it’s limited by how many times you’re going to interact with the chatbot,” said Hao. “Now you have an agent, and the agent is going to be there doing all these tasks – that’s a huge thing from the tokenomics perspective because, all of a sudden, the agent can burn tokens pretty quickly. So, before you know it, it’s already a huge consequence.”
Key playbook elements here include:
- Treat agents as always-on consumers, not as occasional queries.
- Model and simulate worst-case token burn when agents loop, retry, or traverse long workflows, or when incidents spike (e.g., large outages, major attacks).
- Instrument per-agent cost profiles (average and P95/P99 tokens per task).
- Define usage patterns that are allowed for agents (e.g., which workflows can trigger long loops vs which must stay bounded).
4. Guardrails and hard controls
Because of that unpredictability, Hao argued for the need for active control, not just observability: “Another thing is really to have guardrails and controls. When something goes wrong, you have the ability to stop it there right away, rather than having to deal with the mess later.”
In a playbook, that translates into:
- Spending guardrails
- Hard caps per project/agent per day or month.
- Alert thresholds when spend or tokens/time cross certain limits.
- Runtime controls
- Kill switches for specific agents, workflows, or models.
- Max tokens per call, per session, per incident.
- Policy: what gets automatically throttled, paused, or shut off when spend anomalies appear.
- Governance flows
- Clear ownership: who can override caps and under what conditions.
- Incident process for “runaway” agent spend (detection, triage, root cause, remediation).
5. Cost–architecture choices: models and deployment
Hao linked tokenomics to model strategy, especially as enterprises scale: “Frontier models can be very expensive. If you use that in all your tasks, it might actually have some of the tokenomics problem. That’s why we’re also looking at open-source models to specialize in some of those deployment scenarios or different kinds of tasks.”
For the playbook, organizations should:
- Define model selection rules
- Frontier models for high-stakes, high-complexity reasoning (e.g., long-horizon cyber defense).
- Smaller/open models for routine tasks, air-gapped or restricted environments, or high-volume, lower-value operations.
- Encourage tiered routing
- Start with cheaper models, escalate to frontier models only when needed.
- Consider deployment constraints
- On-prem/air-gapped environments where some frontier APIs are unavailable.
- Latency vs cost trade-offs for SOC/SRE scenarios.
6. Enterprise-scale lens: from experiments to full enterprise scale
Hao distinguished between experimentation and scale: “A year ago, they might be thinking about AI as experimentation and cost is usually not a big concern. But now you see all the great potential, you’re going to roll that out for every employee. For some of the larger organizations, they can have 100,000s of employees, and when you do all the math, that cost has become something pretty serious.”
A robust playbook should:
- Explicitly differentiate pilot economics from production economics, with different approval thresholds for “lab” vs “enterprise rollout”.
- Include an industrialization phase. Before scaling to whole-of-enterprise, require cost projections, stress tests, guardrail configurations, and clear value metrics.
- Integrate tokenomics into board-level and C-suite discussions as adoption scales.
In summary
- Value-first framing (why spend tokens at all).
- Strong visibility and fiscal discipline (where, by whom, and for what are tokens burned).
- Explicit handling of agent unpredictability (new cost dynamics vs chatbots).
- Guardrails and controls (alerts, caps, kill switches).
- Model and deployment strategy tightly linked to cost.
- A scale-up discipline for moving from experiments to enterprise-wide rollout.
Jeetu Patel summed it up this way: “That’s the essence of tokenomics in the agentic era: combining architecture, routing, and behavioral controls so that inference at massive scale remains economically sustainable, while still unlocking the value of autonomous agents.”