Token economics: The new reality of enterprise AI cost management

Get in touch

National Solutions Adviser - Data & AI

In this role, Raji works closely with CBS’ clients to leverage both data and artificial intelligence to assist clients to drive outcomes and realise significant business value. He supports and leads a team of data and AI specialists, providing technical analysis and design reviews while also facilitating knowledge sharing and mentoring.

Prior to joining Canon Business Services, Raji held positions of Head of Data and AI, General Manager – Strategy and Innovation and was also a founder of his own AI startup.
 
With more than 25 years in IT, including a decade in data science and AI, Raji has led numerous programs that have transformed businesses through data modernisation, intelligent automation, and AI.

Last updated Friday 31 July 2026
For years, enterprise software pricing followed a familiar logic: buy a licence, assign it to a user, and treat usage as largely irrelevant. AI is changing that model.

As organisations adopt copilots, generative AI tools, agents, and autonomous workflows, software economics are shifting from access-based pricing towards a mix of licences, credits, and consumption-based pricing.

In practical terms, AI cost is increasingly shaped not only by how many people have access but also by how intensively they use it. Model choice, token quantity, conversation history, context retrieval, tool calls, task complexity, and the compute resources required to deliver a response can all influence actual spend.

For leaders, that makes AI both a technology decision and a finance, governance, and operating model issue. "AI cost models are shifting from predictable licence fees to consumption-based usage, token and compute costs," explains Raji Haththotuwegama, National Solutions Advisor, Data & AI, Canon Business Services ANZ (CBS).

Why AI breaks the old software pricing model

Traditional software-as-a-service budgeting was relatively easy to predict. A business could multiply user numbers by licence cost and build an annual budget with reasonable confidence. Heavy and occasional users often carried the same unit cost.

AI introduces a different reality.

Large language models process text and other information as tokens. Input tokens represent the content sent to a model, including prompts, retrieved material, and conversation history. Output tokens represent what the model generates in response.

In token-based pricing models, providers may charge different rates per token for input tokens and output tokens. Costs can also vary according to the selected model, the volume of cached context, the number of tools, calls, and the amount of reasoning or processing involved.

A lightweight prompt may have a negligible token cost. A more complex AI workload involving data retrieval, multiple model calls, external systems, or multi-agent systems may consume far greater token volume and compute resources.

GitHub Copilot, for example, moved its plans to usage-based billing through GitHub AI Credits on 1 June 2026. Credits are consumed according to AI usage, with plans including monthly allowances and reporting to help enterprise teams understand consumption.

Microsoft has also introduced pay-as-you-go services and Copilot Cowork credits alongside fixed licensing. Administrators can connect services to billing policies, assign users or groups, monitor spending, and apply budgets or usage-based access.

This is a subtle but important change. AI systems are behaving less like fixed software subscriptions and more like cloud services or utilities.

That doesn't make consumption-based pricing inherently problematic. It means organisations need a more mature approach to AI token usage and cost optimisation.

Why finance leaders need to pay attention

Many businesses still budget for AI as though it were another fixed SaaS add-on. That assumption becomes harder to sustain as AI adoption expands from individual productivity tools to embedded AI capabilities and agentic workflows.

A power user working with AI throughout the day will not generate the same AI cost as someone who uses it occasionally for drafting or search. An agent researching a topic, retrieving financial data, querying enterprise systems, and coordinating tasks can create materially higher token consumption than a simple chat interaction.

Monthly inference costs may also change as usage patterns evolve. A successful pilot can quickly become a widely used service, while an agent may run far more frequently than enterprise teams originally anticipated.

This creates important questions for CFOs, CIOs, and technology leaders:

  • Which teams and business units are generating the most token usage?
  • Which AI workloads account for the greatest share of cloud spend?
  • What are the main cost drivers behind each AI initiative?
  • Which use cases are producing measurable business value?
  • Which models are being used for tasks that a smaller, lower-cost model could perform?
  • How much unpredictable usage can the organisation absorb?
  • What is the total cost of each AI service once integration, data preparation, monitoring, and maintenance are included?

The answer, according to Raji, isn't to stop AI adoption or minimise every dollar spent. The aim is to achieve stronger cost efficiency by matching AI spend to the value of the work being performed: “Consumption-based AI models can create unexpected cost blowouts if usage isn't monitored and governed.”

Why token economics complicates cost optimisation

Tokens have become an important economic unit in generative AI, particularly when AI systems perform work through multiple prompts, models, and tools.

But token consumption alone doesn't reveal the full cost structure.

AI costs can be distributed across:
  • Model inference and token usage
  • Cloud infrastructure and accelerated computing
  • Data storage and retrieval
  • Middleware and AI gateways
  • Security and compliance controls
  • Data cleaning, labelling, and governance
  • Integration with existing business systems
  • Monitoring and observability
  • Testing and model optimisation
  • Ongoing maintenance and operational overhead
  • Employee training and change management
This is why effective AI cost optimisation should consider total cost of ownership, not simply the advertised price per token.

A lower token cost does not automatically produce a lower total cost if the model generates weak output quality, requires repeated prompts, or creates more manual review. Similarly, a more expensive model may be justified where improved model performance reduces errors, accelerates high-value work, or protects against operational risk.

The right question isn't always, “Which model is cheapest?” It may be, “Which model delivers the best business outcome at an appropriate cost per task?”

This is where token efficiency and model routing become valuable. You can route straightforward tasks to smaller, more cost-effective models, while more capable models are reserved for complex, sensitive, or high-value work.

Good model routing can reduce wasteful spend without compromising output quality. It helps align the marginal cost of each AI interaction with its expected marginal value.

From AI governance to AI financial governance

Most enterprise AI governance programs rightly focus on privacy, data security, compliance, responsible AI and model risk. Those controls remain essential.

But a new governance pillar is emerging alongside them: financial governance.

As usage-based pricing models expand, organisations need policies that define:

  • Who owns the AI budget
  • How cost allocation will work across departments
  • Which teams can approve new AI initiatives
  • How baseline costs and growth assumptions are forecast
  • Where spending limits and quotas are applied
  • How token consumption and cloud costs are monitored
  • When additional capacity can be approved
  • How actual spend is compared with expected business outcomes

Microsoft’s usage-based Copilot Cowork reflects this shift. Administrators can use central cost management tools to allocate Copilot Cowork Credits, apply policy-based limits, view spending, and operate across prepaid and pay-as-you-go pricing models.

This is becoming the AI equivalent of FinOps: a collaborative discipline connecting finance, technology, and business teams to improve cost visibility, resource allocation, and accountability.

“AI is moving from license-based predictability to usage-based consumption," Raji explains. "That means finance, technology, risk, and business leaders need visibility into what is being used, by whom, and at what cost.”

Raji Haththotuwegama, National Solutions Advisor, Data & AI at Canon Business Services ANZ (CBS)


The hidden risk: Running out of AI capacity

One of the less discussed consequences of token economics is operational disruption.

In a fixed-license model, software usually remains available regardless of how heavily it is used. In a consumption-based model, access to some AI capabilities can be shaped by included credits, spending limits, quotas, or billing policies.

GitHub Copilot plans now include monthly AI Credit allowances. Enterprise customers can purchase additional usage and use consumption reporting to understand how credits are being used across their organisation.

Microsoft’s pay-as-you-go services similarly allow administrators to create billing policies, scope services to users or groups, monitor spending, and set optional budgets.

A budget notification is not necessarily the same as a hard service cap, and the effect of reaching a limit will depend on the particular product and configuration. Nevertheless, organisations need to understand precisely what happens when an allowance, quota or payment threshold is reached.

If an AI system is embedded in customer service, knowledge management, software development, or automated business processes, unexpected restrictions can become more than a billing inconvenience. They may delay workflows, reduce productivity, or interrupt an important business function.

AI capacity therefore needs to be actively monitored, much like cloud capacity, storage growth or other critical technology resources.

What a practical AI consumption model looks like

Organisations that handle this transition well are unlikely to be those with the greatest number of specialised AI tools. They will be the ones with the clearest operating model around consumption, cost control, and value.

Establish a cost baseline

Begin by identifying the baseline costs associated with licenses, infrastructure, token usage, data preparation, integration, security, and maintenance.

A total cost model should distinguish between one-off implementation expenses and recurring operating costs. It should also account for the way total cost may change as usage expands.

Allocate ownership

AI spend shouldn't sit as an unexplained line in a central cloud budget.

Cost allocation by business unit, team, project, or AI initiative creates clearer ownership. It also helps finance and technology leaders identify which use cases justify further investment and which require redesign.

Track usage patterns

Unified cost visibility should show token volume, model selection, cost per task, high-consumption users, expensive agents, and unexpected changes in cloud spend.

Traditional monitoring tools may show infrastructure usage without revealing what drove the token count or which business process generated it. AI-specific monitoring or middleware can connect token usage with users, models, applications, and cost centres.

Set alerts and spending limits

Real-time monitoring and automated anomaly detection can alert managers when AI costs escalate unexpectedly.

Useful thresholds may include:
  • Daily or monthly AI spend
  • Token consumption by user or application
  • Rapid changes in inference workloads
  • Unexpected use of high-cost models
  • Unusual tool-call volume
  • Cost per successful transaction
  • Budget burn rate against forecast
  • Early alerts give enterprise teams time to investigate unusual trends and make proactive budget adjustments before AI costs spiral.

Improve resource utilisation

AI cost optimisation should operate close to the inference path—the point at which the system selects a model, sends a request, and generates a response.

Practical strategies may include:
  • Routing routine tasks to smaller models
  • Limiting unnecessary conversation history
  • Reducing excessive prompt or context size
  • Caching frequently reused information
  • Managing output-token limits
  • Batching appropriate AI workloads
  • Reducing repeated or unsuccessful requests
  • Reviewing multi-agent systems for duplicated work
  • Removing underutilised AI tools and infrastructure

These measures support resource optimisation without treating lower consumption as the only measure of success.

Measure cost against business outcomes

Cost reduction is only one possible AI outcome. An initiative may justify higher AI spend if it creates greater revenue, stronger customer service, faster decision-making, or lower operational risk.

Useful performance measures may include:
  • Cost per task completed
  • Cost per customer enquiry resolved
  • Cost per document processed
  • Cost per employee hour saved
  • Cost savings from automation
  • Improvement in output quality
  • Reduction in error or rework rates
  • Revenue or business growth supported
  • Service quality or customer satisfaction improvement
  • Return generated for every AI dollar spent
This shifts the conversation from “How many tokens did we use?” to “What value did that token consumption create?”

AI cost optimisation is most effective when it protects high-impact uses rather than applying blanket restrictions. Organisations should invest first in clearly defined AI initiatives that can produce measurable business outcomes, learn from those implementations and expand incrementally.

Why the uplift process matters

As AI becomes embedded in daily work, extra capacity should not necessarily be treated as an automatic entitlement.

Organisations can introduce a controlled uplift process for users, teams or agents that exceed their baseline allocation. The process need not create unnecessary bureaucracy, but it should answer four questions:
  • What is the business case?
  • Who approves the additional usage?
  • Which budget or cost centre will fund it?
  • What measurable value is expected in return?

An uplift request may also prompt a cost optimisation review. Before raising a limit, the organisation can assess whether excessive token quantity, unsuitable model selection, repeated tool calls, or weak workflow design is creating avoidable AI cost.

This turns AI consumption into a managed enterprise resource rather than an invisible cost leak.

The bigger shift now underway

GitHub Copilot’s move to usage-based billing is one visible sign of a broader industry trend. Microsoft’s growing use of pay-as-you-go AI services and Copilot Credits is another.

The exact pricing models will vary between AI providers. Fixed subscriptions, included allowances, token-based pricing, and usage-based services are likely to coexist.

But the direction is clear: as AI capabilities become more agentic, integrated, and computationally intensive, organisations will need deeper insight into usage patterns, compute costs, and cost structures.

AI adoption is moving from experimental spending towards strategic investment. That transition will demand the same discipline that organisations developed around cloud cost management, but with additional complexity created by models, tokens, agents, data, and variable output quality.

From token spend to business value

The leadership challenge isn't simply to minimise AI spend. It's to connect AI cost with measurable business value.

Organisations that scale AI successfully will be the ones that:
  • Establish AI financial governance early
  • Create shared accountability across finance, technology, risk, and business teams
  • Budget for licenses, token consumption, and total cost of ownership
  • Maintain real-time cost visibility
  • Apply spending limits and anomaly alerts where appropriate
  • Use model optimisation and model routing to improve cost efficiency
  • Prioritise high-impact AI initiatives
  • Invest in data quality, governance, and employee training
  • Connect every material AI investment to a clear business outcome

In the next phase of enterprise AI, intelligence alone will not be enough. Token economics, cost control, and operational discipline will matter just as much.

AI adoption becomes more sustainable when organisations can connect usage, cost, and governance to real business value. Canon Business Services ANZ helps customers build the visibility, controls, and operating models needed to manage AI consumption with confidence as technology and pricing models evolve.

Get in touch to explore how Canon Business Services ANZ can help you govern AI spend, reduce risk, and scale AI in a way that’s practical, measurable, and commercially sound.

Similar Articles

View all

AI agents vs automation

Uncover the key differences between AI agents and automation. Learn how each technology can improve workflows and drive smarter decisions for New Zealand businesses.

AI automation and the future of work

Uncover how AI automation is transforming the future of work in New Zealand. Learn about the latest trends, impacts on jobs, and strategies to adapt.

A guide on AI fraud detection

Explore how AI fraud detection enhances security of businesses in New Zealand. Learn about machine learning algorithms, benefits, challenges, and best practices.

The impact of AI on business productivity

Discover the artificial intelligence's impact on business and how it revolutionises operations. Protect your business data with CBS New Zealand's expert insights now!

What is Security Automation?

Learn how automated security transforms cybersecurity, making it simpler and more efficient. Protect your business data with CBS New Zealand’s expert insights now!

What are the benefits of machine learning in business?

Explore the myriad benefits of machine learning in business. Learn how ML enhances efficiency and drives innovation for sustainable growth in New Zealand. Discover now!

BPO vs RPO: Choose the right outsourcing solution

Learn the key differences between BPO and RPO to make an informed outsourcing choice. Discover benefits, use, and expert insights with Canon Business Services New Zealand!

What are the challenges of AI in financial services

Discover challenges of AI in finance, tackling bias, security, and integration for ethical, efficient financial services. Protect your business data with CBS New Zealand's expert insights now!

How to use Cloud-based AI & Machine Learning for businesses

Unlock the potential of cloud-based AI and machine learning with CBS New Zealand's expert insights now!

Differences between Copilot and ChatGPT

Compare Copilot and ChatGPT to understand their unique capabilities. Explore how each tool can enhance productivity and creativity in different contexts.

Feature comparison between Copilot and Copilot Pro

Compare Microsoft Copilot and Copilot Pro, exploring their features, benefits, and value propositions for organisations in New Zealand. Read more.

The difference between chatbots and AI agents

Explore the differences between Chatbots and AI agents. Understand their unique capabilities and how they enhance user experiences for New Zealand businesses.