When an F1 car speeds around the track, we usually notice the driver, the car, and how fast it is going. Less noticeable is what it takes, and what it costs, to keep it all running.
AI works in much the same way. We usually focus on the answer it gives us, not what is happening behind the scenes or what that response costs to produce. Every time we send AI a prompt, it breaks it down into small units called tokens.
They may be largely invisible to us everyday users, but tokens play an important role in how AI works, how its usage is measured and how businesses pay for it.
When AI Became Conversational
For years, artificial intelligence (AI) worked quietly behind the products people used. Machine learning systems could detect fraud, recommend products and make predictions, but building or adapting them usually required engineers and specialised technical teams.
Large language models (LLMs) changed how people access that capability. Since ChatGPT’s launch in 2022, everyday language has become an interface: people can ask AI to write, summarise, analyse and create without needing to code.
To us, that interaction feels like a conversation. To the AI model, those words and sentences are processed as tokens.
So, What Exactly Is a Token?
Tokens may represent a whole word, part of a word or even punctuation in your prompt. Every 100 words typically translates to around 130 to 140 tokens.
Prompt → Tokens → AI Model → Tokens → Response
Tokens are not intelligence, and more tokens do not automatically mean better work. They are simply units models use to process and generate information.
But tokens have another important role: they make AI usage measurable, almost like a currency. What goes into the model counts as input tokens, while what it generates counts as output tokens.
And because tokens can be counted, they can also be priced.
Paying for Access vs Paying for Usage
If your company pays for software, you probably know the drill: one employee, one account, one monthly fee. This is known as seat-based pricing. Microsoft 365, for example, still prices business plans on a per-user, per-month basis.
With AI, pricing can work differently. When businesses access AI models through APIs or usage-based enterprise plans, charges can depend on the number of tokens processed.
Compare an administrative employee who occasionally uses AI to rewrite emails or summarise short documents with an analyst who regularly uploads lengthy reports, analyses datasets and generates detailed research. Both may occupy one “seat”, but their AI consumption can be vastly different.
With token-based pricing, instead of measuring only who has access, businesses can also be charged for how much AI is actually used.
Here is where it gets slightly more complicated: not all token usage is priced at the same rate.
Different models charge different rates, and input and output tokens can also carry different prices. Think of an F1 car and a Perodua Myvi: they are built for very different levels of performance, and their running costs reflect that. AI models work in much the same way.
So, when it comes to AI costs, it is not just how many tokens are used, but what type of tokens are being processed and by which model.
When AI Usage Starts Adding Up
For an occasional ChatGPT user, a short conversation may use only a few hundred to a few thousand tokens and cost very little. Multiply that across thousands of employees, however, and those small units quickly add up.
But once something becomes easy to measure, there is also a temptation to assume that more must mean better.
As companies raced to adopt AI, token consumption became an easy way to measure whether employees were actually using it. On paper, the logic seemed reasonable: if token consumption was rising, employees must be using AI more. But greater adoption does not automatically mean greater value.
Just as having the fastest car does not guarantee winning the race, using more AI does not guarantee a better outcome. Tokens can measure AI consumption. But does more consumption actually create more value?
Coming up next: Why going faster and spending more does not always get you further, and why tokenmaxxing may not be the answer.
Acknowledgements:
Thank you to the Marylin Isaacs, College Lecturer and the Sunway iLabs team for their invaluable contribution and insights in preparing this article.
References
Camacho, E. (2024). Why I changed how I pitch AI: It’s no longer about saving money, but managing tokens and adoption. CIO.
OpenAI. (2026). Understanding and counting tokens | OpenAI help center. OpenAI Help Center.
Perez, S. (2026). Reid Hoffman weighs in on the ‘tokenmaxxing’ debate. TechCrunch.
Schneider, J., Abraham, J., Linderman, M., & Sastry, N. (2025). The AI-centric imperative: Navigating the next software frontier. McKinsey & Company.


