We fed one and the same document to seven companies' AIs all at once. Looking at the bills that came back, the 'input token count' was different for every company. This was true even though it was the exact same document, without a single character changed.
A token is the unit by which AI breaks up text as it reads, and the fee is charged by how many of these tokens there are. Where to break is decided not by a grammar that people set, but by the chunks that frequently appeared together in the vast body of text used for training. So you won't be far off if you think of it as a unit roughly corresponding to the root of a word.
Each Company Builds Its Own Ruler
The problem is that each company builds that list of chunks with its own hands. On top of that, even within a single company, the list is sometimes rebuilt as model generations change. Counting the same text on a screen that OpenAI made public, the newest generation gives 53, a model a generation or two ahead gives 57, and an older model gives 64. If it spreads this much even within one company, then between companies it goes without saying.
Length has the meter, weight has the gram, common rulers everyone shares. Yet there is no such ruler for tokens. It amounts to each measuring the same text with a different ruler.
When the Per-Unit Price Stayed the Same but Only the Bill Grew
The first place I took real notice of this problem was a talk at a FinOps event. Pointing out why estimating the cost of running AI is so tricky, it said that because each company counts tokens differently, the money that actually goes out widens greatly even with the same price list.
There is an actual case, too. This past April, Anthropic changed the way it counts tokens when it released a new model, and even though the per-unit price on the list was identical to the previous model, the same text produced up to 35% more tokens. Not a single digit of the per-unit price rose, but only the bill grew.
* FinOps : a term for the financial-practice field of tracking and managing the costs that go into cloud and AI.
Measuring Seven Companies and Sixteen Models Side by Side
So we at Polora measured it ourselves. We sent the same question, with the same report attached, to sixteen models across seven companies, splitting it into several rounds, and gathered the token counts each company returned as the basis for billing.
Setting the count OpenAI gave to 1.00, Google 1.05, Mistral 1.06, Deepseek 1.09, xAI 1.15, and Moonshot 1.19 were all roughly alike with one another. Yet Anthropic alone stood at 1.55, sitting far off by itself.
These figures are not a matter of chance that changes with each throw. Across all 121 times we put the same input into two models of the same company and the same generation, the token count did not differ by even a hair.
The Real Price List Comes Only When You Multiply the Per-Unit Price by the Quantity
Why this becomes a money problem is that the unit on the price list is that very token. Even if two companies both write '2 dollars per million tokens,' if one counts the same text at 1.55 times, the money you actually pay is 1.55 times as well. It amounts to having compared the price lists of two countries with different currencies while leaving out the exchange rate.
This is not a story that leads straight to quality. Breaking things into finer pieces has aspects that favor performance, and the companies may have set their per-unit prices with that in mind. The point is that instead of comparing per-unit prices lined up alone, you have to also multiply in how many pieces each company counts your document to see the real amount you will pay. That said, this measurement was based on a document with a lot of Korean, so if the language makeup differs, the multiplier differs too.
When you receive a quote, checking whether the standard for counting the quantity matches on both sides is as much a part of having a real conversation as scrutinizing the per-unit price. It seems the time is coming when AI fees, too, will need a common ruler like the meter for length.
The next piece will cover the results of measuring how far the fee widens when you ask the same question in English versus in Korean. For those who write in Korean, a somewhat unfair number comes out.







