IT Departments: Not Every Department Uses AI the Same Amount. Do Your Chargebacks Reflect That?
Chargeback isn't new, and neither is the gap between what gets measured and what gets estimated. AI just made that gap too big, and now too cheap, to ignore any longer.

IT departments have been doing chargeback since the mainframe - CPU seconds, actual seconds, metered and billed to the department that used them. That was standard practice decades before anyone said "cloud," much less "AI." Chargeback itself is not new, and neither is the idea of measuring usage precisely enough to bill it fairly.
That gap, between what could be measured and what actually was, is not new either. It’s as old as chargeback itself, and for most of IT's history it produced two different kinds of numbers side-by-side.
Some costs were measured precisely, because the line item was big enough or the technology was cheap enough to justify the effort. CPU time, going back to the mainframe. Print volume, tracked per page for decades because the hardware to do it was trivial. Network bandwidth, once tools existed to watch it. Help desk hours, because a ticket system counts itself.
Everything else ran on reasonable estimates, and everyone understood them as estimates, not measurements. Shared infrastructure cost got split by headcount, because nobody was going to install a meter on every employee's desktop. It got split by device count, updated quarterly, a defensible proxy for a mostly static allocation. It got split by seat count on the applications each department licensed, by storage quota assigned rather than storage actually touched, sometimes by nothing more precise than each department's share of the overall budget. None of this was laziness. Measuring precisely would have cost more than the value of getting that particular number exactly right, so an estimate was the rational answer.
AI's rapid growth changed the math on that trade, on both sides at once. The line item got big enough to matter in a way office electricity or shared printer toner never did. And most of the technology to measure AI's usage precisely, request by request, customer by customer, department by department, has been sitting on the shelf for a few years now, largely thanks to OpenTelemetry becoming a standard most modern applications already speak. OpenTelemetry is a shared, open standard, not owned by any single vendor, and it quietly became available while most finance and IT teams were still building their chargeback models the old way. For the first time, costs that used to sit comfortably in the "estimate it" column have crossed into the "you can actually meter this now" column. Most organizations have not caught up to that. The raw data exists. What’s missing is technology that turns that raw data into a number a department head can actually be handed, and can actually defend.
Is it time to get up to date? Here are the six places this shows up:
1. Model selection and routing. Different teams reach for different models for different tasks, and the cost difference between them can be significant. That difference is metered at the API level the moment the call happens. Most finance teams still find out about it the old way, after the fact, from a bill total, rather than from the measurement that already exists. The practical effect on a chargeback model is that the department is billed as one lump sum, so the team that chose the expensive model and the team that didn't get charged the same average, and neither number is actually true.
2. Overages and enforcement. A pilot project left running over a weekend is a metering problem with two parts: what happened, and who caused it. The usage itself is easy, it was recorded in real time the moment it occurred. Attributing it to the specific team or project that left it running is the part that actually makes it a chargeback fix rather than just a louder bill, and it depends on whether the usage data carries that context all the way through, not just on whether an alert exists.
3. Performance and caching. Redundant queries and repeated calls are measurable the instant they happen. The reason they still surprise finance teams is not that the data is missing. It’s that nobody connected the measurement to a department or a decision maker who could act on it.
4. Quality and reliability. This is the one piece that resists metering the way the others do. A bad output costs money later, in rework or a support escalation, and by the time it does, the cost often lands in a different department than the one that caused it. That’s real money, but whose chargeback line it belongs on is genuinely unclear, closer to the "estimate it" side even with good instrumentation everywhere else.
5. Governance and risk. Similarly indirect, and arguably harder to attribute than any of the others. You can measure which data goes to which model provider in real time. But once a compliance gap turns into an actual fine or audit finding, it's much harder to say which team's chargeback line should absorb that cost. By then, it's usually the whole organization's liability, not any single team's.
6. Attributing the cost. It’s not the total bill; central IT can always produce that. The question is whether the number assigned to each department reflects what actually happened, measured, or reflects a proxy dressed up to look like measurement. The headcount, device count, seat count, and storage quota are all still doing quiet work in a lot of chargeback models today. They stand in for numbers that could now be measured directly if anyone built the system to do it.
Six piles, and one clear line running through them. Four of them - model selection, overages, caching, and cost itself, generate a metered event the moment they happen. The data already exists somewhere. Two of them, quality and governance, arrive secondhand, as rework, an escalation, a fine, an audit finding, and are harder to meter directly no matter how good your instrumentation is elsewhere. Even among the four, as the examples above show, recording the event is the easy half. Attributing it to the specific team or decision maker who caused it is usually the part still missing.
That split is worth sitting with before your next budget cycle, because it points to a sharper question than "is AI expensive?" If every AI cost were perfectly metered and attributed tomorrow, would your chargeback model actually be fair?
Probably still no, and the reason has nothing to do with AI. Headcount, device count, seat count, and storage quota were reasonable estimates when they were built, a fair trade given what could be measured at the time, and nobody who built them was cutting corners. AI did not expose them as wrong. It simply became the first line item large enough, and new enough, to make someone ask whether those old proxies still had to be estimates, or whether they could finally be measurements instead. For AI, and for the other costs that generate a metered event the instant they happen, the honest answer is that they can.
To be clear, not everything gets to move from the estimate column to the measured column just because the technology improved. Quality and governance costs are likely to stay proxies for a long time, because their damage shows up somewhere else, later, in a different budget line entirely, and often in a different department than the one that caused it. That’s a reasonable limit, not a failure, and no chargeback model should pretend otherwise.
But for the parts of your infrastructure spend that do generate a metered event the instant they happen (compute, storage, data transfer, API calls, AI or otherwise), there is no longer a good reason to allocate by proxy what you could now attribute directly to the team or project responsible. The mainframe gave way to client server, which gave way to cloud, which is now giving way to whatever AI infrastructure becomes next. The question a chargeback model has to answer never changes: did this actually happen, who caused it, and can you defend that number to the department paying for it? For decades, the honest answer for most infrastructure cost was "we can't measure that precisely, so here is our best proxy." For a meaningful and growing share of it, that’s no longer true. Continuing to guess where you could now measure is a choice, not a limitation.
Request a demo
About the Author
Alan Cox founded Beakpoint after experiencing firsthand the frustration that comes with mysterious cloud costs. As a technology leader who has spent over two decades building and scaling software organizations, he's seen how cloud expenses can spiral out of control.





