← Back to the blog

Aptabit Blog

The real cost of per-token AI: what usage billing hides

Per-seat and per-token pricing looks cheap in a pilot and unpredictable in production. A breakdown of where the budget actually goes.

CostOperationsStrategy

A pilot with twenty users costs almost nothing. The same platform, rolled out to two thousand users with agents running in the background, is a different financial product entirely — and most teams only discover that after the rollout is already committed.

The three cost curves nobody models

Usage-based pricing creates costs that grow with success, not with value:

  • Adoption penalty. Every new user is a new line item, so the finance team ends up rationing the tool that was supposed to raise productivity.
  • Automation penalty. Background agents are the highest-value use case and the highest-volume consumer of tokens. Usage pricing taxes exactly what you want more of.
  • Retry and context tax. Long context windows, retrieval passes and retries multiply consumption invisibly. Nobody budgets for the third attempt.

How to compare properly

When you compare a SaaS assistant with a self-hosted platform, model the second year, not the first month.

DimensionUsage-based SaaSSelf-hosted platform
Cost driverUsers, tokens, featuresFixed license + infrastructure
Budget predictabilityRecalculated every quarterKnown in advance
Cost of scaling agentsGrows linearlyMarginal
Exit costData and workflows are hostagePortable by design

The practical takeaway

The right question is not "what does a million tokens cost?" — it is "what does it cost when this works and everyone uses it?" Fixed-license, self-hosted platforms flip the incentive: you optimize for adoption instead of rationing it.

Next step

Bring this to your own environment — talk to the Aptabit team about your use case.

Book a demo