Guide
How should AI agencies bill clients for OpenAI and Anthropic usage?
Agency founders and finance leads face the same question once OpenAI and Anthropic usage becomes a recurring line item: how should client contracts reflect AI provider costs? There is no single correct answer — the right model depends on predictability, margin, administrative capacity, client expectations, and how material AI spend has become. This guide compares five common approaches and explains when reporting-period cost allocation becomes necessary, regardless of which model you choose.
Should agencies absorb OpenAI costs in their retainer?
Often yes, especially early on when AI spend is small relative to project fees. Absorbing costs simplifies client conversations and keeps invoices clean. The downside is margin risk: if usage grows faster than fees, the agency eats the variance silently until someone notices in a provider dashboard.
Retainer absorption works best when AI usage is predictable or immaterial. It still helps to track usage by client internally even if you do not bill it separately — so you know when the model is breaking.
Should agencies use a monthly AI allowance?
A fixed monthly AI allowance gives clients predictability with a defined cap. Overage can be absorbed, renegotiated, or passed through depending on the contract. Allowances reduce surprise bills but require monitoring — and clear rules for what happens when usage exceeds the cap.
Should AI API costs be passed through at cost?
At-cost pass-through maximizes transparency and aligns client charges with provider spend. Administrative overhead is higher because you must allocate provider rows to clients each reporting period and explain the breakdown. Clients may scrutinize line items more closely than with a bundled retainer.
Should agencies mark up OpenAI or Anthropic usage?
Markup compensates for management, risk, and integration work around AI usage. It is common in agency economics but requires clarity in contracts — clients should understand they are paying above provider rates. Markup does not remove the need to know underlying usage; it adds a layer on top of allocated cost.
When does usage-based client billing make sense?
Usage-based billing fits when AI spend is material, highly variable, and clients expect to pay for what they consume — similar to infrastructure pass-through. It places the highest burden on allocation accuracy and reporting-period reconciliation. Without a reliable client attribution process, usage-based billing creates disputes.
Comparing five billing models
| Model | Client predictability | Agency margin | Admin overhead | Transparency | Allocation need |
|---|---|---|---|---|---|
| Absorb into retainer | High for client | Agency bears variance | Low | Low | Optional |
| Fixed monthly allowance | Medium | Shared variance | Medium | Medium | Useful at overage |
| At-cost pass-through | Low for client | None on usage | Higher | High | Required |
| Marked-up pass-through | Low for client | On usage | Higher | Medium–high | Required |
| Usage-based billing | Lowest | Flexible | Highest | Highest | Required |
How should agencies explain variable AI usage costs to clients?
Frame costs around the reporting period: what changed, what providers were involved, and what is included or excluded. Avoid implying precision your data does not support — if exports show account-level usage rather than per-prompt detail, say so. Clients accept variable bills more easily when the statement is consistent month to month and ties to work they recognize.
What should an AI agency show a client?
A useful client-facing AI cost summary typically includes:
- Reporting period — the dates the statement covers.
- Provider / source — e.g. OpenAI, Anthropic, or other AI providers included.
- Client-attributed amount — spend assigned to this client for the period.
- Allocation notes — brief explanation where shared or rule-based rows were split.
- Unallocated or excluded items — if your process documents organization-level rows not billed to this client.
- Total client AI cost — the figure that supports your invoice line or pass-through discussion.
This is statement-level reporting — not prompt-level attribution unless your source data actually supports that granularity.
When reporting-period allocation becomes necessary
Allocation work intensifies with pass-through, markup, and usage-based models — and whenever AI spend becomes large enough that retainer margin is at risk. Even absorbed models benefit from internal allocation so you can reprice before margin disappears.
Provider dashboards help you gather cost data; they do not replace the client-allocation decision. See how to allocate shared AI API costs across clients for the operational workflow.
Illustrative example (fictional)
A fictional agency bills Client Nova a $500 monthly retainer including AI usage. Internal tracking shows Nova consumed $380 in OpenAI and Anthropic charges last month while another client on a similar retainer consumed $90. The agency keeps Nova's fee unchanged for now but recognizes the retainer model is underpriced for Nova — a signal to renegotiate or move Nova to an allowance with overage pass-through. These figures are illustrative only, not industry benchmarks.
Where AI Billback fits
Once you choose a billing model that ties provider spend to clients, you need a reporting-period workflow to import cost data, attribute rows, review unallocated spend, and produce client-ready statements. AI Billback is one tool for that workflow — it is not endorsed by OpenAI or Anthropic, and it does not choose your pricing model for you.
For agency-specific context, see AI cost billback for agencies.
Practical takeaway
Pick a billing model that matches client expectations and your capacity to administer it. As AI spend grows, internal visibility by client becomes non-optional — regardless of whether clients see a pass-through line or a bundled fee.
Apply your billing model with client attribution
Import provider cost data, attribute spend to clients, and open client-ready AI cost statements for your reporting period.
No proxy · No API keys · No application changes