- Selecting common models from Anthropic, OpenAI, Google, and other providers
- Creating billable metrics, products, and rates based on your configured markup percentage(s)
- Automatically syncing newly released models and provider token price changes to your rate card at your configured markup
Public previewBilling for LLM tokens is in public preview. We will continue to build on the product and there may be quick changes!
Set up with an AI coding agent
Use the Metronome billing for LLM tokens skill to guide an AI coding agent through a sandbox integration. Open your application repository, then copy and paste this prompt into your agent:Use case
Fictional company Designr is an AI-powered design tool. Customers use Designr to generate design assets, including prototypes, mockups, and images. Designr uses common AI models and charges customers a 10% markup on underlying model costs. Designr offers several plan tiers. Its Pro Plan includes 200 Designr Credits — a custom pricing unit — each month. If a customer uses all of their Designr Credits, they can purchase additional credits during the month.Set up your rate card
In Metronome, a rate card is your centralized price book, where you define pricing for all products. When you use billing for LLM tokens, Metronome automatically creates billable metrics, products, and rates for managed AI products based on the markup you enter. You do not need to create these separately. Because Designr uses a custom pricing unit, Designr Credits, first navigate to Offering > Pricing Units > Custom Pricing Units. Click + Add, then create a custom pricing unit named Designr Credits. Next, create your rate card.Note: Billing for LLM tokens uses a USD rate card because provider prices are denominated in USD. To let customers pay in their local currency, see Let customers pay in their local currency.
Create your rate card
- Click Offering in the left-hand sidebar.
- Navigate to the Rate Cards tab and select + Add.
- Enter the rate card name and description, then enable Charge based on AI provider pricing (managed). You can also add a human-readable alias, such as
default_rate_card, to reference the rate card more easily throughout the API. - Select the AI models you want to use. You can select all models from a provider or expand the provider to select individual models.
- Add any other usage-based, subscription, or composite products that you want to include on the same rate card.
- Click Next to proceed to the next page.
Set rates — Custom Pricing Unit
- Under Default markup for future AI models, enter the markup percentage that should automatically apply when new models are added to the rate card.
- In the upper-right corner of the AI models section, select USD. In the dropdown, select the Designr Credits custom pricing unit.
- In the modal, define a conversion rate between USD and Designr Credits.
- Expand each model to verify its distinct rates by author, provider, and token type.
- Click Save.
Set rates — USD
- Under Default markup for future AI models, enter the markup percentage that should automatically apply when new models are added to the rate card.
- Enter markup percentages for each selected model, or use Apply markup to all in the upper-right corner to apply the same markup percentage to every selected model.
- Expand each model to verify its distinct rates by author, provider, and token type.
- Click Save.
Let customers pay in their local currency
Managed AI rate cards use USD, but customers can pay in their local currency, such as EUR, with Stripe Adaptive Pricing. You keep your managed USD rate card, so Metronome continues to sync newly released models and provider price changes. Invoices remain in USD, and Stripe shows customers a converted amount when they pay. To use Adaptive Pricing with Metronome invoices:- Set the customer’s
stripe_collection_methodtosend_invoice. Adaptive Pricing for invoices applies when customers pay on the Stripe hosted invoice page. - Make sure USD is one of your Stripe settlement currencies, because Stripe requires invoices to be issued in a settlement currency. To add USD, see multi-currency settlement.
Define your pricing model
Because Designr’s Pro Plan includes an allocation of 200 Designr Credits per month, you can create a Package to encode the credit allocation alongside the rate card you just created. In Metronome, Packages define customer-facing offerings, such as Pro Plan or Max Plan, and simplify assigning PLG customers to these offerings.Provision customers
You are now ready to assign customers to the Pro Plan. Provisioning a customer with a package creates a contract: a customer-specific agreement that applies the terms from the package. Use the API call below to provision a customer with the Pro Plan:Integrate usage tracking
Ensure that your events follow the format below, withevent_type set to token-billing.
input_tokens: Uncached input tokenscached_input_tokens: Cache-read input tokensoutput_tokens: Tokens in the responsecached_write_tokens: Cache-write tokens. Supported for Anthropic models and OpenAI GPT-5.6+ models only
Map provider token counts
Providers report token usage in different shapes. Some include cached tokens in their total input count, while others report them separately. Before you send an event, convert the provider’s counts so that each token appears in exactly one property. For example, a cache-read token belongs incached_input_tokens only, not in input_tokens as well.
When a provider leaves out an optional count, such as cache writes for a request that didn’t write to the cache, use 0 in your calculation. In the event, you can send that property as 0 or leave it out. Either way, Metronome bills no tokens of that type.
For streaming responses, send the event after the stream finishes and the provider reports final usage. Don’t estimate token counts from partial output.
OpenAI Responses
The Responses APIusage object includes cache reads and cache writes in input_tokens. Subtract them to get ordinary input:
properties:
OpenAI Chat Completions
The Chat Completions APIusage object includes cache reads and cache writes in prompt_tokens. Subtract them to get ordinary input:
stream_options.include_usage to true so that the final chunk includes usage.
For example, given this API response:
properties:
Anthropic Messages
The Messages APIusage object reports ordinary input, cache reads, and cache writes as separate counts. Map them directly:
properties:
Google Gemini generateContent
For text requests without tool use, the UsageMetadata object includes cache reads in promptTokenCount. Subtract them to get ordinary input. Thinking tokens count toward output pricing, so add them to output. Gemini doesn’t report cache writes, so leave out cached_write_tokens:
properties:
model and provider fields match each token count to the correct rate.
Send events in the correct format to Metronome’s /ingest endpoint. Then navigate to the Events page to confirm that the events have matched a billable metric.
Test your end to end flow
To test your integration:- Create a webhook destination by navigating to Developer > Notifications > Webhooks.
- Create a test customer.
- Provision the test customer using the rate card that you just created.
- Send in usage events with the test customer’s
customer_id. - Call the
getInvoiceendpoint to view the customer’s spend, and validate that each line item looks correct based on the markup you configured.