How billing works
Text calls are billed from the provider's measured input, cached input, output, and reasoning tokens—not from the length of the prompt alone.
- The API reserves credits before a call, then settles against actual provider usage.
- Long-context pricing applies only where the selected model declares a separate long-context rate.
- A failed model call releases the full reservation.
Price in practice
Rates are stated per one million tokens, but settlement is rounded for each measured input and output component. This means a small request can still consume credits even when the displayed per-million rate looks fractional.
Cached input and reasoning are distinct usage fields. Do not estimate a final bill by counting prompt characters; use the API usage returned after the request.
Rate card
Questions
How is a token call charged?
Input, cached input, output, and reasoning tokens are measured by the provider. A successful request is charged from actual usage; failed model calls are refunded.
Why can the final cost differ from the estimate?
The estimate reserves enough credit to start safely. The final amount uses the provider's measured usage and returns any unused reservation.
Are credits charged before or after a text response?
Credits are reserved before the provider call and settled from measured usage afterwards. Any unused reservation is returned.
Do cached tokens cost the same as fresh input?
No. Cached input follows the provider rate for cached tokens, which can differ from fresh input.