← Pricing
PricingModels

Text & reasoning

Text, JSON, and multimodal-to-text calls, billed per million tokens.

How billing works

Text calls are billed from the provider's measured input, cached input, output, and reasoning tokens—not from the length of the prompt alone.

  • The API reserves credits before a call, then settles against actual provider usage.
  • Long-context pricing applies only where the selected model declares a separate long-context rate.
  • A failed model call releases the full reservation.

Price in practice

Rates are stated per one million tokens, but settlement is rounded for each measured input and output component. This means a small request can still consume credits even when the displayed per-million rate looks fractional.

Cached input and reasoning are distinct usage fields. Do not estimate a final bill by counting prompt characters; use the API usage returned after the request.

Rate card

ServicePriceHow it is applied
Gemini 3.1 Pro Preview500 in · 3,000 out1,000 in · 4,500 out above the long-context threshold
Gemini 3.8 / 3.7 / 3.6 Flash375 in · 1,875 outPer 1M tokens
Gemini 3.5 Flash375 in · 2,250 outPer 1M tokens
Gemini 3.5 Flash Lite75 in · 625 outPer 1M tokens
Gemini 2.5 Pro313 in · 2,500 out625 in · 3,750 out for long context
Gemini 2.5 Flash75 in · 625 out250 Cr / 1M audio input
Gemini 2.5 Flash Lite25 in · 100 out75 Cr / 1M audio input
Kimi K3750 in · 3,750 outCached input uses its provider rate
Kimi K2.6 / K2.7 Code238 in · 1,000 outCached input uses its provider rate
NVIDIA models25–140 in · 70–420 outRate varies by selected model

Questions

How is a token call charged?

Input, cached input, output, and reasoning tokens are measured by the provider. A successful request is charged from actual usage; failed model calls are refunded.

Why can the final cost differ from the estimate?

The estimate reserves enough credit to start safely. The final amount uses the provider's measured usage and returns any unused reservation.

Are credits charged before or after a text response?

Credits are reserved before the provider call and settled from measured usage afterwards. Any unused reservation is returned.

Do cached tokens cost the same as fresh input?

No. Cached input follows the provider rate for cached tokens, which can differ from fresh input.