Pricing
You pay for tokens. The prices are published.
Every model in the catalog has a list price per million tokens for input, cached input, and output. The same numbers appear on the model page, in the rate sheet below, and on the usage your console reports — so a bill is never a surprise.
Estimate a monthly bill
Pick a model and a token mix. The estimate uses the same list prices as the rate sheet — no rounding tricks, no hidden platform fee in the number.
2,000,000 tokens
500,000 tokens
Estimated spend
$1.90
Input side
$0.95
Output side
$0.95
Estimate only: it uses today’s list prices for the selected model and assumes every token is billed at the input or output rate. Cached input is cheaper than the input rate, and failed requests that return an error status are not billed as completions.
How usage is metered
- One meter on Responses. Each response reports the input and output tokens it used, and that count is what the price applies to.
- Cached input is discounted. When a request reuses a prompt prefix, the cached tokens are billed at the cached-input rate shown for that model.
- Retries do not double-bill. Send an Idempotency-Key and a retried create replays the original response instead of charging again.
- Read your own numbers. The console shows requests, tokens, and settled spend for the workspace, broken down by model.
- These are our own prices. We sell each model at the rate on this page, and the meter bills the same numbers — one rate sheet, no per-model markup table to decode and no fee added at checkout.
Charged against your organization’s keys at https://api.models.sylphx.ai/v1. Keys are prefixed sk-sx- and can be revoked at any time.
The rate sheet
USD per million tokens, per model. Cached input is what you pay when a prefix is reused; the percentage under it is the saving against the input rate. 488.8K requests were metered in the last seven days across this catalog.
“Est. mix / M” is a 3:1 input:output estimate for a million tokens, shown to compare models. Your bill uses the exact input, cached-input, and output counts on each response.
Billing questions
The short answers. For the request-and-response details, start at the quickstart.
What exactly am I charged for?
Is there a subscription or a minimum?
Do failed requests cost anything?
How do prices change?
Where can I see what I have spent?
Lowest listed input price: Qwen: Qwen3.7 Flash at $0.0378 per million tokens.