What's the catch on 5%?
There isn't one. 5% per call is our entire platform fee — the only margin we make. We add it to the underlying model cost on every API call. No monthly fees, no seat fees, no per-model surcharges, no markup on tokens, no minimum. You only pay when a request succeeds. (Adding credit carries a separate 5% top-up fee that covers payment processing — disclosed before you pay, like any online card checkout.)
So what does it actually cost me?
Per call: provider cost × 1.05 — debited from your wallet when the request succeeds, and that 5% is the only fee we earn. Adding credit: a one-time 5% top-up fee covering card processing. Worked example: buy $100 of credit → charged $105 → wallet credited $100 → each call debits provider cost + a 5% platform fee from the wallet.
Is there free credit?
Adding a payment method (no charge) gets you $1 in credit spendable across a curated trial catalog — GLM-5.3 Flash leads the picks, and it also includes GLM-5.2, Kimi K2.7 Code, gpt-oss-120b, FLUX, Whisper Turbo, Aura 2, BGE-M3 and more. The $1 expires 30 days after you add the card. Top up any amount and the full 1000+ catalog opens at pass-through + a 5% platform fee. Topups never expire.
What's the launch offer?
Two tiers, split at $100. Top up $5-$99 through September 30, 2026 and your credit is tripled — top up $5, hold $15; top up $50, hold $150. That bonus is spendable on zai-org/glm-5.3-flash only, expires 90 days after the top-up, and is metered like any other credit; it is capped at $100 of bonus per account over the life of the offer, so a top-up above $50 stops earning more of it. A single top-up of $100 or more takes the other tier instead: no bonus, but zai-org/glm-5.3-flash stops being metered for you until September 30, 2026 — those calls are not billed and do not draw down your balance, so the $100 stays ordinary credit for the full 1050+-model catalog. The threshold is one payment, not a running total: five $20 top-ups do not unlock it, and refunding the payment ends the unlock. Everything lands automatically at checkout — no code — and applies to auto top-ups too.
What counts as a 'successful run'?
A request that returned a usable model response. Fallbacks that succeed bill once, at the model that actually answered. Failed runs and timeouts don't charge.
How do cache hits price?
Exact-match and semantic hits return in under 10ms and get a flat 50% discount on the uncached price. Cache is on by default; override per-request with the "cache" field in the request body.
Can I bring my own provider keys?
Yes, free on every paid account. Route through your Anthropic / OpenAI / Google / etc. key and pay zero per-request fees on that traffic — your existing volume discounts apply.
Sub-accounts and per-user keys?
Yes — one API call mints a scoped key with its own spend cap, rate limit, default tag, and analytics. Ideal for marketplaces or multi-tenant apps that want to meter end users.
How do STT, TTS, embeddings, video, image price?
Every modality uses the same math: underlying model rate plus 5%. Whisper, ElevenLabs, BGE, Flux, Veo — all one simple line item.
Do credits expire?
Topup credits never expire — once you pay, the balance stays. Two exceptions: the $1 card credit expires 30 days after you add a card, and launch-offer bonus credit expires 90 days after the top-up that earned it. Refunds on unused topup balances within 90 days, no questions.
What's in Enterprise?
Evals, guardrails, replay + shadow A/B, and versioned prompt IDs — the gateway primitives. Plus SSO, 99.95% SLA, DPA, SOC 2, private audit export, dedicated endpoint, and a named engineer on Slack Connect. Details on /enterprise.