Pricing
Pay only for tokens you use
No monthly minimums. No tiered cliffs. We charge upstream cost + a flat 30% margin — and show every line of the bill.
Start free. Scale by usage. No surprises.
Pay-as-you-go is the cheapest plan for most teams. Upgrade to enterprise when you need volume discounts or dedicated capacity.
Free Trial
- $10 starter credits (≈ 800K GPT-4o-mini tokens)
- All 50+ models included
- Standard network node
- Email support
Pay-as-you-go
- All 50+ models, all providers
- Priority low-latency node
- Real-time consumption stats
- Stripe, PayPal, crypto top-ups
- Per-key rate limits & alerts
- Full audit logs
- Auto-refill supported
Enterprise
- Volume discounts from 5M+ tokens / day
- Exclusive dedicated node
- Unlimited concurrency
- Custom DPA, BAA, SOC 2 reports
- SLA-backed 99.99% uptime
- Custom rate limits & quotas
- Per-tenant isolated billing
- Dedicated solutions engineer
Estimate your monthly bill in 30 seconds
Three real-world workloads below. Multiply by your own usage for a precise estimate.
Light chatbot
GPT-4o-mini≈ $54 / month
Production SaaS
GPT-4o-mini mix≈ $2,700 / month
Reasoning workload
DeepSeek-Reasoner≈ $329 / month
Cost = tokens × (input rate · % input + output rate · % output). Set hard caps in your dashboard to never exceed budget.
All models, all rates, all in one place
Prices per 1M tokens in USD · charged per request with millisecond-level precision.
OpenAI
Industry-leading general intelligence & reasoning.| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
gpt-4oHot | 128K | $2.500 | $10.000 |
gpt-4o-miniHot | 128K | $0.150 | $0.600 |
o1 | 200K | $15.000 | $60.000 |
o1-mini | 128K | $3.000 | $12.000 |
gpt-4-turbo | 128K | $10.000 | $30.000 |
Anthropic
Long context, careful reasoning, strong code.| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
claude-sonnet-4-5-20250929Hot | 200K | $3.000 | $15.000 |
claude-haiku-4-5-20251001Hot | 200K | $1.000 | $5.000 |
claude-opus-4-1-20250805 | 200K | $15.000 | $75.000 |
claude-3-5-sonnet-20241022 | 200K | $3.000 | $15.000 |
claude-3-5-haiku-20241022 | 200K | $0.800 | $4.000 |
| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
gemini-2.0-flashHot | 1M | $0.100 | $0.400 |
gemini-2.0-flash-lite | 1M | $0.025 | $0.100 |
gemini-1.5-pro | 2M | $1.250 | $5.000 |
DeepSeek
Math & reasoning at aggressive prices.| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
deepseek-chatHot | 32K | $0.270 | $1.100 |
deepseek-reasoner | 64K | $0.550 | $2.190 |
Alibaba (Qwen)
Strong Chinese & English instruction following.| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
Qwen/Qwen2.5-72B-InstructHot | 131K | $0.410 | $1.210 |
Qwen/Qwen2.5-7B-Instruct | 131K | $0.100 | $0.100 |
Qwen/Qwen2-VL-72B-Instruct | 32K | $0.990 | $0.990 |
Meta (Llama)
Open-weight flagship models.| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
meta-llama/llama-3.3-70b-instruct | 128K | $0.590 | $0.790 |
meta-llama/llama-3.1-8b-instruct | 128K | $0.050 | $0.080 |
BAAI (Embeddings)
Best-in-class retrieval embeddings.| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
BAAI/bge-m3 | 8K | $0.100 | — |
BAAI/bge-large-zh-v1.5 | 8K | $0.100 | — |
BAAI/bge-large-en-v1.5 | 8K | $0.100 | — |
Need a model not listed? Request one — we add new flagship models within a week.
Everything you'd ask before adding us to your stack
Pre-funded wallet: top up via Stripe, PayPal, or crypto (USDT/BTC/ETH). Each API call debits your wallet by the exact token cost at published rates, with millisecond precision. No monthly bills, no surprise charges.
Yes — $10 in starter credits at signup (≈ 800K GPT-4o-mini tokens). After depletion, add a payment method to keep using the same API. No credit card required for the trial.
Pay-as-you-go already has the lowest unit price. For sustained usage above 5M tokens / day, contact us for custom enterprise rates — typically 20–40% lower than the published rates.
Yes. Set a per-key rate limit (rpm / tpm) and a per-month hard cap in your dashboard. We'll alert you at 80% and stop accepting requests at 100% — no surprise bills.
If we return 4xx (your error — bad request, insufficient balance, etc.), you are not charged. If we returned a 200/2xx but the response is malformed, we refund the hold automatically.
We add a flat 30% margin on top of upstream cost. The trade-off: one API key, one bill, automatic failover across providers, and 60-second integration. For most teams this trades 30% cost for months of integration engineering.
Start with $10 in free credits
No credit card. Production traffic in 60 seconds.