Credit Pricing
Transparent, per-token billing. 1 Credit = 1 USD = 26,300 VND by default. All prices shown per 1M tokens.
Cost Calculation Formula
cost = (uncached_input / 1M) × input_price + (cached_input / 1M) × cache_price + (output / 1M) × output_price
Example: A 2,000-token conversation (1,500 input + 500 output) with Claude Sonnet 4 costs ~0.013 credits (~340 VNĐ)
Maintained Models
Free upstream — Izzi charges a small maintenance fee (~2,000 VNĐ / 1M tokens)
| Model | Provider | Size | Input / 1M | Output / 1M |
|---|---|---|---|---|
| Qwen3 235B | Cerebras | 235B | $0.08 | $0.14 |
| LLaMA 3.3 70B | Cerebras | 70B | $0.08 | $0.14 |
| LLaMA 3.1 405B | SambaNova | 405B | $0.08 | $0.14 |
| DeepSeek R1 | SambaNova | 671B | $0.08 | $0.14 |
| DeepSeek V3 | SambaNova | 671B | $0.06 | $0.12 |
| LLaMA 3.3 70B | SambaNova | 70B | $0.06 | $0.12 |
| QwQ 32B | SambaNova | 32B | $0.05 | $0.10 |
| Qwen3 32B | SambaNova | 32B | $0.05 | $0.10 |
| Nemotron 70B | NVIDIA | 70B | $0.06 | $0.12 |
| LLaMA 3.3 70B | NVIDIA | 70B | $0.06 | $0.12 |
| Mistral 7B | NVIDIA | 7B | $0.04 | $0.08 |
| LLaMA 3.1 8B | NVIDIA | 8B | $0.04 | $0.08 |
| Nemotron 3 Super 120B | OpenRouter | 120B | $0.08 | $0.14 |
| Devstral 2 123B | OpenRouter | 123B | $0.08 | $0.14 |
| Gemma 3 27B | OpenRouter | 27B | $0.08 | $0.14 |
| LLaMA 3.3 70B Free | OpenRouter | 70B | $0.08 | $0.14 |
| Auto Router | OpenRouter | — | $0.08 | $0.14 |
| Step 3.5 Flash | StepFun | — | $0.07 | $0.30 |
| GLM 4.5 Air | Z.ai | — | $0.07 | $0.30 |
Budget Models
Upstream cost +10% markup — best value for everyday tasks
| Model | Provider | Upstream Price | Izzi Input / 1M | Izzi Output / 1M |
|---|---|---|---|---|
| GPT-4o Mini | OpenAI | $0.15 / $0.60 | $0.165 | $0.660 |
| GPT-4.1 Mini | OpenAI | $0.40 / $1.60 | $0.440 | $1.760 |
| Gemini 2.5 Flash-Lite | $0.13 / $0.75 | $0.143 | $0.825 | |
| Gemini 2.5 Flash | $0.30 / $2.50 | $0.330 | $2.750 | |
| Grok 4.1 Fast | xAI | $0.21 / $0.53 | $0.231 | $0.583 |
Standard Models
Balanced performance and quality — includes approved flat-retail routes
| Model | Provider | Upstream Price | Izzi Input / 1M | Izzi Output / 1M |
|---|---|---|---|---|
| GPT-5.6 Luna | Codex LB | Flat retail | $0.045627 | $0.045627 |
| Claude Haiku 4.5 | Claude | $0.80 / $4.00 | $0.880 | $4.40 |
| GPT-4.1 | OpenAI | $2.00 / $8.00 | $2.20 | $8.80 |
| GPT-5.1 | OpenAI | $1.00 / $8.00 | $1.10 | $8.80 |
Premium Models
Top-tier models — flat-retail SmartRouter routes are marked in the table
| Model | Provider | Upstream Price | Izzi Input / 1M | Izzi Output / 1M |
|---|---|---|---|---|
| GPT-5.6 Sol | Codex LB | Flat retail | $0.064639 | $0.064639 |
| GPT-5.6 Terra | Codex LB | Flat retail | $0.045627 | $0.045627 |
| Grok 4.5 High | 9Router | Flat retail | $0.057034 | $0.057034 |
| Claude Sonnet 4.6 | Claude | $3.00 / $15.00 | $3.30 | $16.50 |
| Claude Opus 4.6 | Claude | $5.00 / $25.00 | $5.50 | $27.50 |
| Claude Sonnet 4.5 | Claude | $3.00 / $15.00 | $3.30 | $16.50 |
| Claude Sonnet 4 | Claude | $3.00 / $15.00 | $3.30 | $16.50 |
| Claude Opus 4 | Claude | $5.00 / $25.00 | $5.50 | $27.50 |
| GPT-5.4 | OpenAI | $2.50 / $15.00 | $2.75 | $16.50 |
| GPT-5.2 | OpenAI | $1.75 / $14.00 | $1.93 | $15.40 |
| Grok 4 | xAI | $3.00 / $15.00 | $3.30 | $16.50 |
Frequently Asked Questions
What is 1 Credit?
1 Credit = 1 USD = 26,300 VNĐ theo tỉ giá mặc định. Tài khoản mới nhận 5 credits miễn phí khi đăng ký.
Why are "free" models not free?
Izzi proxies free models through premium infrastructure (SambaNova, NVIDIA NIM, Cerebras, OpenRouter). The maintenance fee ($0.04–$0.08 / 1M tokens) covers server costs, monitoring, and reliability. This is 50–100× cheaper than paid models.
How is the +10% markup calculated?
Most paid models use upstream price × 1.10; rows marked Flat retail use the approved fixed SmartRouter rate. This supports API key management, rate limit optimization, failover routing, monitoring, and operational maintenance. Availability remains subject to upstream provider and infrastructure outages.
Do you offer discounts?
Yes! When Izzi finds cheaper upstream sources (e.g., via optimized routing or alternative providers), we pass the savings to users — up to 50% discount on select models. Check the pricing page for current promotions.
What about cached tokens?
Cache pricing is model-specific. GPT-5.6 and Grok 4.5 High use the same approved flat rate for input, output, and cached tokens. Other models may retain a discounted cache rate; see the current model row or API catalog.