Pay per token, nothing else.
Every rate below is per million tokens. Hosted inference does not require an hourly GPU rental or reserved capacity.
- Billed per 1M tokens
- Cached input rates
- Streaming supported
- Usage-based billing
Rates per 1M tokens
One rate card for every request: streaming or not, tools or not. Cached input is the rate for prompt prefixes served from cache.
| Model | Context | Input | Cached input | Output |
|---|---|---|---|---|
| GPT-OSS 120Bopenrelay/gpt-oss-120b | 128K | $0.15 | $0.015 | $0.60 |
| GPT-OSS 20Bopenrelay/gpt-oss-20b | 128K | $0.05 | $0.005 | $0.20 |
| Gemma 4 31B NVFP4 32Kopenrelay/gemma-4-31b-nvfp4-32k | 32K | $0.081 | n/a | $0.306 |
| Gemma 4 31Bopenrelay/gemma-4-31b | 32K | $0.14 | n/a | $0.40 |
| GLM 5.2openrelay/glm-5.2 | 1M | $1.82 | $0.338 | $5.72 |
| GLM 5.3 Flashopenrelay/glm-5.3-flash | 256K | $0.195 | $0.039 | $0.65 |
| DeepSeek V4-Flashopenrelay/deepseek-v4-flash | 256K | $0.22 | $0.22 | $0.66 |
| DeepSeek-OCR 2openrelay/deepseek-ocr-2 | 8K | $0.039 | n/a | $0.039 |
| Embeddings | Max input | Input | Notes |
|---|---|---|---|
| BGE-M3openrelay/bge-m3 | 8K | $0.013 | Input-only metering via /v1/embeddings |
All prices in USD per 1M tokens. You pay only for the tokens you use, metered per request. More models are on the way; browse the inference catalog or request a model.
The models behind the rates
GPT-OSS 120B
OpenAI
- Context
- 128K
- Parameters
- 120B
- Input / 1M
- $0.15
- Output / 1M
- $0.60
OpenAI's 120B open-weight model, with reasoning, tool calling, and structured output support.
GPT-OSS 20B
OpenAI
- Context
- 128K
- Parameters
- 20B
- Input / 1M
- $0.05
- Output / 1M
- $0.20
The smaller GPT-OSS variant, with tool calling and structured output support at lower token rates.
Gemma 4 31B NVFP4 32K
- Context
- 32K
- Parameters
- 31B
- Input / 1M
- $0.081
- Output / 1M
- $0.306
An NVFP4-quantized Gemma 4 31B deployment with a 32K-token context window and lower token rates.
Gemma 4 31B
- Context
- 32K
- Parameters
- 31B
- Input / 1M
- $0.14
- Output / 1M
- $0.40
Google's dense 31B model with text and image input, reasoning, tool calling, and a 32K-token context window.
GLM 5.2
Zhipu
- Context
- 1M
- Parameters
- 744B MoE
- Input / 1M
- $1.82
- Output / 1M
- $5.72
Zhipu's GLM model for reasoning, coding, and tool use in English and Chinese, with a 1M-token context window.
GLM 5.3 Flash
Zhipu
- Context
- 256K
- Parameters
- 320B MoE (18B active)
- Input / 1M
- $0.195
- Output / 1M
- $0.65
Zhipu's natively multimodal GLM 5.3 Flash, served on our own H100s. Hybrid linear and sparse attention keeps long-context serving cheap, and it takes text or images in.
DeepSeek V4-Flash
DeepSeek
- Context
- 256K
- Parameters
- 284B (13B active)
- Input / 1M
- $0.22
- Output / 1M
- $0.66
DeepSeek's V4-Flash mixture-of-experts with 13B active parameters, reasoning and tool calling, served on our own B200s at NVFP4.
DeepSeek-OCR 2
DeepSeek
- Context
- 8K
- Parameters
- -
- Input / 1M
- $0.039
- Output / 1M
- $0.039
DeepSeek's second-generation OCR model. Reads document images (scans, receipts, screenshots, tables) and returns structured markdown that preserves headings, tables, and layout.
Metered per request, itemized per token
Cached input
Prompt prefixes served from cache bill at the cached input rate: 10x cheaper than fresh input on GPT-OSS models. It applies automatically, with no configuration. How to structure prompts so they hit the cache.
Request features
Supported request features do not add a separate fee. Each request is metered on input and output tokens at the model's rates.
Batch at 50% off
Large offline jobs can run on the Batch Inference API at half the per-token rate, with results delivered within 24 hours. Rate-for-rate comparisons against Gemini's batch mode and Bedrock batch inference are in the provider guides.
Send your first request today.
Grab an API key, point your OpenAI SDK at OpenRelay, and pay only for what you run.