Models ready to run
Choose a model, call the API, and pay only for the tokens you use.

GPT-OSS 120B
$0.15 / $0.60
input / output per 1M
OpenAI's 120B open-weight model, with reasoning, tool calling, and structured output support.
Text · Reasoning · Code
Model page
GPT-OSS 20B
$0.05 / $0.20
input / output per 1M
The smaller GPT-OSS variant, with tool calling and structured output support at lower token rates.
Text · Code
Model pageGemma 4 31B NVFP4 32K
$0.081 / $0.306
input / output per 1M
An NVFP4-quantized Gemma 4 31B deployment with a 32K-token context window and lower token rates.
Text · Vision · Reasoning
Model pageGemma 4 31B
$0.14 / $0.40
input / output per 1M
Google's dense 31B model with text and image input, reasoning, tool calling, and a 32K-token context window.
Text · Vision · Reasoning
Model pageGLM 5.2
$1.82 / $5.72
input / output per 1M
Zhipu's GLM model for reasoning, coding, and tool use in English and Chinese, with a 1M-token context window.
Text · Reasoning · Code
Model pageGLM 5.3 Flash
$0.195 / $0.65
input / output per 1M
Zhipu's natively multimodal GLM 5.3 Flash, served on our own H100s. Hybrid linear and sparse attention keeps long-context serving cheap, and it takes text or images in.
Text · Vision · Reasoning · Code

DeepSeek V4-Flash
$0.22 / $0.66
input / output per 1M
DeepSeek's V4-Flash mixture-of-experts with 13B active parameters, reasoning and tool calling, served on our own B200s at NVFP4.
Text · Reasoning · Code

DeepSeek-OCR 2
$0.039 / $0.039
input / output per 1M
DeepSeek's second-generation OCR model. Reads document images (scans, receipts, screenshots, tables) and returns structured markdown that preserves headings, tables, and layout.
Vision
Model pageDon't see the model you need?
Need a guaranteed SLA for a model?
The pay-per-token catalog uses shared capacity. Contact us for dedicated capacity with contractual availability and performance targets.
- Availability
- Uptime terms and service credits are agreed in the contract for dedicated capacity.
- Performance
- Time-to-first-token and throughput targets are sized to the selected model and traffic profile.
- Support
- Support channels and response times are defined as part of the deployment agreement.
Start calling models in minutes
Grab an API key, pick a model, and send your first request.