Traffic
Traffic
Representative requests
Pick a model. Set your throughput. We'll profile your workload for a lower-cost serving configuration.
We profile your workload and match it to the right serving configuration.
Representative requests
Route repeated prefixes together.
Prefill and decode phases.
Sustained generation target
Example profiling path. The selected approach depends on model and workload behavior.
Compare published token rates against your workload and cache profile.
Rates & assumptionsChoose a model
Across all requests, not one response.
Eligible input only. Example, not a forecast.
Assumptions
Assumes 4:1 input/output · 24h/day · 30 days.
Fleet output
Across all requests, not one response.
Public-rate estimate
$16,070
per month · $192,845 annualized
50% cache scenario
$2,074
lower per month in this scenario
Final pricing follows your workload benchmark.
USD per 1M tokens: $0.09 fresh input · $0.05 cached input · $0.34 output. Route: deepinfra/turbo, FP4. Model pricing · Provider rate source
Costs use aggregate output TPS × active hours × 3,600 × 30 days, plus input at your selected ratio. Annualized values use twelve 30-day months. Both scenarios use the same provider’s rates and output volume. Bar segments represent fresh input (gray), cached input (green), and output (black).
OpenRouter publishes a cache-read rate for this provider endpoint. Its endpoint metadata reported implicit caching as unavailable at retrieval. Do not assume a cache hit. Apply this rate only to input confirmed eligible by the provider's cache behavior and request configuration. Cache hits are an editable assumption, not a prediction. No separate cache-write rate is published. Additional cache-write, storage, network, platform fees, and taxes are excluded, not assumed free.
Market-reference provider route, not an OpenRelay availability or price quote. These token-charge scenarios are not your actual bill, measured savings, or an SLA. Final pricing and service targets require a workload benchmark. sales@openrelay.inc
Not sure which to pick? Start with tokens per second. Let us profile your workload and lower the bill.
TPS
Set a sustained output rate for the workload you actually run.
TTFT + ITL
Set response targets for chat, copilots, and user-facing agents.
Tokens / window
Set a token volume and the time by which it must finish.
Approaches benchmarked per workload.
Test cache-aware routing and shared-prefix reuse.
Benchmark kernels, quantization, and batch shape.
Match prefill and decode phases to measured demand.
Send your model, traffic shape, and target SLA. We will define a representative benchmark.
Email sales@openrelay.incFinal pricing and SLA follow a workload benchmark.