Same workload.A smaller token bill.

Pick a model. Set your throughput. We'll profile your workload for a lower-cost serving configuration.

LOWER COSTSAME PERFORMANCEYOUR WORKLOAD

From workload to serving plan.

We profile your workload and match it to the right serving configuration.

Traffic

Representative requests

8K in650 out

Cached context

Route repeated prefixes together.

Matched compute

Prefill and decode phases.

Output SLA

Sustained generation target

TPSDefault

Example profiling path. The selected approach depends on model and workload behavior.

Model your inference costs.

Compare published token rates against your workload and cache profile.

Rates & assumptions

Choose a model

Across all requests, not one response.

Eligible input only. Example, not a forecast.

50%
0%100%

Assumptions

Assumes 4:1 input/output · 24h/day · 30 days.

Optional service target

Fleet output

Across all requests, not one response.

Fresh inputCached inputOutput

Public-rate estimate

$16,070

per month · $192,845 annualized

50% cache scenario

$2,074

lower per month in this scenario

Reference · 0% input cache hits$18,144
Scenario · 50% input cache hits$16,070
Keep the model. Lower the cost.

Final pricing follows your workload benchmark.

DeepInfra via OpenRouter · 2026-09-06 · Rates & assumptions

USD per 1M tokens: $0.09 fresh input · $0.05 cached input · $0.34 output. Route: deepinfra/turbo, FP4. Model pricing · Provider rate source

Costs use aggregate output TPS × active hours × 3,600 × 30 days, plus input at your selected ratio. Annualized values use twelve 30-day months. Both scenarios use the same provider’s rates and output volume. Bar segments represent fresh input (gray), cached input (green), and output (black).

OpenRouter publishes a cache-read rate for this provider endpoint. Its endpoint metadata reported implicit caching as unavailable at retrieval. Do not assume a cache hit. Apply this rate only to input confirmed eligible by the provider's cache behavior and request configuration. Cache hits are an editable assumption, not a prediction. No separate cache-write rate is published. Additional cache-write, storage, network, platform fees, and taxes are excluded, not assumed free.

Market-reference provider route, not an OpenRelay availability or price quote. These token-charge scenarios are not your actual bill, measured savings, or an SLA. Final pricing and service targets require a workload benchmark. sales@openrelay.inc

Start with the SLA that matters.

Not sure which to pick? Start with tokens per second. Let us profile your workload and lower the bill.

Recommended

TPS

Tokens per second

Set a sustained output rate for the workload you actually run.

Best starting point

TTFT + ITL

Latency

Set response targets for chat, copilots, and user-facing agents.

User experience

Tokens / window

Batch deadline

Set a token volume and the time by which it must finish.

Completion time

Where we look for savings.

Approaches benchmarked per workload.

Context reuse

Test cache-aware routing and shared-prefix reuse.

Kernel + batch tuning

Benchmark kernels, quantization, and batch shape.

Fit the hardware

Match prefill and decode phases to measured demand.

Let's find the expensive tokens.

Send your model, traffic shape, and target SLA. We will define a representative benchmark.

Email sales@openrelay.inc

Final pricing and SLA follow a workload benchmark.

Workload analysis
Serving configuration
Final pricing and SLA