GPU Router
Loading models…

Fetching the model catalog.

DEVELOPER DOCUMENTATION

Pricing & free models

Know the benchmark. Know the price you can actually call.

Four prices, one token basis

  • Direct list: the model’s direct price reported by the market source.
  • OpenRouter: the lowest combined input + output price from one active Standard provider endpoint, identified alongside the price. Flex, batch and priority rates are excluded. Provider discounts are already included. If endpoint prices are unavailable, the comparison explicitly labels its catalog fallback. These are base token rates; long-context and request-specific charges can differ.
  • Market price: the cheapest observed qualified offer for the selected token basis.
  • Our price (estimated): the selected gateway’s current eligible route. When labeled “budget estimate,” it includes a routing price buffer and is a ceiling for budgeting; the final billed cost can be lower. The workload savings calculation uses this estimate when available. Market observations can be cheaper and do not guarantee execution at that price.

Discount calculation

discount = (1 − price / direct price) × 100

The default comparison is 1M input plus 1M output tokens. Use the workload calculator for your own token mix. Charts label current reference prices explicitly; they are not fabricated historical observations.

The 99.5% free threshold

Both input and output rates must be no more than 0.5% of the corresponding nonzero direct prices. Qualification uses exact micro-dollar amounts before display rounding. Unapproved, stale, missing-reference, and empty-capacity offers cannot qualify.

“Free candidate” means the market price meets the threshold. A request is free only when the gateway exposes an active sponsored route with zero customer charge. A candidate never silently becomes a paid request.

Admin control

Administrators enable free-model visibility and access in Settings → Admin. It starts disabled. Turning it off hides free listings and promotions and blocks new sponsored calls; requests already admitted may finish. Provider eligibility and spending limits always apply.

Calling a free model

Choose a model ID ending in :free from the gateway’s live catalog. Both customer token prices are $0. An API key is still required; an active stake or credit balance is not. The sponsor pays the provider using a separate daily allowance.

Free calls allow up to 2,048 output tokens, 16,384 estimated input tokens, and 100 attempts per wallet per UTC day, shared across its keys. A depleted sponsor allowance returns 429. A vanished qualifying route returns 404. Neither error switches to paid inference.

The sponsor reserves the executable route ceiling, which can exceed the observed qualifying floor because the venue uses whole-percent discount paths. That difference is funded by the sponsor. A trusted-provider allow-list stays attached to dispatch.

Settlement

Paid calls reserve credits before dispatch. Final provider receipts take precedence over token-price estimates. Missing usage or interrupted streams may require conservative settlement. Check the request receipt for its cost basis.

Documentation | GPU Router