Media

Money

Dispatch

Company

WAVE · Inference

Every model that matters, one POST away, priced to the token.

Inference is the model funnel on WAVE, media infrastructure for the agentic internet. One OpenAI-compatible endpoint fronts every route below. You ask for a task; the funnel picks a route that clears the measured bar and charges you that route's rate.

GET /v1/model/infolive prices, 2026-09-03
bandUSD per million input tokensinout own hardware$0.0000$0.0000 under $1$0.4000$1.2000 $0.5900$0.7900 $0.6000$0.6000 $0.6267$1.8800 $0.6667$2.0000 $0.8333$2.5000 $1 to $2$1.1867$3.5600 $1.2500$10.0000 $2 and up$2.0000$6.0000 $2.0000$6.0000 $3.3333$10.0000 $3.3333$10.0000 $3.3333$10.0000 $5.0000$15.0000

Fifteen routes, one endpoint, a 12.5x spread between the cheapest paid route and the priciest. Routing is the product: you pay the floor that still clears the bar, and one route runs on our own hardware at zero marginal cost.

$ curl -sH "Authorization: Bearer $KEY" https://inference.wave.online/v1/models | jq '.data | length' 15

15routes behind one OpenAI-compatible endpoint
1,048,576tokens, the largest live context window on the funnel
0.59 sround trip on a sample completion, 12 tokens billed

One call shape, whichever route answers

The request is the OpenAI chat-completions shape and so is the response, including the error bodies. Point an existing client at this base URL, change the key, and nothing else moves. There is no SDK to adopt and nothing to host.

callPOST /v1/chat/completions {"model","messages"}
catalogueGET /v1/models every route the funnel will answer on
pricesGET /v1/model/info input and output cost per token, per route
authAuthorization: Bearer <key> virtual keys, issued by the gateway
OpenAI-compatible15 routesautomatic failoverper-token metering

The routing decision is measured, not guessed

Most gateways route by margin or by a table someone typed. This funnel routes on a measured floor-to-ceiling profile per task class, so "cheapest sufficient" is a number you can read back. A new route is admitted by measurement or not at all, and the price it carries is the price you see above.

Fallback degrades, it does not stop

Cooldowns, retries, and fallback chains belong to the funnel. When a route drops, the next one on the chain serves the same request under the same call shape. Your request either completes or you get an OpenAI-compatible error that says which stage failed.

$ curl -s https://inference.wave.online/health/readiness
{"status":"healthy","db":"connected"}

Metered to the token

Every completion is logged against your key: tokens in, tokens out, spend to eight decimals. Virtual keys carry their own budgets and expirations, so one tenant cannot surprise another, and reading your spend is a query rather than a support ticket. Auth, entitlement, and metering settle through api.wave.online; this spoke renders the front door and forwards the rest.

Get an API key Read the docs