Every model that matters, one POST away, priced to the token.
Inference is the model funnel on WAVE, media infrastructure for the agentic internet. One OpenAI-compatible endpoint fronts every route below. You ask for a task; the funnel picks a route that clears the measured bar and charges you that route's rate.
Fifteen routes, one endpoint, a 12.5x spread between the cheapest paid route and the priciest. Routing is the product: you pay the floor that still clears the bar, and one route runs on our own hardware at zero marginal cost.
$ curl -sH "Authorization: Bearer $KEY" https://inference.wave.online/v1/models | jq '.data | length' 15
One call shape, whichever route answers
The request is the OpenAI chat-completions shape and so is the response, including the error bodies. Point an existing client at this base URL, change the key, and nothing else moves. There is no SDK to adopt and nothing to host.
The routing decision is measured, not guessed
Most gateways route by margin or by a table someone typed. This funnel routes on a measured floor-to-ceiling profile per task class, so "cheapest sufficient" is a number you can read back. A new route is admitted by measurement or not at all, and the price it carries is the price you see above.
Fallback degrades, it does not stop
Cooldowns, retries, and fallback chains belong to the funnel. When a route drops, the next one on the chain serves the same request under the same call shape. Your request either completes or you get an OpenAI-compatible error that says which stage failed.
$ curl -s https://inference.wave.online/health/readiness
{"status":"healthy","db":"connected"}
Metered to the token
Every completion is logged against your key: tokens in, tokens out, spend to eight decimals. Virtual keys carry their own budgets and expirations, so one tenant cannot surprise another, and reading your spend is a query rather than a support ticket. Auth, entitlement, and metering settle through api.wave.online; this spoke renders the front door and forwards the rest.