The WAVE Dispatch Engine
One engine decides how work gets done: which tier runs this, at what cost, on whose infrastructure. We return the decision in milliseconds; your stack runs the work. The prompt never has to leave your network for us to route it.
14 capabilities, 5 priced
Every one is the same call for a person with a key and an agent with a wallet.
Included with a key9 paths answer without a payment step
Each chip opens the machine-readable document for that path.
The terms are public before the first call
$ curl -s -D - -o /dev/null https://dispatch.wave.online/.well-known/x402 HTTP/2 200content-type: application/json
The discovery document answers unauthenticated. An agent reads the terms there before it spends anything. Open the discovery document
| Field | Observed |
|---|---|
| Status | 200OK |
| content-type | application/json |
| Surface | dispatch.wave.online/.well-known/x402 |
| Document | Open the discovery document |
Solved once, for every kind of work
Classification, cost-aware routing, the confidence gate, and decision metering belong to the engine. Add a model, a pool, or a provider, and it inherits all of it: no second identity system, no second meter, no second compliance posture.
9 the hard parts6 live · 2 building · 1 planned
| Name | Status | What it solves |
|---|---|---|
| 01Edge classify | Live | A prompt is classified in embedding space at the edge and answered with a route, not a completion: which tier should run this work. Milliseconds, and no inference performed to decide. |
| 02The decision is the unit | Live | One routing decision is one metered unit, $0.001 over x402, or a card subscription. WAVE bills the orchestration layer and never the inference, so local hops we never executed are never token-billed. |
| 03Agent-payable | Live | Every WAVE capability is callable with a key and payable by an autonomous agent over HTTP-402 on the same route, and a routing decision is no exception. High-frequency, low-stakes decisions are exactly where per-call settlement has to work without friction, and the live terms are public at dispatch.wave.online. |
| 04Control plane / data plane | Live | The wall is the product. Keys, prompts, and inference stay on your infrastructure; the edge returns a decision and receives nothing it did not need. BYO-infra and BYO-keys are architectural, not a setting you can forget to switch on. |
| 05Sovereign, fail-closed | Live | Run it fully local and it fails closed: if local inference fails, you get an error, not a silent forward of your prompt to a frontier provider. The quiet leak is the failure mode worth engineering against. |
| 06Reach it from anywhere | Live | Thin clients in five languages, a CLI, an MCP tool surface, and an OpenAI-compatible proxy: point an agent you already run at a new base URL and it routes. |
| 07Named profiles | Building | Fast, Expert, Heavy, and Code as explicit local → heavy → frontier chains, so a route resolves to a named policy instead of a single hop. Defined and scaffolded; not yet served by the live decision path. |
| 08Published economics | Building | The proxy logs every request so you can measure your own local-able fraction before enabling offload. Measuring your own traffic ships today. A published WAVE benchmark with methodology does not exist yet. |
| 09One plane for the AI surface | Planned | Voice, captions, transcription, sentiment, and search inherit this plane rather than inventing a parallel one, the same routing policy, the same meter, the same agent path. |
Where the work actually runs
The wall is the product. We return a decision; your infrastructure does the work.
The control plane
- Classifies the turn at the edge
- Returns a route and a confidence margin
- Checks the licence or the x402 payment
- Meters exactly one decision
The data plane
- Runs the model — local pool, heavy box, or frontier
- Holds every API key
- Sees the prompt and the output
- Never hands us either one
Why this is an engine and not an AI feature
Once media can move and value can settle, something still has to decide what computation each request actually needs. Most systems answer by default, send everything to a frontier model, pay frontier prices, and let every prompt leave the building. That default is a policy choice, not a technical necessity, and the Dispatch Engine exists to make it optional. It does not replace models. It decides which one earns the turn.