OpenGatewayDocs

Docs. One OpenAI- and Anthropic-compatible endpoint across providers.

If your client speaks the OpenAI API, it speaks OpenGateway: swap the base URL, set your key, and pick a model. The reference below covers keys, streaming, tools, routing, errors, limits and cost tracking.

Quickstart. One request, three clients.

Base URL https://opengateway.gitlawb.com/v1, and your key as Authorization: Bearer ogw_live_…. Create keys in the console. Model ids are exact: the model catalog has copy-paste ids, live pricing and capabilities.

curl https://opengateway.gitlawb.com/v1/chat/completions \
  -H "authorization: Bearer ogw_live_…" \
  -H "content-type: application/json" \
  -d '{
    "model": "mimo-v2.5-pro",
    "messages": [
      {"role": "user", "content": "hello, gateway"}
    ]
  }'

Reference. Everything the endpoint does, topic by topic.

Authentication

Every /v1/* request needs a key in the Authorization header. Keys look like ogw_live_ plus 32 hex characters; only a hash is stored server-side, and you can hold up to 20 active keys, one per app or environment, each revocable on its own. Requests without a valid key get 401 api_key_required.

Anthropic-compatible endpoint

The gateway also speaks the Anthropic Messages API: POST /v1/messages on https://opengateway.gitlawb.com. Point any Anthropic SDK (or Claude Code's ANTHROPIC_BASE_URL) at the gateway with your ogw_live_ key; the SDK's native x-api-key header is accepted. Any claude-* model id routes through smart routing (billed at the serving model's rate; the x-gateway-served-model header names what answered), or pass one of the catalog's exact ids to pin a model. max_tokens is required, per the Anthropic spec; streaming, tool use and system prompts are translated. count_tokens is not supported.

Models and ids

Send the model string exactly as listed in the catalog: full provider/model ids (MiMo also accepts the bare mimo-v2.5-pro short form). Ids ending in :free bill $0 and work with a zero credit balance, rate limited upstream. Retired free ids alias to their paid siblings, so old configs keep working, billed at the paid rate.

Streaming

Set stream: true for standard OpenAI server-sent events. The gateway opens the stream immediately and, while a slow reasoning model thinks, emits keepalive chunks: ordinary chat.completion.chunk events with an empty delta and ids like chatcmpl-gateway-keepalive-…. OpenAI SDKs handle them; hand-rolled parsers should tolerate zero-delta chunks (don't treat the first chunk as the first token).

Upstream failures mid-stream arrive as a data: {"error": …} event followed by [DONE]. Reasoning models want max_tokens of 300 or more; smaller budgets can be consumed entirely by thinking.

Tool calling

Standard OpenAI tools and function calling pass straight through on supporting models. The catalog says which (so does supports_tools on GET /v1/models). Smart routing never sends a request with tools to a model whose tool support is unverified.

Smart routing

Send model: "auto" (or gitlawb/auto) and the gateway scores your request, starts on the cheapest capable model, and escalates on 429, 5xx and timeouts; you are billed at the serving model's rate. Steer it with a route object in the body: priority is cost, balanced or quality, and max_cost_usd sets a price ceiling, as below. The x-gateway-served-model response header names the model that answered, and your decision log is at console, routing.

{
  "model": "auto",
  "route": { "priority": "cost", "max_cost_usd": 0.01 },
  "messages": [ … ]
}

Errors

Error statuses and codes
StatusCodeMeaning
401api_key_requiredThe Authorization header is missing.
401api_key_invalidThe key is malformed, unknown or revoked.
402insufficient_creditsThe balance is at or below zero on a paid model. Top up, or use a :free model.
402free_quota_exceededThe free window is used up; the body includes the reset time.
429gateway_rate_limit_exceededA per-IP or per-key limit was hit. Honour retry-after.
5xxupstream_unavailableA provider-side failure. Safe to retry with backoff.

Error bodies are OpenAI-shaped ({"error": {…}}), so SDK error handling works unchanged.

Free usage

Two ways to use the gateway with a zero credit balance: :free models (currently Nemotron 3 Ultra, $0, rate limited upstream), and a free allowance on xiaomi/mimo-v2.5-pro of 30 requests per 5-hour window (windows reset at 00, 05, 10, 15 and 20 UTC), subject to a shared daily budget. When a window is used up you get 402 free_quota_exceeded with the reset time; topping up credits removes the cap.

Pay per request with USDC (x402)

Agents can pay per request in USDC on Base, with no sign-up and no API key. Send a completion request with no credential header; the gateway replies 402 with a PAYMENT-REQUIRED header (the standard x402 challenge: exact scheme, Base mainnet, the USDC price for that model). Sign the EIP-3009 authorization, retry with a PAYMENT-SIGNATURE header, and you are served; the response carries PAYMENT-RESPONSE with the settlement transaction. It works on both /v1/chat/completions and /v1/messages, and any x402 client library handles the handshake for you.

Prices are flat per model with a bounded output: each model's x402_price (USD and max_output_tokens) is published on GET /v1/models, so max_tokens is clamped to that ceiling and n is fixed at 1. Smart routing (auto) and :freemodels are not payable this way; pick an explicit paid model. Set your client's spend control (most default to a low per-payment cap). For steady use, prepaid credits are simpler; you can top up in USDC there too.

Rate limits

600 requests a minute per IP and 240 a minute per key; both return 429 with retry-after. Need more? Email support@gitlawb.com.

Cost tracking

Non-streaming responses carry x-gateway-cost-usd (what this completion billed) and x-gateway-balance-usd (approximate remaining balance, cached for about 10 seconds); below $1 an x-gateway-balance-warning header appears, your cue to top up before a long run hits a 402. GET /v1/credits returns your balance, and GET /v1/usage/me your usage (?days=N, ?key_id=self, ?rollup=daily for lifetime totals). The same figures are in the console under usage and credits.

Ready to ship? Get a key, or write to us if you are stuck.