Docs. One OpenAI- and Anthropic-compatible endpoint across providers.
If your client speaks the OpenAI API, it speaks OpenGateway: swap the base URL, set your key, and pick a model. The reference below covers keys, streaming, tools, routing, errors, limits and cost tracking.
Quickstart. One request, three clients.
Base URL https://opengateway.gitlawb.com/v1, and your key as Authorization: Bearer ogw_live_…. Create keys in the console. Model ids are exact: the model catalog has copy-paste ids, live pricing and capabilities.
curl https://opengateway.gitlawb.com/v1/chat/completions \
-H "authorization: Bearer ogw_live_…" \
-H "content-type: application/json" \
-d '{
"model": "mimo-v2.5-pro",
"messages": [
{"role": "user", "content": "hello, gateway"}
]
}'Reference. Everything the endpoint does, topic by topic.
Authentication
Every /v1/* request needs a key in the Authorization header. Keys look like ogw_live_ plus 32 hex characters; only a hash is stored server-side, and you can hold up to 20 active keys, one per app or environment, each revocable on its own. Requests without a valid key get 401 api_key_required.
Anthropic-compatible endpoint
The gateway also speaks the Anthropic Messages API: POST /v1/messages on https://opengateway.gitlawb.com. Point any Anthropic SDK (or Claude Code's ANTHROPIC_BASE_URL) at the gateway with your ogw_live_ key; the SDK's native x-api-key header is accepted. Any claude-* model id routes through smart routing (billed at the serving model's rate; the x-gateway-served-model header names what answered), or pass one of the catalog's exact ids to pin a model. max_tokens is required, per the Anthropic spec; streaming, tool use and system prompts are translated. count_tokens is not supported.
Models and ids
Send the model string exactly as listed in the catalog: full provider/model ids (MiMo also accepts the bare mimo-v2.5-pro short form). Ids ending in :free bill $0 and work with a zero credit balance, rate limited upstream. Retired free ids alias to their paid siblings, so old configs keep working, billed at the paid rate.
Streaming
Set stream: true for standard OpenAI server-sent events. The gateway opens the stream immediately and, while a slow reasoning model thinks, emits keepalive chunks: ordinary chat.completion.chunk events with an empty delta and ids like chatcmpl-gateway-keepalive-…. OpenAI SDKs handle them; hand-rolled parsers should tolerate zero-delta chunks (don't treat the first chunk as the first token).
Upstream failures mid-stream arrive as a data: {"error": …} event followed by [DONE]. Reasoning models want max_tokens of 300 or more; smaller budgets can be consumed entirely by thinking.
Tool calling
Standard OpenAI tools and function calling pass straight through on supporting models. The catalog says which (so does supports_tools on GET /v1/models). Smart routing never sends a request with tools to a model whose tool support is unverified.
Smart routing
Send model: "auto" (or gitlawb/auto) and the gateway scores your request, starts on the cheapest capable model, and escalates on 429, 5xx and timeouts; you are billed at the serving model's rate. Steer it with a route object in the body: priority is cost, balanced or quality, and max_cost_usd sets a price ceiling, as below. The x-gateway-served-model response header names the model that answered, and your decision log is at console, routing.
{
"model": "auto",
"route": { "priority": "cost", "max_cost_usd": 0.01 },
"messages": [ … ]
}Errors
| Status | Code | Meaning |
|---|---|---|
| 401 | api_key_required | The Authorization header is missing. |
| 401 | api_key_invalid | The key is malformed, unknown or revoked. |
| 402 | insufficient_credits | The balance is at or below zero on a paid model. Top up, or use a :free model. |
| 402 | free_quota_exceeded | The free window is used up; the body includes the reset time. |
| 429 | gateway_rate_limit_exceeded | A per-IP or per-key limit was hit. Honour retry-after. |
| 5xx | upstream_unavailable | A provider-side failure. Safe to retry with backoff. |
Error bodies are OpenAI-shaped ({"error": {…}}), so SDK error handling works unchanged.
Free usage
Two ways to use the gateway with a zero credit balance: :free models (currently Nemotron 3 Ultra, $0, rate limited upstream), and a free allowance on xiaomi/mimo-v2.5-pro of 30 requests per 5-hour window (windows reset at 00, 05, 10, 15 and 20 UTC), subject to a shared daily budget. When a window is used up you get 402 free_quota_exceeded with the reset time; topping up credits removes the cap.
Pay per request with USDC (x402)
Agents can pay per request in USDC on Base, with no sign-up and no API key. Send a completion request with no credential header; the gateway replies 402 with a PAYMENT-REQUIRED header (the standard x402 challenge: exact scheme, Base mainnet, the USDC price for that model). Sign the EIP-3009 authorization, retry with a PAYMENT-SIGNATURE header, and you are served; the response carries PAYMENT-RESPONSE with the settlement transaction. It works on both /v1/chat/completions and /v1/messages, and any x402 client library handles the handshake for you.
Prices are flat per model with a bounded output: each model's x402_price (USD and max_output_tokens) is published on GET /v1/models, so max_tokens is clamped to that ceiling and n is fixed at 1. Smart routing (auto) and :freemodels are not payable this way; pick an explicit paid model. Set your client's spend control (most default to a low per-payment cap). For steady use, prepaid credits are simpler; you can top up in USDC there too.
Rate limits
600 requests a minute per IP and 240 a minute per key; both return 429 with retry-after. Need more? Email support@gitlawb.com.
Cost tracking
Non-streaming responses carry x-gateway-cost-usd (what this completion billed) and x-gateway-balance-usd (approximate remaining balance, cached for about 10 seconds); below $1 an x-gateway-balance-warning header appears, your cue to top up before a long run hits a 402. GET /v1/credits returns your balance, and GET /v1/usage/me your usage (?days=N, ?key_id=self, ?rollup=daily for lifetime totals). The same figures are in the console under usage and credits.