OpenGatewayModels

Models. The gateway's catalog, with the rate billed right now.

Send the model id exactly as listed, or send auto and let the router pick the cheapest model expected to handle the request.

The catalog. USD per 1M tokens, billed from a credit balance.

Models in the gateway's catalog, with rates per million tokens in US dollars
ModelInputOutputContextToolsSpeed
Auto (smart routing)autoSmart routing: picks the cheapest capable model and tries the next one up on failureBilled at the rate of the model that served it
MiMo V2.5-Proxiaomi/mimo-v2.5-proGeneral large language model$0.522$1.04262kYesMedium
MiMo V2.5xiaomi/mimo-v2.5Multimodal understanding model$0.168$0.336262kYesFast
Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-liteLow-latency multimodal model, 1M context$0.30$1.801MYesFast
MiniMax M3minimax/minimax-m3Agentic reasoning model$0.36$1.44205kYesMedium
Qwen 3.7 Maxqwen/qwen3.7-maxFlagship coding model$1.50$4.50262kYesSlow
Kimi K3moonshotai/kimi-k3Frontier multimodal reasoning, 1M context$3.60$18.001MYesSlow
GLM 5.2z-ai/glm-5.2Agentic coding & reasoning, 1M context$1.68$5.281MYesSlow
Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:freeFrontier reasoning MoE, 1M context, free (rate limited)No credits needed$0$0131kNoSlow
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55bFrontier reasoning MoE, 1M context$0.72$4.32512kNoSlow
Ling 3.0 Flashinclusionai/ling-3.0-flashToken-efficient agentic MoE, 262k context$0.072$0.216262kYesFast
Tencent HY3tencent/hy3Multimodal model, 262k context$0.24$0.96262kYesMedium
Macaron V1 Tallmindai/macaron-v1-tallRouted MoL personal-intelligence model, 262k context$0.54$3.12262kYesFast
Macaron V1 Ventimindai/macaron-v1-ventiFlagship routed MoL on GLM-5.2, 1M context$1.80$5.401MYesSlow

Live from the gateway's model list, GET /v1/models, re-read every five minutes. The rate billed from a credit balance right now, promotions included.

One key for the whole catalog. Swap the base URL and ship.