Liquid Inference is an LLM router built around a single promise: the lowest price for every prompt. You send an ordinary chat request and providers compete to answer it, so your request is completed at the lowest price that still satisfies the requirements you set. It is aimed at developers, engineers and teams who reach large language models through an API, from someone wiring up a single coding agent to organisations pushing large volumes of inference, and its purpose is to turn model access into a competitive market rather than a fixed rate card. Signing up is free with an email address, and you receive an API key to use with one base URL.
The problem it addresses is that inference is normally bought at list price. Conventional access means choosing a vendor and accepting whatever rate card that vendor publishes, even though the underlying capacity is supplied by many providers whose costs and spare capacity change constantly. Providers drop their prices when they have spare capacity, but under a published rate card that movement never reaches the buyer. Liquid Inference moves that dynamic into the request path. Instead of negotiating contracts, you let providers bid for each individual prompt and you pay only for the tokens the model produces. Because routing rules constrain model, region, speed and minimum quality, the competition happens strictly inside boundaries you define.
The mechanism is an instant auction for every request, described on the site in three steps. First, providers set their prices: each provider sets a price per token for each model and can change it at any time. Second, the cheapest match answers: you set the rules, covering model, region, speed and minimum quality, and the cheapest provider that meets them answers your request. Third, you pay for what you use: before the model starts you know the most a request can cost, and you are billed only for the tokens it produces, with the cap fixed before the first token. The practical effect is that cost exposure is known in advance instead of being discovered on an invoice later.
Routing is yours to control. You choose the model, region, speed and minimum quality, and only providers that meet those requirements compete, which means the auction can never push a request outside your constraints. The site also supports routing rule presets and Auto routing algorithms for teams that would rather not define every rule by hand. Alongside routing sits a price limit on every request: the maximum a request can cost is fixed before the model starts. Together, the two controls make cheapness one dimension of the decision rather than the only one, since quality, geography and latency stay in your hands.
The catalogue is broad. Liquid Inference covers hundreds of open- and closed-weight models with full multi-modal support. The site lists model families with their counts and lowest current offers, including GPT with 55 models, Qwen with 53, Gemini with 46, GLM with 31, Deepseek with 28, Gemma with 25, Kimi with 18, Grok with 16, Claude with 16, Mistral with 14, O with 9 and Nemotron with 8, alongside Ernie, Mimo, Llama, Muse, Seed, Doubao, Minimax, Aion, Phi, Nova, Command, Mercury, Hermes, Longcat, Ling, Ministral, Step, Hunyuan, Claw, Fugu, Nex, Reka, Granite, Synth, Laguna, Agnes, Schematron, Solar, Unslopnemo, Remm, Weaver, Sonar, Morph, Inkling and Relace. A live offer board shows the cheapest offers on the book, headed in the example shown by Nemotron: Nano 9B V2 at 0.01 and DeepSeek V3.1 at 0.02 USD per million output tokens.
Compatibility is deliberately conservative: one key covers OpenAI and Anthropic APIs. OpenAI Chat Completions and Anthropic Messages are served on one base URL, so your existing code works unchanged. In practice you change the base URL in any client that speaks either API and keep everything else the same; the documented edit is a single line pointing base_url at the router with your Liquid API key. That means the router can be adopted without a rewrite, and the same key serves both API dialects.
Because the interface is standard, the product plugs into the AI coding tools teams already use. The site lists Claude Code, Codex, Pi, Oh My Pi, Cursor, OpenCode, Cline, Roo Code, Kilo Code, Goose, Open WebUI, Cherry Studio and n8n as supported clients, together with the OpenAI SDK, Anthropic SDK, Claude Agent SDK, OpenAI Agents SDK, Vercel AI SDK, LiteLLM, LangChain and curl. Fully compatible with agentic coding tools and full multi-modal support are both stated up front, so an agent that already speaks the OpenAI or Anthropic API can be pointed at the router without changes to its internal logic.
Transparency sits alongside the pricing model. Billing is itemized, so you can see what every request cost, line by line, and reconcile spend per request rather than per invoice. Under market data, Liquid Inference publishes a public price history: you can download live and past prices for every model and send work when prices are low. On the provider-trust side, an architect tests every provider regularly with standard benchmarks and publishes the results. Customers see a provider's price, speed and measured quality, and the same rules apply to every provider.
The platform is two-sided. If you already run models, you can sell inference on them: your software posts a price and changes it whenever your costs do, and when you win a request you are paid for the tokens you serve. Providers set their own price with no fixed rate card, can change it as often as they like, control load by stopping quotes when full, and lower price when they have spare capacity. Winning is on price and quality, since customers see price, speed and measured quality under the same rules. Records are clear, with every job and payment checkable against the provider's own logs. Registration takes four steps: create an account, register the deployment you want to serve, an operator reviews it and lists it, and your agent connects and starts quoting. Registration is free, and Liquid Inference charges its fee to the customer rather than to the provider.
For the buyer, the outcome is straightforward: competitive lowest marginal cost on each prompt, a hard price ceiling known before the first token, and the freedom to set constraints on model, region, speed and quality. You get one key and one base URL across OpenAI and Anthropic APIs, itemized records of what each request cost, and public price history that lets you schedule flexible work for the moments when providers are cheapest. Free signup with an email address lowers the barrier to testing, and the first 500 users receive $20 of free inference.
The use cases follow directly from those mechanics. Agentic coding tools such as Claude Code, Codex, Cursor, Cline and OpenCode can be repointed at the router so each prompt is served by the cheapest provider that meets your rules, with multi-modal requests handled by the same key. Teams planning a budget can pick a model and a monthly output volume of 1M, 10M, 100M or 1B tokens and read the estimated cost per month from the lowest offer on the book, remembering that input tokens are priced separately. Cost-sensitive batch work can be timed against the public price history. Workloads with regional, latency or quality requirements can specify them and let the market satisfy them.
Pricing is usage-based: you are billed per token produced, with input tokens priced separately, and there is no rate card to accept. Signing up is free with email, and the first 500 users get $20 of free inference; referrals earn 20 percent of referred fees as free inference and 10 percent for second-level referrals. On the supply side, registration is free for providers and the platform's fee is charged to the customer. The product is delivered as a web sign-up and API service reached through a single base URL, https://router.inference.ai.exchange/v1, with integrations spanning coding agents, chat UIs, automation tools, SDKs and frameworks.
In summary, Liquid Inference turns AI inference into an auction. Providers post and revise prices per token, your routing rules decide who is eligible, the cheapest eligible provider answers each prompt, and a fixed cap tells you the most a request can cost before the model starts. You pay only for the tokens produced, see the cost of every request line by line, and can join the market from the other side to sell capacity on the models you already run. The value proposition is the one in the title: the lowest price for every prompt.