Architect Financial Technologies launched Liquid Inference a real‑time exchange that runs an auction for every LLM request Developers keep their existing OpenAI or Anthropic API calls only swap the base URL and let multiple providers bid to serve each prompt The buyer sets rules such as cost caps latency limits zero data retention region restrictions or allow‑lists Providers that meet the rules compete on price the lowest qualifying offer wins and the maximum price is locked before the first token is generated Billing is metered usage only with no hidden markups
For buyers this means predictable pricing instant access to spare GPU capacity and full control over data handling without managing individual contracts Live order books show quotes and cleared trades giving market transparency rare in LLM APIs The platform offers free email signup with the first 500 users receiving $20 of inference credit and a referral program that returns 20% of referred fees plus 10% on second‑level referrals
Providers onboard in minutes through a simple app register models via REST or WebSocket and can adjust quotes based on their own costs Payouts go through Stripe with itemized records for every job allowing them to monetize idle GPU cycles only when they choose
Liquid Inference differs from other routers by using a per‑request auction instead of weighted load balancing supporting both OpenAI and Anthropic formats offering detailed market data and preserving zero‑data‑retention and regional controls It solves the pain points of cost uncertainty vendor lock‑in and operational overhead for teams building agentic coding tools autonomous agents or any application that calls large language models at scale
#AI #Product #LLM #Inference #CostSaving #Developers