Every model you can run here, priced per million tokens, and 3.2M more to read about, download or deploy to a GPU. One key, and one conversation that follows you when you switch.
Routes each request to the cheapest model that can handle it — one that sees images if you send one, calls tools if you send tools, and fits your conversation — and moves up to a stronger model if that one fails. No fixed price: you pay for whichever models ran, at their listed prices, capped by max_spend (default $0.10 per request). Without a verify schema it escalates on errors and refusals only, not on answer quality.
Every model you can run here, priced per million tokens, and 3.2M more to read about, download or deploy to a GPU. One key, and one conversation that follows you when you switch.
Routes each request to the cheapest model that can handle it — one that sees images if you send one, calls tools if you send tools, and fits your conversation — and moves up to a stronger model if that one fails. No fixed price: you pay for whichever models ran, at their listed prices, capped by max_spend (default $0.10 per request). Without a verify schema it escalates on errors and refusals only, not on answer quality.
Searching the library…