Models
The StrataCode model catalog — GLM-5.2, GLM-4.5-Air, DeepSeek V4 Flash, Qwen3.5 Plus, plus the auto router and a free lane.
Every model on the endpoint
Every model speaks the OpenAI API and is billed per token at the rate shown. Prefer stratacode/auto unless you have a reason to pin one — it routes each request to the cheapest model that can handle it, and bills you at that model’s rate.
Catalogue
- StrataCode Auto (router) stratacode/auto
router → GLM-5.2 / GLM-4.5-Air; billed at the model that answered
- GLM-5.2 stratacode/glm-5.2
US-hosted (DeepInfra primary, Z.ai fallback)
- GLM-4.5-Air stratacode/glm-4.5-air
Novita
- DeepSeek V4 Flash stratacode/deepseek-v4-flash
DeepSeek direct (upstream wiring: R3.1)
- Qwen3.5 Plus stratacode/qwen3.5-plus
Alibaba DashScope (upstream wiring: R3.1)
- StrataCode Free stratacode/free
free lane (subsidized)
Rates
- GLM-5.2 id stratacode/glm-5.2 input $1.40 output $4.40 vs direct at list
- GLM-4.5-Air id stratacode/glm-4.5-air input $0.20 output $1.10 vs direct at list
- DeepSeek V4 Flash id stratacode/deepseek-v4-flash input $0.14 output $0.28 vs direct at list
- Qwen3.5 Plus id stratacode/qwen3.5-plus input $0.40 output $2.40 vs direct at list
- StrataCode Free id stratacode/free input $0.00 output $0.00 vs direct at list
Per million tokens, in USD. Cached input is billed at each model's published cached rate — $0.26 for GLM-5.2. These are the vendors' list prices: you never pay more here than going direct, and the router is what makes a given task cost less.
Sourcing and availability
GLM models are served through US-based inference hosts — DeepInfra primary, with failover — rather than a single origin. This is the standard, compliance-aware way to serve GLM, and it keeps a lane up if any one host degrades. Each host must pass our quality bench before it enters rotation, so a cheaper host never means quietly worse output.
We do not list models whose supply or compliance status is unsettled. Kimi is not in the catalogue. See trust and sourcing for the full policy.