An LLM routing policy is the rule that turns the model field of a request into the model and provider that actually answer it. Most products never write that rule down: the model id in the code is the policy. Once a gateway sits in front of the providers, two policies become available behind the same endpoint, and the choice between them changes how you reproduce, evaluate, attribute and debug every call. This post lays out both, when each is right, and what has to be logged so that either can be trusted.
Two policies behind one endpoint
The first policy is a pinned model id: the request names a model, and the router sends it to a provider that serves that model. On Runix Router, an explicit id pins the model, and failover stays within the providers that serve it. If the first provider fails mid-request, the retry goes to another provider of the same model, never to a different model.
The second policy is auto: the request leaves the choice to the router. On Router, auto chooses by cost, health and the policy attached to the key, and a key's allowed models are the outer boundary of any choice made on its behalf. The application sends one string and receives an answer; which model produced it is recorded in the gateway's request-level trail: who called what, when, with which model.
The two are not exclusive: the model field travels with every request, so one key can pin on one route and send auto on another. What matters is that each route chooses deliberately, because the trade-offs run in opposite directions.
curl https://api.router.runixcloud.io/v1/chat/completions \
-H "Authorization: Bearer $RUNIX_API_KEY" \
-d '{"model": "auto", "messages": [{"role": "user", "content": "..."}]}'
curl https://api.router.runixcloud.io/v1/chat/completions \
-H "Authorization: Bearer $RUNIX_API_KEY" \
-d '{"model": "MODEL_ID", "messages": [{"role": "user", "content": "..."}]}'
When a pinned model id is right
Pin when the output depends on the model in a way you have measured: prompts tuned against one model's habits, structured output validated against one model, tool-calling formats, and anything a regulator or a customer contract requires you to name. An evaluation harness is pinned too, because an evaluation that does not name the model is measuring the router.
Pinning the model does not pin the provider. The same model is often served by more than one provider, first-party and through clouds that resell it, and those providers differ in region, data terms, latency and version lag. If your data terms depend on where the request runs, pin that in the key's routing policy as well, not only in the model string.
A pinned id also has an expiry date. Providers retire ids on a published schedule, and an id written as a literal in your code turns their calendar into your deploy calendar. Keep the pin in configuration the gateway reads, not in the application binary.
When auto is right
Send auto when several models clear your quality bar and the differences between them are cost and availability. Classification, extraction into a validated schema, summarisation of internal text and background enrichment usually qualify. For this traffic the question of which model has no product answer, so letting cost and health decide is the honest policy, not the lazy one.
auto can also widen the failover set, where the router's policy allows it. A pinned request can only move between providers of one model; a router that chose the model for an auto request may choose again when every provider of the first is rate-limiting or down, within the key's allowed models. Whether a router does this is a question to ask it; Runix documents failover as provider to provider. The price is that the model under a route can change without a deploy of yours; the rest of this post is about living with that.
What changes with an LLM routing policy
Four things look different depending on the policy a route uses. None is a reason to avoid either policy; each is a reason to record one more field.
Reproducibility
Under a pinned id you can re-run a request against the same model, subject to the provider's own snapshot changes behind that id. Under auto the request alone does not determine the model, so the model field of the chat completion response, not the request, is the ground truth for what ran, provided the router fills it with the answering model rather than echoing auto; check that once. Store both, on every request, before you need them.
Evaluation
A pinned route is evaluated against one model. An auto route is evaluated against the whole set it may choose from, and that set changes whenever a model is added to the key's allowed list or to the catalogue. Treat an addition as a model change, and watch the signals that move when a model change makes things worse before and after it.
Cost attribution
With a pinned id the rate is known from the catalogue before the request is sent. With auto the rate depends on the router's billing rule: Runix Router bills an auto request at the highest list price among the models the key may choose from, as the catalogue states; another router may bill at the answering model's price. Either way the ledger must record the answering model and the rate applied at call time, under the rule the router actually applies, not a price looked up later.
Incident response
A provider incident on a pinned route is your incident unless another provider serves that model. On an auto route the router moves traffic away, which keeps users unaware and, unless you watch for it, keeps you unaware too. The signal is the share of requests answered by each model, per key, per hour: when it shifts without a change of yours, a provider is having a bad hour on your behalf.
Keep both honest with request ids and logs
Both policies are only auditable if one record per request carries the whole story. Runix Router returns a request id on every response and an OpenAI-style error envelope on failures; that id is the join key between your application log and the gateway's ledger. The record worth keeping:
{
"request_id": "...",
"key_id": "...",
"model_requested": "auto",
"model_answered": "MODEL_ID",
"provider": "PROVIDER",
"attempts": 2,
"usage": {"prompt_tokens": N, "completion_tokens": N},
"finish_reason": "stop"
}
Two alerts fall out of it. On pinned routes, alert on attempts per successful request, because a rising count means failover is working and the pinned model's providers are degrading. On auto routes, alert on the model mix, because a shift is the only visible trace of a decision the router made for you. What to log for LLM traffic covers the fields to redact, and cost attribution at call time covers why the rate has to be stored with the record.
Whatever the policy, the application should never learn which provider answered; that is the point of a gateway. The logs are where that knowledge lives, and the request id is how an engineer gets it back at three in the morning.
Runix Router is in early access. It accepts auto or an explicit model id on every request, applies allowed models and routing per key, and returns a request id on every response; the request shape is in the Router quickstart. There is no published SLA during early access; the failover mechanism is documented rather than promised.
Questions this raises
Does a pinned model id also pin the provider?
No. A pinned id fixes which model answers, and several providers can serve the same model, so failover can move between them. Pin the provider in the key's routing policy as well if your data terms or your evaluation depend on it.
Can auto and a pinned model id be mixed in one application?
Yes, per request. The model field is sent with every call, so a product can pin the model on its user-facing route and send auto on background work under the same key.
What has to be logged to audit an auto decision later?
The request id, the model string you sent, the model and provider that answered, the number of attempts and the usage block, all on one record with the rate applied at call time.