Single-tenant model serving explained: what a dedicated LLM inference endpoint changes

A dedicated LLM inference endpoint is one where the model, the serving engine and the accelerators behind the URL serve one tenant. A shared endpoint is one where a provider multiplexes many customers across one fleet and hands each of them a rate limit. The difference sounds administrative, but it changes who controls latency, when the model changes, where the data goes, and who is on call. This post sets out each of those without figures, because every one of them depends on your traffic.

Shared endpoint, dedicated endpoint: the definitions

On a shared endpoint you send requests to a model id and the provider decides which replica answers, how your requests are batched with everyone else's, and when the model behind the id is updated or retired. You pay per token, which means idle costs nothing and your peak competes with other customers' peaks.

On a dedicated endpoint, capacity is reserved: specific accelerators run specific weights for you alone. You pay for the capacity whether or not it is busy, and your requests compete only with your own. Single-tenant is the same arrangement described from the isolation side rather than the capacity side.

Isolation: what a dedicated LLM inference endpoint separates

Three kinds of isolation come with it, and they are worth separating because teams usually want one and get all three.

Latency under your own load, and capacity planning

Modern serving engines batch requests continuously: new sequences join the batch as others finish, and the accelerator works on all of them at once. Throughput rises with concurrency, and so does per-request latency, until memory or compute saturates and requests start to queue. Where that knee sits depends on the model, the hardware, and the shape of your prompts and outputs.

Capacity planning is the exercise of finding that knee for your traffic before your users do. The inputs are concrete:

The last point is where a dedicated endpoint differs most from a shared one. On a shared endpoint the provider's rate limits are the backpressure; on your own endpoint there are none until you set them, and a client that retries hard can take the endpoint down with its own traffic. Set per-client limits and timeouts as deliberately as you would for a hosted provider.

Quantised variants and the batch you actually run

A quantised variant stores the weights, and sometimes the key-value cache, at lower precision. It needs less memory, so a larger batch fits on the same hardware and throughput per accelerator usually rises, although dequantisation has a cost of its own once the batch is large enough to be compute-bound. It is also a different model from the full-precision one, and the difference shows up unevenly: some tasks are unaffected, others degrade in ways a generic benchmark does not catch.

The rule that follows is to evaluate the exact variant you serve on your own held-out set, and serve it only if quality holds. The memory saving is real; whether it is free is a question your evaluation answers, not the quantisation method's documentation.

Where it sits: an OpenAI-compatible API, optionally behind a gateway

Serving engines such as vLLM and Text Generation Inference expose an HTTP API shaped like the OpenAI chat completions endpoint, so an application that already calls a hosted model can call the dedicated one by changing a base URL and a model id. Since the compatibility surface is not a contract, test the parameters you rely on, the usage fields and the way streams end before you switch.

Putting a gateway in front adds one more hop and three things that hop can do: present hosted and dedicated models at one endpoint, fail over from the dedicated model to a hosted one when it is down or over capacity, and apply per-key limits and routing without touching the application. Runix Router is designed for that position, and a Runix Models endpoint can sit behind it next to hosted providers.

What it costs you in operations

A dedicated endpoint moves work from the provider's operations team to yours. The list is not exotic, but it is yours:

None of this argues against dedicated serving. It argues for choosing it when the data, the volume or the control actually require it, and for counting the operations when you do.

Runix Models serves tuned open-weight models single-tenant behind an OpenAI-compatible API, in your cloud account or on capacity arranged per engagement, optionally behind Runix Router, with quantised variants only where the evaluation shows quality holds. It is in early access, by engagement.

Questions this raises

Is a dedicated endpoint always lower latency than a shared one?

Not inherently. It makes latency a function of your own load instead of everyone else's, which makes it predictable and plannable. Under-provisioned dedicated capacity queues just like a busy shared endpoint.

Are rate limits still needed on a dedicated endpoint?

Yes. A shared provider's limits are the backpressure you lose when you run your own; without per-client limits and timeouts, one retrying client can saturate the endpoint.

Can a dedicated model and hosted models share one API?

Yes, if the serving engine exposes an OpenAI-compatible API and a gateway presents both at one endpoint. Test the compatibility surface before switching traffic.

Related to this post: Runix Models. Tell us what you are building and we reply within one business day.