Skip to content

AI transformation · Model serving

Model Serving and Inference Sizing

Open models give you control over data, cost, and model choice, but their memory and latency depend on model size, precision, context length, and concurrency. Silex sizes, serves, and benchmarks open models on vLLM and Red Hat OpenShift AI, from single-GPU models to frontier open-weight models on NVIDIA Blackwell systems, and shows you the calculation behind every GPU we recommend.

The sizing math

KV cache per token
2 × layers × KV heads × head dimension × bytes per element
KV cache total
per-token size × tokens per sequence × concurrent sequences
GPU memory
weights + KV cache total + overhead

Agentic operations on Blackwell

Sized for agentic operations on NVIDIA Blackwell.

Silex engineers have designed and tested deployments on NVIDIA B200 and B300 systems, sizing self-hosted GLM-5.3, GLM-5.3-Flash, and Kimi K3 for agentic operations use cases such as triage, remediation, and patching.

Workload1,000+Tasks per day across the agentic operations use cases we sized. Each task makes many model calls with long contexts.
Target utilization60%+GPU utilization target, so the hardware you buy spends its time on work while agents still meet their latency targets.
HardwareB200/B300NVIDIA Blackwell systems, sized against the same targets as every other deployment we recommend.
Models sizedGLM-5.3GLM-5.3-FlashKimi K3

From tasks to GPUs.

Agents do not send one request per ticket. Silex sizes an agentic workload from the task, so the plan accounts for every model call an agent makes while it investigates, drafts, and verifies.

  1. 01Tasks per dayCounted by use case: triage, remediation, patching, and requests.
  2. 02Calls per taskThe model calls an agent makes to investigate, draft, and verify.
  3. 03Tokens per callInput and output tokens, including tool results and long context.
  4. 04Peak token rateTokens per second in the busiest hour, when work arrives in bursts.
  5. 05GPUs at the targetThe GPU type and count that run the peak at 60 percent or higher utilization and meet latency targets.

The memory math

How precision changes the GPU count.

A dense model shows the calculation most clearly. For Llama 3.3 70B, weights and KV cache set the memory requirement, and FP8 halves both, so the same workload fits on half the GPUs.

BF16About 243 GB in totalFour 80 GB GPUs, tensor parallelism 4
Weightsabout 141 GBKV cacheabout 86 GB
GPU 180 GB
GPU 280 GB
GPU 380 GB
GPU 480 GB
FP8About 122 GB in totalTwo 80 GB GPUs, tensor parallelism 2
Weightsabout 71 GBKV cacheabout 43 GB
GPU 180 GB
GPU 280 GB
Model weightsKV cacheOverhead

Three 80 GB GPUs hold 240 GB, which is less than the 243 GB total, so the BF16 deployment needs four. FP8 weights and an FP8 KV cache bring the total to about 122 GB, which fits on two.

Assumptions: Llama 3.3 70B has 80 layers, 8 KV heads, and head dimension 128, which gives 327,680 bytes of KV cache per token at BF16. The load is 64 concurrent sequences of 4,096 tokens. Bars are drawn to scale, and all figures are approximate.

Sizing method

We publish the method, not only the answer.

Every sizing starts from your models and your service targets, and ends with a benchmark on the recommended hardware.

  1. 01

    Inputs and targets

    We start from your models, input and output lengths, peak concurrency, and targets for time to first token and inter-token latency.

  2. 02

    Memory at each precision

    We calculate weights plus KV cache at BF16, FP8, and 4-bit precision from each model’s layer count, KV heads, and head dimension.

  3. 03

    Parallelism and GPU type

    We choose the GPU type, tensor parallelism, and replica count that hold the memory total and meet the latency targets.

  4. 04

    Benchmark

    We load-test the deployment to find the highest request rate that still meets your targets, and confirm the replica count from that result.

Bill of materials

GPU type and count, servers, and platform software, with the memory calculation and assumptions that justify each line.

Benchmark report

Latency and throughput at each load level, with the scripts and raw data so your team can run the benchmark again.

Small and large models

Model size decides where the sizing effort goes.

One GPU per replica

1B to about 30B parameters

These models fit on one GPU, so sizing depends on concurrency, context length, and replica count. Silex sets the replica count by dividing peak demand by the load one replica sustains in a benchmark.

Tensor parallelism in one server

Dense 70B to 120B parameters

These models are split across several GPUs in one server with tensor parallelism. FP8 often halves the GPU count, as the worked example above shows.

Every expert held in memory

Mixture-of-experts models

Memory must hold every expert, even though each token uses only a few. gpt-oss-120b fits on one 80 GB GPU in MXFP4, while a 671B model at FP8 needs about 671 GB for its weights alone.

Serving stack

The serving stack from GPU to gateway.

Silex configures every layer against the same targets, so the memory plan, the serving engine, the autoscaler, and the gateway quotas agree.

AI gatewayKong AI Gateway
Routing, token quotas, guardrails, and audit for self-hosted and hosted models
PlatformRed Hat OpenShift AI
Model catalog, Models-as-a-Service quotas, and access control
Kubernetes servingKServe · llm-d · KEDA
Replicas, KV-cache-aware routing, and autoscaling on queue depth
DistributionRed Hat AI Inference · LLM Compressor
A supported vLLM build and quantized model checkpoints
EnginevLLM
Continuous batching, PagedAttention, and prefix caching
AcceleratorsNVIDIA B200 and B300 · AMD GPUs · Dell · Supermicro
GPU type and count from the sizing calculation

Evidence

Work you can check.

Blackwell and agentic operations sizing

Silex engineers have designed and tested deployments on NVIDIA B200 and B300 systems, sizing GLM-5.3, GLM-5.3-Flash, and Kimi K3 for agentic operations workloads of more than 1,000 tasks per day at a target GPU utilization of 60 percent or higher.

Published benchmark, May 2025

Silex published a benchmark of Llama 3.3 70B on four-GPU A100 and H100 systems, and the scripts are public.

Read the benchmark 

Self-hosted and hosted models behind one gateway

In Silex’s gateway demonstration environment, a Kong AI gateway routes to a self-hosted gpt-oss-120b alongside hosted models.

Multi-model serving on Kubernetes

Silex has served vLLM inside a Triton ensemble on Kubernetes GPU nodes, and its hands-on workshops serve open models with vLLM.

What we deliver

Sized, served, and measured.

Sizing

Sizing with the calculation shown

A recommended GPU type and count for your models and targets, with the memory calculation and every assumption written out.

Serving

Model serving on OpenShift AI

Models served on Red Hat OpenShift AI and Red Hat AI Inference with vLLM, KServe, and llm-d, deployed and configured as code.

Tuning

Quantization and tuning

FP8, FP4, and INT4 variants, speculative decoding, and prefix caching, with an accuracy check after each change.

Benchmarks

Benchmarks against your targets

Load tests against your time to first token and inter-token latency targets, delivered with the scripts and raw data.

Operations

Autoscaling and routing

Autoscaling on queue depth and KV cache use, multi-model routing, and Models-as-a-Service quotas through the AI gateway.

Hardware

Servers and GPUs

Dell and Supermicro servers with NVIDIA or AMD GPUs, supplied by Silex and sized from the same calculation.

Platforms

What we build with.

Premier Partner

Platinum Partner

GPU servers

Certified Delivery Partner

vLLM · Red Hat OpenShift AI · Red Hat AI Inference · KServe · llm-d · LLM Compressor · GuideLLM · KEDA · Ollama · llama.cpp · NVIDIA Blackwell (B200, B300) and AMD GPUs

Questions buyers ask

Questions

How do you size GPUs for agentic operations?

We start from the tasks, not the requests. For each use case we estimate tasks per day, model calls per task, and tokens per call, find the peak token rate, and size the GPUs to run it at the target utilization while meeting latency targets. Silex has sized GLM-5.3, GLM-5.3-Flash, and Kimi K3 this way on NVIDIA B200 and B300 systems for workloads of more than 1,000 tasks per day at 60 percent or higher utilization.

How many GPUs does a 70B model need?

It depends on precision, context length, and concurrency. For Llama 3.3 70B serving 64 concurrent 4,096-token sequences, BF16 weights take about 141 GB and the KV cache about 86 GB. With about 16 GB of overhead, the total of about 243 GB needs four 80 GB GPUs. FP8 weights and an FP8 KV cache bring the total to about 122 GB, which fits on two.

Does quantization reduce quality?

FP8 results are usually close to BF16. Silex tests 4-bit variants on your own tasks before recommending them and checks accuracy after every quantization change.

Do we need a large model?

Many tasks do not need one. Models in the 8B to 30B range handle many tasks well, so Silex evaluates candidate models on your tasks before sizing any hardware.

Can one GPU serve several models or fine-tuned variants?

It can. vLLM serves many fine-tuned LoRA adapters on one base model and selects the adapter for each request, and several small models can share one GPU.

Can self-hosted and cloud models share one set of controls?

They can. The AI gateway routes requests to self-hosted models and to hosted providers such as Amazon Bedrock and Azure OpenAI, and applies the same authentication, token quotas, guardrails, and audit logging to both.

Is NVIDIA required?

NVIDIA GPUs are not required. vLLM and Red Hat AI Inference also run on AMD Instinct GPUs, and Silex sizes deployments for either.

Next step

Request an Inference Sizing Assessment.

A GPU and platform sizing for your target models, context lengths, concurrency, and latency targets, validated with benchmarks. For hands-on practice, your team can start with the lab Serving and Sizing Open Models with vLLM on OpenShift AI.