Sizing
Sizing with the calculation shown
A recommended GPU type and count for your models and targets, with the memory calculation and every assumption written out.
AI transformation · Model serving
Open models give you control over data, cost, and model choice, but their memory and latency depend on model size, precision, context length, and concurrency. Silex sizes, serves, and benchmarks open models on vLLM and Red Hat OpenShift AI, from single-GPU models to frontier open-weight models on NVIDIA Blackwell systems, and shows you the calculation behind every GPU we recommend.
The sizing math
Agentic operations on Blackwell
Silex engineers have designed and tested deployments on NVIDIA B200 and B300 systems, sizing self-hosted GLM-5.3, GLM-5.3-Flash, and Kimi K3 for agentic operations use cases such as triage, remediation, and patching.
Agents do not send one request per ticket. Silex sizes an agentic workload from the task, so the plan accounts for every model call an agent makes while it investigates, drafts, and verifies.
The memory math
A dense model shows the calculation most clearly. For Llama 3.3 70B, weights and KV cache set the memory requirement, and FP8 halves both, so the same workload fits on half the GPUs.
Three 80 GB GPUs hold 240 GB, which is less than the 243 GB total, so the BF16 deployment needs four. FP8 weights and an FP8 KV cache bring the total to about 122 GB, which fits on two.
Assumptions: Llama 3.3 70B has 80 layers, 8 KV heads, and head dimension 128, which gives 327,680 bytes of KV cache per token at BF16. The load is 64 concurrent sequences of 4,096 tokens. Bars are drawn to scale, and all figures are approximate.
Sizing method
Every sizing starts from your models and your service targets, and ends with a benchmark on the recommended hardware.
We start from your models, input and output lengths, peak concurrency, and targets for time to first token and inter-token latency.
We calculate weights plus KV cache at BF16, FP8, and 4-bit precision from each model’s layer count, KV heads, and head dimension.
We choose the GPU type, tensor parallelism, and replica count that hold the memory total and meet the latency targets.
We load-test the deployment to find the highest request rate that still meets your targets, and confirm the replica count from that result.
GPU type and count, servers, and platform software, with the memory calculation and assumptions that justify each line.
Latency and throughput at each load level, with the scripts and raw data so your team can run the benchmark again.
Small and large models
One GPU per replica
These models fit on one GPU, so sizing depends on concurrency, context length, and replica count. Silex sets the replica count by dividing peak demand by the load one replica sustains in a benchmark.
Tensor parallelism in one server
These models are split across several GPUs in one server with tensor parallelism. FP8 often halves the GPU count, as the worked example above shows.
Every expert held in memory
Memory must hold every expert, even though each token uses only a few. gpt-oss-120b fits on one 80 GB GPU in MXFP4, while a 671B model at FP8 needs about 671 GB for its weights alone.
Serving stack
Silex configures every layer against the same targets, so the memory plan, the serving engine, the autoscaler, and the gateway quotas agree.
Evidence
Silex engineers have designed and tested deployments on NVIDIA B200 and B300 systems, sizing GLM-5.3, GLM-5.3-Flash, and Kimi K3 for agentic operations workloads of more than 1,000 tasks per day at a target GPU utilization of 60 percent or higher.
Silex published a benchmark of Llama 3.3 70B on four-GPU A100 and H100 systems, and the scripts are public.
Read the benchmarkIn Silex’s gateway demonstration environment, a Kong AI gateway routes to a self-hosted gpt-oss-120b alongside hosted models.
Silex has served vLLM inside a Triton ensemble on Kubernetes GPU nodes, and its hands-on workshops serve open models with vLLM.
What we deliver
Sizing
A recommended GPU type and count for your models and targets, with the memory calculation and every assumption written out.
Serving
Models served on Red Hat OpenShift AI and Red Hat AI Inference with vLLM, KServe, and llm-d, deployed and configured as code.
Tuning
FP8, FP4, and INT4 variants, speculative decoding, and prefix caching, with an accuracy check after each change.
Benchmarks
Load tests against your time to first token and inter-token latency targets, delivered with the scripts and raw data.
Operations
Autoscaling on queue depth and KV cache use, multi-model routing, and Models-as-a-Service quotas through the AI gateway.
Hardware
Dell and Supermicro servers with NVIDIA or AMD GPUs, supplied by Silex and sized from the same calculation.
Platforms
Premier Partner
Platinum Partner
GPU servers
Certified Delivery Partner
vLLM · Red Hat OpenShift AI · Red Hat AI Inference · KServe · llm-d · LLM Compressor · GuideLLM · KEDA · Ollama · llama.cpp · NVIDIA Blackwell (B200, B300) and AMD GPUs
Questions buyers ask
We start from the tasks, not the requests. For each use case we estimate tasks per day, model calls per task, and tokens per call, find the peak token rate, and size the GPUs to run it at the target utilization while meeting latency targets. Silex has sized GLM-5.3, GLM-5.3-Flash, and Kimi K3 this way on NVIDIA B200 and B300 systems for workloads of more than 1,000 tasks per day at 60 percent or higher utilization.
It depends on precision, context length, and concurrency. For Llama 3.3 70B serving 64 concurrent 4,096-token sequences, BF16 weights take about 141 GB and the KV cache about 86 GB. With about 16 GB of overhead, the total of about 243 GB needs four 80 GB GPUs. FP8 weights and an FP8 KV cache bring the total to about 122 GB, which fits on two.
FP8 results are usually close to BF16. Silex tests 4-bit variants on your own tasks before recommending them and checks accuracy after every quantization change.
Many tasks do not need one. Models in the 8B to 30B range handle many tasks well, so Silex evaluates candidate models on your tasks before sizing any hardware.
It can. vLLM serves many fine-tuned LoRA adapters on one base model and selects the adapter for each request, and several small models can share one GPU.
They can. The AI gateway routes requests to self-hosted models and to hosted providers such as Amazon Bedrock and Azure OpenAI, and applies the same authentication, token quotas, guardrails, and audit logging to both.
NVIDIA GPUs are not required. vLLM and Red Hat AI Inference also run on AMD Instinct GPUs, and Silex sizes deployments for either.
Next step
A GPU and platform sizing for your target models, context lengths, concurrency, and latency targets, validated with benchmarks. For hands-on practice, your team can start with the lab Serving and Sizing Open Models with vLLM on OpenShift AI.