Skip to main content

LLM Inference

inference workload

Curated

Serving a large language model to answer requests. Latency- and throughput-sensitive; the model weights plus the KV cache must fit in GPU memory.

Classification

Category
inference
Compute profile
memory-bound
Scaling
horizontal

Infrastructure considerations

Qualitative — no fabricated numbers
Interconnect sensitivity
medium
Network sensitivity
low

Curated · Memory · Model weights + KV cache; grows with context length and concurrent requests.

Curated · Storage · Weights loaded once at startup; modest fast local storage for the checkpoint.

Typical frameworks

Curated — not a catalog relationship
vLLMTensorRT-LLMSGLang

Relevant models

Models that run this workload

Data class

Curated · The workload taxonomy is authored qualitative reference, not an ingested source. Sensitivities and profiles are classifications, not measurements; a workload's concrete VRAM comes from the specific model it runs.

Explore this infrastructure

Connected by real relationships