LLM Inference
inference workload
Serving a large language model to answer requests. Latency- and throughput-sensitive; the model weights plus the KV cache must fit in GPU memory.
Classification
- Category
- inference
- Compute profile
- memory-bound
- Scaling
- horizontal
Infrastructure considerations
Qualitative — no fabricated numbers- Interconnect sensitivity
- medium
- Network sensitivity
- low
Curated · Memory · Model weights + KV cache; grows with context length and concurrent requests.
Curated · Storage · Weights loaded once at startup; modest fast local storage for the checkpoint.
Typical frameworks
Curated — not a catalog relationshipvLLMTensorRT-LLMSGLang
Relevant models
Models that run this workloadData class
Curated · The workload taxonomy is authored qualitative reference, not an ingested source. Sensitivities and profiles are classifications, not measurements; a workload's concrete VRAM comes from the specific model it runs.
