AI Workload Catalog
11 workload types and the infrastructure each one needs. Qualitative, curated reference — concrete VRAM comes from the model.
- LLM InferenceServing a large language model to answer requests. Latency- and throughput-sensitive; the model weights plus the KV cache must fit in GPU memory.inferencememory-boundhorizontal
- LLM TrainingTraining a large language model from scratch. Weights, gradients and optimizer state multiply memory demand; almost always multi-GPU and multi-node.trainingcompute-boundhorizontal
- Fine-tuningAdapting a pretrained model to a task or domain. Full fine-tuning approaches training cost; parameter-efficient methods (LoRA/QLoRA) fit far smaller.trainingmixedsingle-node
- Batch InferenceRunning a model over a large offline dataset. Throughput matters far more than latency, so large batches keep the GPU busy and cost per item low.inferencecompute-boundbatch
- EmbeddingsEncoding text or other inputs into vectors. Smaller models, high throughput; often CPU-viable but GPU-accelerated at scale.datacompute-boundbatch
- RAG (Retrieval-Augmented Generation)Combining a retrieval step over a vector store with LLM generation. Infrastructure spans an embedding/retrieval path and an LLM inference path.inferencemixedhorizontal
- Image GenerationDiffusion or transformer image models. Iterative denoising is compute-heavy; VRAM scales with resolution and batch.generationcompute-boundhorizontal
- Video GenerationTemporal generative models. The most VRAM- and compute-intensive generation workload; frames add a time dimension to image generation.generationcompute-boundvertical
- Speech-to-TextTranscribing audio to text. Streaming or batch; encoder-decoder models with modest weights but latency sensitivity when real-time.audiomixedhorizontal
- Text-to-SpeechSynthesizing speech from text. Latency-sensitive for interactive use; models are typically small relative to LLMs.audiomixedhorizontal
- Computer VisionClassification, detection or segmentation over images. Well-optimized, high-throughput; a broad range of model sizes.visioncompute-boundbatch
