Batch Inference
inference workload
Running a model over a large offline dataset. Throughput matters far more than latency, so large batches keep the GPU busy and cost per item low.
Classification
- Category
- inference
- Compute profile
- compute-bound
- Scaling
- batch
Infrastructure considerations
Qualitative — no fabricated numbers- Interconnect sensitivity
- low
- Network sensitivity
- low
Curated · Memory · Weights + large batch activations; batch size traded against available memory.
Curated · Storage · High-throughput read of the input dataset and write of results.
Typical frameworks
Curated — not a catalog relationshipRayvLLMPyTorch
Relevant models
Models that run this workloadData class
Curated · The workload taxonomy is authored qualitative reference, not an ingested source. Sensitivities and profiles are classifications, not measurements; a workload's concrete VRAM comes from the specific model it runs.
