Skip to main content

Batch Inference

inference workload

Curated

Running a model over a large offline dataset. Throughput matters far more than latency, so large batches keep the GPU busy and cost per item low.

Classification

Category
inference
Compute profile
compute-bound
Scaling
batch

Infrastructure considerations

Qualitative — no fabricated numbers
Interconnect sensitivity
low
Network sensitivity
low

Curated · Memory · Weights + large batch activations; batch size traded against available memory.

Curated · Storage · High-throughput read of the input dataset and write of results.

Typical frameworks

Curated — not a catalog relationship
RayvLLMPyTorch

Relevant models

Models that run this workload

Data class

Curated · The workload taxonomy is authored qualitative reference, not an ingested source. Sensitivities and profiles are classifications, not measurements; a workload's concrete VRAM comes from the specific model it runs.

Explore this infrastructure

Connected by real relationships