Skip to main content

LLM Training

training workload

Curated

Training a large language model from scratch. Weights, gradients and optimizer state multiply memory demand; almost always multi-GPU and multi-node.

Classification

Category
training
Compute profile
compute-bound
Scaling
horizontal

Infrastructure considerations

Qualitative — no fabricated numbers
Interconnect sensitivity
high
Network sensitivity
high

Curated · Memory · Weights + gradients + optimizer state (often 3-4x the weights); sharded across GPUs.

Curated · Storage · High-throughput shared storage for datasets and frequent checkpoints.

Typical frameworks

Curated — not a catalog relationship
PyTorchDeepSpeedMegatronRay

Relevant models

Models that run this workload

Data class

Curated · The workload taxonomy is authored qualitative reference, not an ingested source. Sensitivities and profiles are classifications, not measurements; a workload's concrete VRAM comes from the specific model it runs.